meshopt_optimizeVertexCache vs meshopt_optimizeVertexCacheFifo vs meshopt_optimizeVertexCacheStrip: Choosing the Right Vertex Cache Optimizer
The three functions represent different trade-offs between mesh processing speed and rendering cache efficiency—meshopt_optimizeVertexCache maximizes GPU vertex cache hits using an LRU model, meshopt_optimizeVertexCacheStrip prioritizes triangle strip length for better compression, and meshopt_optimizeVertexCacheFifo delivers the fastest preprocessing using a simplified FIFO cache approximation.
The meshoptimizer library provides three distinct algorithms for reordering index buffers to minimize vertex shader invocations on the GPU. While all three functions—declared in [src/meshoptimizer.h](https://github.com/zeux/meshoptimizer/blob/master/src/meshoptimizer.h)—share the goal of improving post-transform cache locality, they target different hardware cache behaviors and pipeline requirements, as implemented in [src/vcacheoptimizer.cpp](https://github.com/zeux/meshoptimizer/blob/master/src/vcacheoptimizer.cpp).
Understanding the Three Cache Optimization Models
meshopt_optimizeVertexCache: Maximum Quality with LRU
meshopt_optimizeVertexCache implements a classic least-recently-used (LRU) cache optimization algorithm, making it the default choice for production rendering pipelines. This function reorders indices to keep recently transformed vertices resident in the GPU cache as long as possible, minimizing redundant vertex shader executions.
According to the meshoptimizer source code, this variant provides the highest cache hit rates but requires the most computation time. The function signature expects the destination buffer, source indices, index count, and vertex count:
meshopt_optimizeVertexCache(
unsigned int* destination,
const unsigned int* indices,
size_t index_count,
size_t vertex_count);
The declaration resides at approximately line 207 of the header file.
meshopt_optimizeVertexCacheStrip: Compression-Friendly Ordering
meshopt_optimizeVertexCacheStrip modifies the cost function to prioritize triangle strip length alongside cache efficiency. This approach produces index buffers that compress better when using strip-based encoding or delta-compression algorithms, though it yields inferior cache performance compared to the LRU variant.
As noted in the source analysis, this optimizer runs roughly three times faster than the default implementation but sacrifices some vertex cache locality. The function shares the same parameters as the standard variant, declared at approximately line 216:
meshopt_optimizeVertexCacheStrip(
unsigned int* destination,
const unsigned int* indices,
size_t index_count,
size_t vertex_count);
Use this variant when your pipeline includes additional index compression steps or when you need faster preprocessing during iterative asset development.
meshopt_optimizeVertexCacheFifo: Fast FIFO Approximation
meshopt_optimizeVertexCacheFifo targets a first-in-first-out (FIFO) cache model rather than LRU, significantly simplifying the optimization heuristic. This approximation runs approximately three times faster than meshopt_optimizeVertexCache, making it ideal for build-time tools that require rapid iteration or preprocessing of massive asset libraries.
The function accepts an additional cache_size parameter, which should be set slightly smaller than the target GPU's actual cache size to prevent cache thrashing:
meshopt_optimizeVertexCacheFifo(
unsigned int* destination,
const unsigned int* indices,
size_t index_count,
size_t vertex_count,
unsigned int cache_size);
Declared at approximately line 227, this variant appears in the test suite at [tools/vcachetester.cpp](https://github.com/zeux/meshoptimizer/blob/master/tools/vcachetester.cpp), demonstrating its use for benchmarking and validation scenarios.
Performance and Quality Comparison
When selecting a vertex cache optimizer, consider the following characteristics:
meshopt_optimizeVertexCache— Best cache efficiency, slowest processing. Use for final production assets where runtime rendering performance is critical.meshopt_optimizeVertexCacheStrip— Balanced approach, ~3× faster than default. Use when index buffer compression ratio matters as much as cache performance.meshopt_optimizeVertexCacheFifo— Fastest preprocessing, ~3× faster than default with modest quality loss. Use for quick iteration during development or when processing massive datasets where optimization time dominates.
Practical Implementation Examples
The following example demonstrates how to invoke all three optimizers on the same source geometry, selecting the appropriate variant based on your pipeline stage:
#include "meshoptimizer.h"
#include <vector>
// Original triangle mesh data
std::vector<unsigned int> indices = {0, 1, 2, 2, 1, 3, /* ... */};
size_t index_count = indices.size();
size_t vertex_count = 1000; // Total unique vertices in the mesh
// Buffers for optimized results
std::vector<unsigned int> optimized_lru(index_count);
std::vector<unsigned int> optimized_strip(index_count);
std::vector<unsigned int> optimized_fifo(index_count);
// 1. Production quality: Maximum cache efficiency
meshopt_optimizeVertexCache(
optimized_lru.data(),
indices.data(),
index_count,
vertex_count);
// 2. Compression-oriented: Better strip generation
meshopt_optimizeVertexCacheStrip(
optimized_strip.data(),
indices.data(),
index_count,
vertex_count);
// 3. Fast iteration: FIFO approximation with 16-entry cache
meshopt_optimizeVertexCacheFifo(
optimized_fifo.data(),
indices.data(),
index_count,
vertex_count,
16);
Key Source Files
The implementation and usage of these optimizers span several critical files in the repository:
- [
src/meshoptimizer.h](https://github.com/zeux/meshoptimizer/blob/master/src/meshoptimizer.h) — Public API declarations for all three optimization functions (lines 207, 216, and 227 respectively). - [
src/vcacheoptimizer.cpp](https://github.com/zeux/meshoptimizer/blob/master/src/vcacheoptimizer.cpp) — Core implementation containing the LRU, Strip, and FIFO optimization algorithms. - [
tools/vcachetester.cpp](https://github.com/zeux/meshoptimizer/blob/master/tools/vcachetester.cpp) — Test harness specifically exercising the FIFO variant and benchmarking cache performance. - [
demo/main.cpp](https://github.com/zeux/meshoptimizer/blob/master/demo/demo/main.cpp) — Real-world demonstration showing integration into a complete mesh processing pipeline.
Summary
- Use
meshopt_optimizeVertexCachefor shipping assets where maximum GPU vertex cache efficiency is required, accepting slower offline processing. - Use
meshopt_optimizeVertexCacheStripwhen you need faster preprocessing and your pipeline benefits from improved triangle strip continuity for compression. - Use
meshopt_optimizeVertexCacheFifofor rapid iteration during development or when processing massive asset libraries, tuning thecache_sizeparameter to match your target hardware.
Frequently Asked Questions
Which vertex cache optimizer should I use for real-time game rendering?
For runtime game assets, use meshopt_optimizeVertexCache as your default choice. The LRU-based algorithm provides the highest post-transform cache hit rates, directly reducing vertex shader invocations and improving frame rates. Only switch to the FIFO variant if profiling shows that mesh preprocessing time creates a bottleneck in your asset build pipeline.
How does the cache_size parameter affect meshopt_optimizeVertexCacheFifo?
The cache_size parameter defines the number of entries in the simulated FIFO cache. You should set this value slightly smaller than your target GPU's actual vertex cache size—typically 16 to 24 entries—to prevent the optimizer from assuming it can retain more vertices than the hardware actually supports, which would cause cache thrashing and degrade performance.
Can I use meshopt_optimizeVertexCacheStrip for general geometry optimization?
While meshopt_optimizeVertexCacheStrip produces valid cache-friendly orderings, it generates inferior cache hit rates compared to the standard meshopt_optimizeVertexCache function. Reserve this variant for specific scenarios where you subsequently compress the index buffer using strip-based encoding schemes, as the improved strip continuity can yield better compression ratios that offset the modest loss in cache efficiency.
What is the actual performance difference between the three optimizers?
According to the meshoptimizer implementation, both meshopt_optimizeVertexCacheStrip and meshopt_optimizeVertexCacheFifo execute approximately three times faster than the standard meshopt_optimizeVertexCache function. However, the standard LRU optimizer produces superior vertex cache locality, resulting in fewer vertex shader invocations during actual GPU rendering, which typically outweighs the offline processing cost for final assets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →