What GPU Architectures Does zeux/meshoptimizer's Vertex Cache Optimization Target?
The vertex cache optimizer in meshoptimizer is tuned for GPUs with cache profiles similar to consumer-grade NVIDIA and AMD graphics processors, assuming a maximum cache size of 16 entries.
The zeux/meshoptimizer repository provides high-performance mesh optimization algorithms for real-time rendering and game engine pipelines. Its vertex cache optimization specifically targets the cache behaviors found in modern desktop GPUs to minimize the Average Cache Miss Ratio (ACMR) on the most common consumer hardware.
Target GPU Architecture Profile
The optimizer is designed for GPUs exhibiting FIFO (first-in-first-out) vertex cache behavior characteristic of mainstream graphics hardware.
NVIDIA and AMD Consumer GPUs
According to the source code comment in src/vcacheoptimizer.cpp (lines 22-25):
"Tuned to minimize the ACMR of a GPU that has a cache profile similar to NVidia and AMD"
This explicit tuning targets the post-transform vertex cache implementations found in NVIDIA GeForce and AMD Radeon series GPUs. The algorithm assumes cache hit patterns and replacement policies that match these architectures, which represent the dominant hardware in PC gaming and professional visualization markets.
Cache Size Specifications
The implementation supports a maximum cache size of 16 entries, which corresponds to the typical vertex cache depth in modern consumer-grade GPUs. The optimizer does not target specialized architectures with substantially different cache sizes—such as legacy hardware or mobile GPUs with tile-based architectures—because these exhibit divergent memory access patterns that would require different optimization heuristics.
Implementation Details
In src/vcacheoptimizer.cpp, the meshopt_optimizeVertexCache function implements a scoring algorithm optimized for the NVIDIA/AMD cache profile. The function automatically applies hardware-specific heuristics without requiring manual platform detection.
For scenarios requiring explicit cache size configuration while maintaining the same NVIDIA/AMD-tuned optimization strategy, the library exposes meshopt_optimizeVertexCacheFifo.
Code Examples
Standard Vertex Cache Optimization
The following example demonstrates the default optimization path tuned for NVIDIA/AMD GPUs:
#include "meshoptimizer.h"
#include <vector>
int main()
{
// Example index buffer (triangle list)
std::vector<unsigned int> indices = {
0, 1, 2,
2, 3, 0,
0, 2, 4
};
// Allocate destination buffer matching source size
std::vector<unsigned int> optimized(indices.size());
// Run the optimizer – it uses the default score table tuned for NVIDIA/AMD GPUs
meshopt_optimizeVertexCache(
optimized.data(),
indices.data(),
indices.size(),
/*vertex_count=*/5
);
// optimized now contains index order with reduced ACMR
return 0;
}
Custom Cache Size Configuration
To specify a non-standard cache size while maintaining the NVIDIA/AMD optimization profile:
meshopt_optimizeVertexCacheFifo(
optimized.data(),
indices.data(),
indices.size(),
/*vertex_count=*/5,
/*cache_size=*/32); // any size >= 3; the algorithm still assumes an NVIDIA/AMD-like profile
Key Source Files
The vertex cache optimization implementation resides in specific files within the repository:
src/vcacheoptimizer.cpp: Contains the optimizer implementation and the explicit comment regarding NVIDIA/AMD cache profile targetingsrc/meshoptimizer.h: Public header declaringmeshopt_optimizeVertexCache,meshopt_optimizeVertexCacheStrip, andmeshopt_optimizeVertexCacheFifotools/vcachetuner.cpp: Benchmarking utility for evaluating cache optimization across different cache sizes
Summary
- Target hardware: NVIDIA and AMD consumer GPUs with FIFO vertex cache behavior
- Cache assumptions: Up to 16 entries maximum, matching typical GeForce and Radeon implementations
- Primary source:
src/vcacheoptimizer.cppimplements the NVIDIA/AMD-tuned scoring algorithm - API entry points:
meshopt_optimizeVertexCachefor standard use,meshopt_optimizeVertexCacheFifofor custom cache sizes - Limitations: Does not optimize for mobile tile-based architectures or legacy GPUs with divergent cache behaviors
Frequently Asked Questions
Does meshoptimizer support Intel GPUs?
The optimizer functions on Intel GPUs because modern Intel graphics implement similar FIFO vertex cache behaviors. However, the algorithm remains specifically tuned for NVIDIA/AMD cache hit patterns, which may yield suboptimal ACMR results on Intel hardware compared to the targeted architectures.
Can I use a custom cache size for specialized hardware?
Yes, call meshopt_optimizeVertexCacheFifo with your specific cache_size parameter. While this allows experimentation with different cache depths, the underlying algorithm retains its NVIDIA/AMD cache behavior assumptions and does not adapt to fundamentally different cache replacement policies found in specialized or mobile architectures.
What does ACMR mean in vertex cache optimization?
Average Cache Miss Ratio (ACMR) quantifies the efficiency of vertex reuse by measuring how frequently the GPU must re-process vertices versus accessing them from the post-transform cache. Lower ACMR values indicate better cache utilization, directly correlating with reduced memory bandwidth and improved rendering performance in geometry-bound scenarios.
Why doesn't meshoptimizer target mobile GPU architectures?
Mobile GPUs typically employ tile-based deferred rendering (TBDR) architectures with vertex processing units that behave differently from the immediate-mode rendering pipelines of desktop NVIDIA and AMD cards. The cache hit patterns, memory hierarchies, and vertex reuse characteristics in TBDR systems require different optimization strategies than the FIFO cache model implemented in src/vcacheoptimizer.cpp.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →