# What GPU Architectures Does zeux/meshoptimizer's Vertex Cache Optimization Target?

> Discover which GPU architectures zeux/meshoptimizer's vertex cache optimization targets. Optimize your graphics for NVIDIA and AMD GPUs with a 16-entry cache.

- Repository: [Arseny Kapoulkine/meshoptimizer](https://github.com/zeux/meshoptimizer)
- Tags: internals
- Published: 2026-07-11

---

**The vertex cache optimizer in meshoptimizer is tuned for GPUs with cache profiles similar to consumer-grade NVIDIA and AMD graphics processors, assuming a maximum cache size of 16 entries.**

The `zeux/meshoptimizer` repository provides high-performance mesh optimization algorithms for real-time rendering and game engine pipelines. Its vertex cache optimization specifically targets the cache behaviors found in modern desktop GPUs to minimize the Average Cache Miss Ratio (ACMR) on the most common consumer hardware.

## Target GPU Architecture Profile

The optimizer is designed for GPUs exhibiting **FIFO (first-in-first-out)** vertex cache behavior characteristic of mainstream graphics hardware.

### NVIDIA and AMD Consumer GPUs

According to the source code comment in [`src/vcacheoptimizer.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/vcacheoptimizer.cpp) (lines 22-25):

> "Tuned to minimize the ACMR of a GPU that has a cache profile similar to NVidia and AMD"

This explicit tuning targets the post-transform vertex cache implementations found in **NVIDIA GeForce** and **AMD Radeon** series GPUs. The algorithm assumes cache hit patterns and replacement policies that match these architectures, which represent the dominant hardware in PC gaming and professional visualization markets.

### Cache Size Specifications

The implementation supports a **maximum cache size of 16 entries**, which corresponds to the typical vertex cache depth in modern consumer-grade GPUs. The optimizer does not target specialized architectures with substantially different cache sizes—such as legacy hardware or mobile GPUs with tile-based architectures—because these exhibit divergent memory access patterns that would require different optimization heuristics.

## Implementation Details

In [`src/vcacheoptimizer.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/vcacheoptimizer.cpp), the `meshopt_optimizeVertexCache` function implements a scoring algorithm optimized for the NVIDIA/AMD cache profile. The function automatically applies hardware-specific heuristics without requiring manual platform detection.

For scenarios requiring explicit cache size configuration while maintaining the same NVIDIA/AMD-tuned optimization strategy, the library exposes `meshopt_optimizeVertexCacheFifo`.

## Code Examples

### Standard Vertex Cache Optimization

The following example demonstrates the default optimization path tuned for NVIDIA/AMD GPUs:

```cpp
#include "meshoptimizer.h"
#include <vector>

int main()
{
    // Example index buffer (triangle list)
    std::vector<unsigned int> indices = {
        0, 1, 2,
        2, 3, 0,
        0, 2, 4
    };

    // Allocate destination buffer matching source size
    std::vector<unsigned int> optimized(indices.size());

    // Run the optimizer – it uses the default score table tuned for NVIDIA/AMD GPUs
    meshopt_optimizeVertexCache(
        optimized.data(),
        indices.data(),
        indices.size(),
        /*vertex_count=*/5
    );

    // optimized now contains index order with reduced ACMR
    return 0;
}

```

### Custom Cache Size Configuration

To specify a non-standard cache size while maintaining the NVIDIA/AMD optimization profile:

```cpp
meshopt_optimizeVertexCacheFifo(
    optimized.data(),
    indices.data(),
    indices.size(),
    /*vertex_count=*/5,
    /*cache_size=*/32);   // any size >= 3; the algorithm still assumes an NVIDIA/AMD-like profile

```

## Key Source Files

The vertex cache optimization implementation resides in specific files within the repository:

- **[`src/vcacheoptimizer.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/vcacheoptimizer.cpp)**: Contains the optimizer implementation and the explicit comment regarding NVIDIA/AMD cache profile targeting
- **[`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h)**: Public header declaring `meshopt_optimizeVertexCache`, `meshopt_optimizeVertexCacheStrip`, and `meshopt_optimizeVertexCacheFifo`
- **[`tools/vcachetuner.cpp`](https://github.com/zeux/meshoptimizer/blob/main/tools/vcachetuner.cpp)**: Benchmarking utility for evaluating cache optimization across different cache sizes

## Summary

- **Target hardware**: NVIDIA and AMD consumer GPUs with FIFO vertex cache behavior
- **Cache assumptions**: Up to 16 entries maximum, matching typical GeForce and Radeon implementations
- **Primary source**: [`src/vcacheoptimizer.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/vcacheoptimizer.cpp) implements the NVIDIA/AMD-tuned scoring algorithm
- **API entry points**: `meshopt_optimizeVertexCache` for standard use, `meshopt_optimizeVertexCacheFifo` for custom cache sizes
- **Limitations**: Does not optimize for mobile tile-based architectures or legacy GPUs with divergent cache behaviors

## Frequently Asked Questions

### Does meshoptimizer support Intel GPUs?

The optimizer functions on Intel GPUs because modern Intel graphics implement similar FIFO vertex cache behaviors. However, the algorithm remains specifically tuned for NVIDIA/AMD cache hit patterns, which may yield suboptimal ACMR results on Intel hardware compared to the targeted architectures.

### Can I use a custom cache size for specialized hardware?

Yes, call `meshopt_optimizeVertexCacheFifo` with your specific `cache_size` parameter. While this allows experimentation with different cache depths, the underlying algorithm retains its NVIDIA/AMD cache behavior assumptions and does not adapt to fundamentally different cache replacement policies found in specialized or mobile architectures.

### What does ACMR mean in vertex cache optimization?

**Average Cache Miss Ratio (ACMR)** quantifies the efficiency of vertex reuse by measuring how frequently the GPU must re-process vertices versus accessing them from the post-transform cache. Lower ACMR values indicate better cache utilization, directly correlating with reduced memory bandwidth and improved rendering performance in geometry-bound scenarios.

### Why doesn't meshoptimizer target mobile GPU architectures?

Mobile GPUs typically employ tile-based deferred rendering (TBDR) architectures with vertex processing units that behave differently from the immediate-mode rendering pipelines of desktop NVIDIA and AMD cards. The cache hit patterns, memory hierarchies, and vertex reuse characteristics in TBDR systems require different optimization strategies than the FIFO cache model implemented in [`src/vcacheoptimizer.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/vcacheoptimizer.cpp).