# How `meshopt_optimizeVertexFetch` Improves GPU Memory Bandwidth Through Vertex Fetch Optimization

> Optimize GPU memory bandwidth with meshopt_optimizeVertexFetch. Discover how reordering vertex data eliminates over-fetch and maximizes cache line utilization for faster rendering.

- Repository: [Arseny Kapoulkine/meshoptimizer](https://github.com/zeux/meshoptimizer)
- Tags: deep-dive
- Published: 2026-07-12

---

**Vertex fetch optimization reduces GPU memory bandwidth by reordering vertex data to match the index buffer access pattern, enabling sequential reads that maximize cache line utilization and eliminate over-fetch.**

The `meshopt_optimizeVertexFetch` function in the [meshoptimizer](https://github.com/zeux/meshoptimizer) library restructures vertex buffers to align with GPU memory subsystem requirements. When rendering, vertex shaders read attributes through a memory controller optimized for sequential access. By reordering vertices in the exact order they are referenced by the optimized index buffer, this function transforms random memory access into cache-friendly sequential streams, directly reducing memory transactions and cache misses.

## The GPU Memory Bandwidth Bottleneck

Modern GPUs fetch vertex attributes via a memory subsystem that operates most efficiently with sequential, cache-aligned access patterns. When vertex data is scattered randomly in memory, each vertex shader invocation may trigger a separate cache line fetch, loading 64 or 128 bytes to access only 32–64 bytes of actual vertex data. This **over-fetch** wastes bandwidth and pollutes the cache with bytes that will never be consumed by the shader, especially problematic for meshes with multiple attributes (normals, tangents, UVs, and colors).

## How Vertex Fetch Optimization Works

The `meshopt_optimizeVertexFetch` algorithm, implemented in [`src/vfetchoptimizer.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/vfetchoptimizer.cpp) (lines 29-70), constructs a **vertex-fetch remap table** by scanning the index buffer and recording the first encounter of each vertex index. It then reorders the vertex buffer so vertices appear in the same sequence the GPU will access them during rendering, updating the indices in-place to point to the new locations.

This reordering produces a sequentially accessed vertex buffer that allows the memory controller to:
- **Fetch entire cache lines** containing multiple required vertices in a single transaction rather than individual fetches for scattered vertices
- **Avoid over-fetch** by ensuring nearly all loaded bytes are consumed by the vertex shader, reducing wasted bandwidth
- **Increase cache reuse** by placing vertices from spatially adjacent triangles in contiguous memory locations, allowing a single cache line to serve multiple vertex shader invocations before eviction

## Measuring Over-Fetch Reduction

The `meshopt_analyzeVertexFetch` function defined in [`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h) (lines 51-56) provides quantitative metrics for optimization quality. The `overfetch` statistic specifically measures wasted bandwidth—lower values indicate that the GPU reads fewer unused bytes per vertex, confirming that the vertex fetch optimization is effectively reducing memory bandwidth requirements. According to the meshoptimizer source code, this metric helps verify that vertex data is now accessed in the same order it is stored.

## Pipeline Integration

Vertex fetch optimization achieves maximum benefit when performed **after** vertex-cache optimization (`meshopt_optimizeVertexCache`) and overdraw optimization (`meshopt_optimizeOverdraw`). Because the final triangle order determines the optimal vertex sequence, reordering triangles for temporal locality must precede vertex buffer reordering. The function also performs vertex deduplication, returning a final vertex count (via the `unique` return value) that may be lower than the input when unused vertices exist in the buffer.

## Implementation Example

The following complete workflow demonstrates vertex fetch optimization applied after cache and overdraw optimization, including compact vertex buffer generation:

```cpp
// 1. Generate a compact index buffer (optional but recommended)
size_t indexCount = faces * 3;
std::vector<unsigned int> remap(indexCount);
size_t vertexCount = meshopt_generateVertexRemap(
    remap.data(), nullptr, indexCount,
    unindexedVertices.data(), indexCount, sizeof(Vertex));

// 2. Apply the remap to obtain a compact vertex buffer
std::vector<Vertex> vertices(vertexCount);
meshopt_remapVertexBuffer(vertices.data(), unindexedVertices.data(),
    indexCount, sizeof(Vertex), remap.data());

// 3. Reorder the indices and vertex buffer for optimal fetch
std::vector<unsigned int> indices(indexCount);
meshopt_optimizeVertexCache(indices.data(), indices.data(),
    indexCount, vertexCount);               // vertex-cache step
meshopt_optimizeOverdraw(indices.data(), indices.data(),
    indexCount, &vertices[0].x, vertexCount, sizeof(Vertex), 1.05f); // optional overdraw step

size_t unique = meshopt_optimizeVertexFetch(
    vertices.data(),               // destination (can be the same buffer)
    indices.data(),                // will be updated in-place
    indexCount,
    vertices.data(),               // original vertex buffer (same as destination)
    vertexCount,
    sizeof(Vertex));               // stride of a single vertex

printf("Removed %zu unused vertices, final vertex count = %zu\n",
       vertexCount - unique, unique);

```

For multiple vertex streams, use `meshopt_optimizeVertexFetchRemap` to obtain a remap table without modifying the buffers directly, then apply `meshopt_remapVertexBuffer` to each stream individually.

## Summary

- Vertex fetch optimization reduces GPU memory bandwidth by ensuring vertex data is accessed sequentially rather than randomly
- The algorithm in [`src/vfetchoptimizer.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/vfetchoptimizer.cpp) builds a remap table based on first-occurrence during index buffer scanning (lines 29-70)
- Reordering eliminates over-fetch and maximizes cache line utilization, as measured by `meshopt_analyzeVertexFetch` and its `overfetch` metric
- Always execute after vertex-cache and overdraw optimizations to ensure the vertex order matches the final triangle order
- The function supports in-place operation (source and destination may be the same buffer) and removes unused vertices, returning the final unique vertex count

## Frequently Asked Questions

### What is vertex fetch optimization?

Vertex fetch optimization is the process of reordering vertex buffer data to match the sequence in which the GPU requests vertices during rendering. By aligning memory layout with access patterns, it minimizes cache misses and reduces memory bandwidth consumption, particularly for attribute-heavy meshes.

### How does `meshopt_optimizeVertexFetch` differ from vertex cache optimization?

Vertex cache optimization (`meshopt_optimizeVertexCache`) reorders the index buffer to maximize post-transform cache hits, while vertex fetch optimization reorders the vertex buffer itself to improve pre-transform memory access patterns. Fetch optimization should always run after cache optimization because the final triangle order determines the optimal vertex sequence.

### Can I use vertex fetch optimization with multiple vertex streams?

Yes. While the standard function works on interleaved vertex data via the stride parameter, you can use `meshopt_optimizeVertexFetchRemap` to generate a remap table, then apply `meshopt_remapVertexBuffer` to each separate attribute stream (positions, normals, UVs) individually. This ensures all streams remain synchronized while achieving optimal fetch patterns.

### How do I verify that vertex fetch optimization improved my mesh?

Call `meshopt_analyzeVertexFetch` before and after optimization as defined in [`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h). Compare the `overfetch` metric—a significantly lower value indicates reduced wasted bandwidth and confirms that vertices are now stored in a more cache-friendly order that matches the index buffer traversal pattern.