How `meshopt_optimizeVertexFetch` Improves GPU Memory Bandwidth Through Vertex Fetch Optimization
Vertex fetch optimization reduces GPU memory bandwidth by reordering vertex data to match the index buffer access pattern, enabling sequential reads that maximize cache line utilization and eliminate over-fetch.
The meshopt_optimizeVertexFetch function in the meshoptimizer library restructures vertex buffers to align with GPU memory subsystem requirements. When rendering, vertex shaders read attributes through a memory controller optimized for sequential access. By reordering vertices in the exact order they are referenced by the optimized index buffer, this function transforms random memory access into cache-friendly sequential streams, directly reducing memory transactions and cache misses.
The GPU Memory Bandwidth Bottleneck
Modern GPUs fetch vertex attributes via a memory subsystem that operates most efficiently with sequential, cache-aligned access patterns. When vertex data is scattered randomly in memory, each vertex shader invocation may trigger a separate cache line fetch, loading 64 or 128 bytes to access only 32–64 bytes of actual vertex data. This over-fetch wastes bandwidth and pollutes the cache with bytes that will never be consumed by the shader, especially problematic for meshes with multiple attributes (normals, tangents, UVs, and colors).
How Vertex Fetch Optimization Works
The meshopt_optimizeVertexFetch algorithm, implemented in src/vfetchoptimizer.cpp (lines 29-70), constructs a vertex-fetch remap table by scanning the index buffer and recording the first encounter of each vertex index. It then reorders the vertex buffer so vertices appear in the same sequence the GPU will access them during rendering, updating the indices in-place to point to the new locations.
This reordering produces a sequentially accessed vertex buffer that allows the memory controller to:
- Fetch entire cache lines containing multiple required vertices in a single transaction rather than individual fetches for scattered vertices
- Avoid over-fetch by ensuring nearly all loaded bytes are consumed by the vertex shader, reducing wasted bandwidth
- Increase cache reuse by placing vertices from spatially adjacent triangles in contiguous memory locations, allowing a single cache line to serve multiple vertex shader invocations before eviction
Measuring Over-Fetch Reduction
The meshopt_analyzeVertexFetch function defined in src/meshoptimizer.h (lines 51-56) provides quantitative metrics for optimization quality. The overfetch statistic specifically measures wasted bandwidth—lower values indicate that the GPU reads fewer unused bytes per vertex, confirming that the vertex fetch optimization is effectively reducing memory bandwidth requirements. According to the meshoptimizer source code, this metric helps verify that vertex data is now accessed in the same order it is stored.
Pipeline Integration
Vertex fetch optimization achieves maximum benefit when performed after vertex-cache optimization (meshopt_optimizeVertexCache) and overdraw optimization (meshopt_optimizeOverdraw). Because the final triangle order determines the optimal vertex sequence, reordering triangles for temporal locality must precede vertex buffer reordering. The function also performs vertex deduplication, returning a final vertex count (via the unique return value) that may be lower than the input when unused vertices exist in the buffer.
Implementation Example
The following complete workflow demonstrates vertex fetch optimization applied after cache and overdraw optimization, including compact vertex buffer generation:
// 1. Generate a compact index buffer (optional but recommended)
size_t indexCount = faces * 3;
std::vector<unsigned int> remap(indexCount);
size_t vertexCount = meshopt_generateVertexRemap(
remap.data(), nullptr, indexCount,
unindexedVertices.data(), indexCount, sizeof(Vertex));
// 2. Apply the remap to obtain a compact vertex buffer
std::vector<Vertex> vertices(vertexCount);
meshopt_remapVertexBuffer(vertices.data(), unindexedVertices.data(),
indexCount, sizeof(Vertex), remap.data());
// 3. Reorder the indices and vertex buffer for optimal fetch
std::vector<unsigned int> indices(indexCount);
meshopt_optimizeVertexCache(indices.data(), indices.data(),
indexCount, vertexCount); // vertex-cache step
meshopt_optimizeOverdraw(indices.data(), indices.data(),
indexCount, &vertices[0].x, vertexCount, sizeof(Vertex), 1.05f); // optional overdraw step
size_t unique = meshopt_optimizeVertexFetch(
vertices.data(), // destination (can be the same buffer)
indices.data(), // will be updated in-place
indexCount,
vertices.data(), // original vertex buffer (same as destination)
vertexCount,
sizeof(Vertex)); // stride of a single vertex
printf("Removed %zu unused vertices, final vertex count = %zu\n",
vertexCount - unique, unique);
For multiple vertex streams, use meshopt_optimizeVertexFetchRemap to obtain a remap table without modifying the buffers directly, then apply meshopt_remapVertexBuffer to each stream individually.
Summary
- Vertex fetch optimization reduces GPU memory bandwidth by ensuring vertex data is accessed sequentially rather than randomly
- The algorithm in
src/vfetchoptimizer.cppbuilds a remap table based on first-occurrence during index buffer scanning (lines 29-70) - Reordering eliminates over-fetch and maximizes cache line utilization, as measured by
meshopt_analyzeVertexFetchand itsoverfetchmetric - Always execute after vertex-cache and overdraw optimizations to ensure the vertex order matches the final triangle order
- The function supports in-place operation (source and destination may be the same buffer) and removes unused vertices, returning the final unique vertex count
Frequently Asked Questions
What is vertex fetch optimization?
Vertex fetch optimization is the process of reordering vertex buffer data to match the sequence in which the GPU requests vertices during rendering. By aligning memory layout with access patterns, it minimizes cache misses and reduces memory bandwidth consumption, particularly for attribute-heavy meshes.
How does meshopt_optimizeVertexFetch differ from vertex cache optimization?
Vertex cache optimization (meshopt_optimizeVertexCache) reorders the index buffer to maximize post-transform cache hits, while vertex fetch optimization reorders the vertex buffer itself to improve pre-transform memory access patterns. Fetch optimization should always run after cache optimization because the final triangle order determines the optimal vertex sequence.
Can I use vertex fetch optimization with multiple vertex streams?
Yes. While the standard function works on interleaved vertex data via the stride parameter, you can use meshopt_optimizeVertexFetchRemap to generate a remap table, then apply meshopt_remapVertexBuffer to each separate attribute stream (positions, normals, UVs) individually. This ensures all streams remain synchronized while achieving optimal fetch patterns.
How do I verify that vertex fetch optimization improved my mesh?
Call meshopt_analyzeVertexFetch before and after optimization as defined in src/meshoptimizer.h. Compare the overfetch metric—a significantly lower value indicates reduced wasted bandwidth and confirms that vertices are now stored in a more cache-friendly order that matches the index buffer traversal pattern.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →