How to Decode Meshlet Data on the GPU Using Write-Combined Memory

Use meshopt_decodeMeshletRaw to stream decoded vertices and triangles directly into write-combined mapped GPU buffers, eliminating CPU-side copies and maximizing throughput.

The meshoptimizer library provides optimized routines for encoding mesh geometry into compact meshlets. When you need to decode meshlet data on the GPU using write-combined memory, the library exposes a low-level API that writes directly into mapped GPU memory regions without intermediate staging buffers.

The Two Decoding Paths

According to the zeux/meshoptimizer source code, the library offers two distinct entry points for meshlet decoding:

meshopt_decodeMeshlet – A high-level convenience function that internally allocates temporary buffers and copies data before producing the final result. This approach incurs an extra memory copy, making it unsuitable for direct GPU streaming.

meshopt_decodeMeshletRaw – A low-level decoder that expects pre-allocated, aligned destination buffers. It writes vertices and triangle indices directly into the supplied memory addresses, making it the only variant compatible with write-combined GPU mappings.

Understanding Write-Combined Memory

Write-combined memory is a GPU buffer mapping that disables CPU caching for writes. This allows the CPU to stream large datasets without triggering cache-coherency stalls or polluting the CPU cache hierarchy.

When destination buffers are mapped as write-combined, meshopt_decodeMeshletRaw can stream decoded mesh data straight into the mapped region. This zero-copy path eliminates the need for CPU-side staging buffers and reduces memory bandwidth pressure.

Implementation Steps for GPU Decoding

To decode meshlet data directly into GPU memory, follow this sequence:

  1. Allocate GPU buffers for vertices and triangles with sufficient size and 16-byte alignment for SIMD-friendly operations.
  2. Map the buffers with write-combined flags (e.g., GL_MAP_WRITE_BIT | GL_MAP_UNSYNCHRONIZED_BIT in OpenGL or D3D11_MAP_WRITE_DISCARD in DirectX).
  3. Call meshopt_decodeMeshletRaw passing the raw pointers obtained from the mapping. The function writes directly into the WC-mapped region.
  4. Unmap the buffers once decoding completes; the GPU can consume the data immediately without additional copying.

Code Examples

OpenGL with Write-Combined Mapping

// Create GPU buffers
GLuint vbo, ibo;
glGenBuffers(1, &vbo);
glGenBuffers(1, &ibo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, vertexBufferSize, nullptr, GL_DYNAMIC_DRAW);
glBindBuffer(GL_ELEMENT_ARRAY_BUFFER, ibo);
glBufferData(GL_ELEMENT_ARRAY_BUFFER, indexBufferSize, nullptr, GL_DYNAMIC_DRAW);

// Map with write-combined hints
void* vtxPtr = glMapBufferRange(GL_ARRAY_BUFFER, 0, vertexBufferSize,
    GL_MAP_WRITE_BIT | GL_MAP_UNSYNCHRONIZED_BIT);
void* idxPtr = glMapBufferRange(GL_ELEMENT_ARRAY_BUFFER, 0, indexBufferSize,
    GL_MAP_WRITE_BIT | GL_MAP_UNSYNCHRONIZED_BIT);

// Decode directly into mapped GPU memory
int result = meshopt_decodeMeshletRaw(
    static_cast<unsigned int*>(vtxPtr),  // destination vertices
    vertexCount,
    static_cast<unsigned int*>(idxPtr),  // destination triangles
    triangleCount,
    encodedData,                          // packed meshlet source
    encodedSize);

// Unmap for GPU consumption
glUnmapBuffer(GL_ARRAY_BUFFER);
glUnmapBuffer(GL_ELEMENT_ARRAY_BUFFER);

DirectX 11 with Discard/No-Overwrite Flags

// Create dynamic buffers
ID3D11Buffer* vertexBuf = nullptr;
ID3D11Buffer* indexBuf = nullptr;
D3D11_BUFFER_DESC desc = {};
desc.ByteWidth = vertexBufferSize;
desc.Usage = D3D11_USAGE_DYNAMIC;
desc.BindFlags = D3D11_BIND_VERTEX_BUFFER;
desc.CPUAccessFlags = D3D11_CPU_ACCESS_WRITE;
device->CreateBuffer(&desc, nullptr, &vertexBuf);

desc.ByteWidth = indexBufferSize;
desc.BindFlags = D3D11_BIND_INDEX_BUFFER;
device->CreateBuffer(&desc, nullptr, &indexBuf);

// Map for write-combined access
D3D11_MAPPED_SUBRESOURCE mappedVtx;
context->Map(vertexBuf, 0, D3D11_MAP_WRITE_DISCARD, 0, &mappedVtx);
D3D11_MAPPED_SUBRESOURCE mappedIdx;
context->Map(indexBuf, 0, D3D11_MAP_WRITE_DISCARD, 0, &mappedIdx);

// Stream decode into GPU memory
int rc = meshopt_decodeMeshletRaw(
    static_cast<unsigned int*>(mappedVtx.pData), vertexCount,
    static_cast<unsigned int*>(mappedIdx.pData), triangleCount,
    encodedData, encodedSize);

// Release mapping
context->Unmap(vertexBuf, 0);
context->Unmap(indexBuf, 0);

CPU Reference Implementation

// Allocate aligned host memory for testing
unsigned int* vtx = (unsigned int*)_aligned_malloc(vertexCount * sizeof(Vertex), 16);
unsigned int* idx = (unsigned int*)_aligned_malloc(triangleCount * sizeof(uint32_t), 16);

// Decode without GPU involvement
int rc = meshopt_decodeMeshletRaw(vtx, vertexCount, idx, triangleCount,
                                  encodedData, encodedSize);

Technical Requirements and Source Locations

The decoder implementation imposes specific alignment and format constraints:

  • Destination alignment: Buffers must be aligned to 16 bytes for optimal SIMD performance in the decode loops.
  • Vertex stride: Must be a multiple of 4 bytes as expected by the decoder.
  • Index format: Triangle indices are stored as bytes (1 byte) or half-words (2 bytes) depending on the meshlet configuration.

The function signature is declared in src/meshoptimizer.h at line 349, while the actual SIMD-optimized decoding logic resides in src/meshletcodec.cpp at line 1015. A practical integration example appears in demo/main.cpp at line 871, demonstrating how the library is invoked after loading a meshlet-encoded file.

Summary

  • Use meshopt_decodeMeshletRaw instead of the high-level meshopt_decodeMeshlet when targeting GPU memory.
  • Map buffers as write-combined to avoid cache coherency penalties and enable zero-copy streaming.
  • Maintain 16-byte alignment on destination buffers to ensure SIMD optimizations in the decoder.
  • Unmap immediately after decoding to make data available for GPU consumption.

Frequently Asked Questions

What is the difference between meshopt_decodeMeshlet and meshopt_decodeMeshletRaw?

The high-level meshopt_decodeMeshlet function internally copies data to temporary buffers before writing results, which prevents direct use with GPU-mapped memory. The meshopt_decodeMeshletRaw variant writes directly into caller-provided buffers, making it suitable for write-combined mappings and zero-copy GPU workflows.

Why does the destination buffer need 16-byte alignment?

The decoder uses SIMD instructions (SSE/AVX) to unpack and write vertex and index data. Aligned loads and stores in the implementation found in src/meshletcodec.cpp require 16-byte alignment for optimal performance and to avoid segmentation faults on certain architectures.

Can I use write-combined memory with the high-level decoder?

No. The high-level meshopt_decodeMeshlet performs internal memory allocations and copies that assume cacheable CPU memory. Only meshopt_decodeMeshletRaw supports writing directly into pre-mapped GPU buffers without intermediate staging.

What GPU APIs support write-combined memory mappings?

OpenGL supports write-combined hints via glMapBufferRange with GL_MAP_UNSYNCHRONIZED_BIT. DirectX 11 provides D3D11_MAP_WRITE_DISCARD and D3D11_MAP_WRITE_NO_OVERWRITE. Vulkan and DirectX 12 offer similar memory host-visible flags that achieve the same uncached write semantics required for efficient streaming.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →