# How to Decode Meshlet Data on the GPU Using Write-Combined Memory

> Decode meshlet data on GPU with write-combined memory for maximum throughput. Stream decoded vertices and triangles directly into GPU buffers, eliminating CPU copies.

- Repository: [Arseny Kapoulkine/meshoptimizer](https://github.com/zeux/meshoptimizer)
- Tags: how-to-guide
- Published: 2026-07-12

---

**Use `meshopt_decodeMeshletRaw` to stream decoded vertices and triangles directly into write-combined mapped GPU buffers, eliminating CPU-side copies and maximizing throughput.**

The **meshoptimizer** library provides optimized routines for encoding mesh geometry into compact meshlets. When you need to decode meshlet data on the GPU using write-combined memory, the library exposes a low-level API that writes directly into mapped GPU memory regions without intermediate staging buffers.

## The Two Decoding Paths

According to the zeux/meshoptimizer source code, the library offers two distinct entry points for meshlet decoding:

**`meshopt_decodeMeshlet`** – A high-level convenience function that internally allocates temporary buffers and copies data before producing the final result. This approach incurs an extra memory copy, making it unsuitable for direct GPU streaming.

**`meshopt_decodeMeshletRaw`** – A low-level decoder that expects pre-allocated, aligned destination buffers. It writes vertices and triangle indices directly into the supplied memory addresses, making it the only variant compatible with write-combined GPU mappings.

## Understanding Write-Combined Memory

Write-combined memory is a GPU buffer mapping that disables CPU caching for writes. This allows the CPU to stream large datasets without triggering cache-coherency stalls or polluting the CPU cache hierarchy.

When destination buffers are mapped as write-combined, `meshopt_decodeMeshletRaw` can stream decoded mesh data straight into the mapped region. This zero-copy path eliminates the need for CPU-side staging buffers and reduces memory bandwidth pressure.

## Implementation Steps for GPU Decoding

To decode meshlet data directly into GPU memory, follow this sequence:

1. **Allocate GPU buffers** for vertices and triangles with sufficient size and **16-byte alignment** for SIMD-friendly operations.
2. **Map the buffers** with write-combined flags (e.g., `GL_MAP_WRITE_BIT | GL_MAP_UNSYNCHRONIZED_BIT` in OpenGL or `D3D11_MAP_WRITE_DISCARD` in DirectX).
3. **Call `meshopt_decodeMeshletRaw`** passing the raw pointers obtained from the mapping. The function writes directly into the WC-mapped region.
4. **Unmap the buffers** once decoding completes; the GPU can consume the data immediately without additional copying.

## Code Examples

### OpenGL with Write-Combined Mapping

```cpp
// Create GPU buffers
GLuint vbo, ibo;
glGenBuffers(1, &vbo);
glGenBuffers(1, &ibo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, vertexBufferSize, nullptr, GL_DYNAMIC_DRAW);
glBindBuffer(GL_ELEMENT_ARRAY_BUFFER, ibo);
glBufferData(GL_ELEMENT_ARRAY_BUFFER, indexBufferSize, nullptr, GL_DYNAMIC_DRAW);

// Map with write-combined hints
void* vtxPtr = glMapBufferRange(GL_ARRAY_BUFFER, 0, vertexBufferSize,
    GL_MAP_WRITE_BIT | GL_MAP_UNSYNCHRONIZED_BIT);
void* idxPtr = glMapBufferRange(GL_ELEMENT_ARRAY_BUFFER, 0, indexBufferSize,
    GL_MAP_WRITE_BIT | GL_MAP_UNSYNCHRONIZED_BIT);

// Decode directly into mapped GPU memory
int result = meshopt_decodeMeshletRaw(
    static_cast<unsigned int*>(vtxPtr),  // destination vertices
    vertexCount,
    static_cast<unsigned int*>(idxPtr),  // destination triangles
    triangleCount,
    encodedData,                          // packed meshlet source
    encodedSize);

// Unmap for GPU consumption
glUnmapBuffer(GL_ARRAY_BUFFER);
glUnmapBuffer(GL_ELEMENT_ARRAY_BUFFER);

```

### DirectX 11 with Discard/No-Overwrite Flags

```cpp
// Create dynamic buffers
ID3D11Buffer* vertexBuf = nullptr;
ID3D11Buffer* indexBuf = nullptr;
D3D11_BUFFER_DESC desc = {};
desc.ByteWidth = vertexBufferSize;
desc.Usage = D3D11_USAGE_DYNAMIC;
desc.BindFlags = D3D11_BIND_VERTEX_BUFFER;
desc.CPUAccessFlags = D3D11_CPU_ACCESS_WRITE;
device->CreateBuffer(&desc, nullptr, &vertexBuf);

desc.ByteWidth = indexBufferSize;
desc.BindFlags = D3D11_BIND_INDEX_BUFFER;
device->CreateBuffer(&desc, nullptr, &indexBuf);

// Map for write-combined access
D3D11_MAPPED_SUBRESOURCE mappedVtx;
context->Map(vertexBuf, 0, D3D11_MAP_WRITE_DISCARD, 0, &mappedVtx);
D3D11_MAPPED_SUBRESOURCE mappedIdx;
context->Map(indexBuf, 0, D3D11_MAP_WRITE_DISCARD, 0, &mappedIdx);

// Stream decode into GPU memory
int rc = meshopt_decodeMeshletRaw(
    static_cast<unsigned int*>(mappedVtx.pData), vertexCount,
    static_cast<unsigned int*>(mappedIdx.pData), triangleCount,
    encodedData, encodedSize);

// Release mapping
context->Unmap(vertexBuf, 0);
context->Unmap(indexBuf, 0);

```

### CPU Reference Implementation

```cpp
// Allocate aligned host memory for testing
unsigned int* vtx = (unsigned int*)_aligned_malloc(vertexCount * sizeof(Vertex), 16);
unsigned int* idx = (unsigned int*)_aligned_malloc(triangleCount * sizeof(uint32_t), 16);

// Decode without GPU involvement
int rc = meshopt_decodeMeshletRaw(vtx, vertexCount, idx, triangleCount,
                                  encodedData, encodedSize);

```

## Technical Requirements and Source Locations

The decoder implementation imposes specific alignment and format constraints:

- **Destination alignment**: Buffers must be aligned to **16 bytes** for optimal SIMD performance in the decode loops.
- **Vertex stride**: Must be a multiple of 4 bytes as expected by the decoder.
- **Index format**: Triangle indices are stored as bytes (1 byte) or half-words (2 bytes) depending on the meshlet configuration.

The function signature is declared in **[`src/meshoptimizer.h`](https://github.com/zeux/meshoptimizer/blob/main/src/meshoptimizer.h)** at line 349, while the actual SIMD-optimized decoding logic resides in **[`src/meshletcodec.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/meshletcodec.cpp)** at line 1015. A practical integration example appears in **[`demo/main.cpp`](https://github.com/zeux/meshoptimizer/blob/main/demo/main.cpp)** at line 871, demonstrating how the library is invoked after loading a meshlet-encoded file.

## Summary

- **Use `meshopt_decodeMeshletRaw`** instead of the high-level `meshopt_decodeMeshlet` when targeting GPU memory.
- **Map buffers as write-combined** to avoid cache coherency penalties and enable zero-copy streaming.
- **Maintain 16-byte alignment** on destination buffers to ensure SIMD optimizations in the decoder.
- **Unmap immediately** after decoding to make data available for GPU consumption.

## Frequently Asked Questions

### What is the difference between meshopt_decodeMeshlet and meshopt_decodeMeshletRaw?

The high-level `meshopt_decodeMeshlet` function internally copies data to temporary buffers before writing results, which prevents direct use with GPU-mapped memory. The `meshopt_decodeMeshletRaw` variant writes directly into caller-provided buffers, making it suitable for write-combined mappings and zero-copy GPU workflows.

### Why does the destination buffer need 16-byte alignment?

The decoder uses SIMD instructions (SSE/AVX) to unpack and write vertex and index data. Aligned loads and stores in the implementation found in [`src/meshletcodec.cpp`](https://github.com/zeux/meshoptimizer/blob/main/src/meshletcodec.cpp) require 16-byte alignment for optimal performance and to avoid segmentation faults on certain architectures.

### Can I use write-combined memory with the high-level decoder?

No. The high-level `meshopt_decodeMeshlet` performs internal memory allocations and copies that assume cacheable CPU memory. Only `meshopt_decodeMeshletRaw` supports writing directly into pre-mapped GPU buffers without intermediate staging.

### What GPU APIs support write-combined memory mappings?

OpenGL supports write-combined hints via `glMapBufferRange` with `GL_MAP_UNSYNCHRONIZED_BIT`. DirectX 11 provides `D3D11_MAP_WRITE_DISCARD` and `D3D11_MAP_WRITE_NO_OVERWRITE`. Vulkan and DirectX 12 offer similar memory host-visible flags that achieve the same uncached write semantics required for efficient streaming.