How to Integrate meshopt_encodeMeshlet/meshopt_decodeMeshlet for Mesh Shader Pipelines

Integrating meshopt_encodeMeshlet and meshopt_decodeMeshlet into a mesh shader pipeline requires building locality-optimized meshlets with meshopt_buildMeshlets, encoding them offline using meshopt_encodeMeshlet, and decoding them at runtime either via meshopt_decodeMeshlet on the CPU or the provided meshletdec.slang shader on the GPU.

The meshoptimizer library provides a frameless codec specifically designed for modern mesh shader pipelines in DirectX 12 and Vulkan. When you integrate meshopt_encodeMeshlet and meshopt_decodeMeshlet, you create a compact binary representation of meshlets that achieves 5-7 bits per triangle for triangle data and 9-12 bits per triangle including vertex references, significantly reducing GPU memory bandwidth during rendering.

Understanding the Meshlet Codec Architecture

The codec operates on a frameless binary format defined in src/meshoptimizer.h (lines 720-728) and implemented in src/meshletcodec.cpp. Unlike traditional compression formats, the encoded stream contains no version header, ensuring forward compatibility across all library versions.

The architecture requires three distinct data buffers:

  • Meshlet descriptors (meshopt_Meshlet): Store vertex_offset, vertex_count, triangle_offset, and triangle_count
  • Meshlet vertices: Array of global vertex indices referenced by the meshlet
  • Meshlet triangles: Micro-index buffer containing 3 bytes per triangle that index into the local meshlet vertices

Step 1: Building Locality-Optimized Meshlets

Before encoding, generate meshlets using the API declared in src/meshoptimizer.h and implemented in src/meshletutils.cpp. The meshlets must be locality-optimized using meshopt_optimizeMeshletLevel to maximize compression efficiency.

#include "meshoptimizer.h"

// Configure meshlet limits (must match shader constraints)
const size_t max_vertices_per_meshlet = 128;
const size_t max_triangles_per_meshlet = 128;

std::vector<meshopt_Meshlet> meshlets;
std::vector<unsigned int> meshlet_vertices(indices.size());
std::vector<unsigned char> meshlet_triangles(indices.size());

size_t meshlet_count = meshopt_buildMeshlets(
    meshlets.data(),
    meshlet_vertices.data(),
    meshlet_triangles.data(),
    indices.data(),
    indices.size(),
    positions, vertex_count,
    sizeof(vec3),            // Position stride
    max_vertices_per_meshlet,
    max_triangles_per_meshlet,
    0.0f);                   // cone_weight (0 disables cone culling)

// Optimize for locality before encoding
for (size_t i = 0; i < meshlet_count; i++) {
    meshopt_optimizeMeshletLevel(
        &meshlet_vertices[meshlets[i].vertex_offset],
        &meshlet_triangles[meshlets[i].triangle_offset],
        meshlets[i].triangle_count,
        meshlets[i].vertex_count);
}

Step 2: Encoding Meshlets with meshopt_encodeMeshlet

Allocate worst-case buffer sizes using meshopt_encodeMeshletBound, then encode each meshlet individually. You must store the vertex_count, triangle_count, and returned encoded_size alongside the compressed payload.

// Allocate worst-case buffer
std::vector<unsigned char> encode_buffer(meshopt_encodeMeshletBound(max_vertices_per_meshlet, max_triangles_per_meshlet));

// Storage for your asset format
struct MeshletHeader {
    uint8_t vertexCount;
    uint8_t triangleCount;
    uint16_t encodedSize;
};
std::vector<MeshletHeader> headers;
std::vector<unsigned char> payload_stream;

for (const meshopt_Meshlet& m : meshlets) {
    size_t encoded_size = meshopt_encodeMeshlet(
        encode_buffer.data(),
        encode_buffer.size(),
        &meshlet_vertices[m.vertex_offset],
        m.vertex_count,
        &meshlet_triangles[m.triangle_offset],
        m.triangle_count);
    
    // Store header and payload
    headers.push_back({
        static_cast<uint8_t>(m.vertex_count),
        static_cast<uint8_t>(m.triangle_count),
        static_cast<uint16_t>(encoded_size)
    });
    
    payload_stream.insert(payload_stream.end(), 
                         encode_buffer.begin(), 
                         encode_buffer.begin() + encoded_size);
}

Step 3: Decoding Meshlets for GPU Rendering

You have two decoding strategies depending on your pipeline architecture: CPU-side preprocessing or GPU-side on-demand decompression.

CPU Decoding with meshopt_decodeMeshlet

Use this approach when decompressing at load time before uploading to GPU buffers. The decoder automatically deduces element sizes from pointer types: 2 bytes for 16-bit vertex indices and 3 bytes for packed triangles.

// Prepare output buffers
std::vector<uint16_t> decoded_vertices(meshlet.vertex_count);
std::vector<uint8_t> decoded_triangles(meshlet.triangle_count * 3);

int result = meshopt_decodeMeshlet(
    decoded_vertices.data(), meshlet.vertex_count,
    decoded_triangles.data(), meshlet.triangle_count,
    encoded_payload.data(), encoded_size);

assert(result == 0); // 0 indicates success

GPU Decoding Using meshletdec.slang

For direct GPU-side decompression, reference the demo/meshletdec.slang shader provided in the repository. This SLANG/HLSL implementation decodes the same frameless format with a 32-bit output layout, matching the raw decoder layout used by the C API.

// Simplified HLSL adaptation from meshletdec.slang
struct MeshletHeader {
    uint vertexCount;
    uint triangleCount;
    uint payloadOffset;
};

StructuredBuffer<MeshletHeader> gHeaders : register(t0);
ByteAddressBuffer gPayload : register(t1);

[numthreads(1,1,1)]
void DecodeMeshlet(uint meshletId : SV_GroupID)
{
    MeshletHeader hdr = gHeaders[meshletId];
    
    // Calculate payload address based on accumulated offsets
    uint payloadAddr = hdr.payloadOffset;
    
    // Decode loop follows bit-packing layout from meshletcodec.cpp
    // See meshletdec.slang for exact bit manipulation operations
}

Pipeline Integration Checklist

Complete your integration by following this asset flow:

  1. Preprocessing (Offline): Build meshlets, optimize locality, encode with meshopt_encodeMeshlet, and serialize headers plus payloads to custom container format
  2. Upload: Create GPU buffers containing:
    • Array of MeshletHeader structures (or equivalent)
    • Contiguous block of encoded meshlet payloads
  3. Runtime: Bind buffers as shader resources. In the mesh shader:
    • Read header to determine counts
    • Call decode routine (inline the meshletdec.slang logic or use lookup tables)
    • Emit primitives using decoded micro-indices

Reference implementation tools/codecbench.cpp demonstrates the complete build-encode-decode cycle for testing your integration.

Summary

  • Build meshlets using meshopt_buildMeshlets followed by meshopt_optimizeMeshletLevel to ensure compressibility
  • Encode offline with meshopt_encodeMeshletBound allocating buffers and meshopt_encodeMeshlet generating frameless compressed payloads
  • Store metadata alongside encoded bytes: vertex_count, triangle_count, and encoded_size
  • Decode on CPU using meshopt_decodeMeshlet for automatic pointer-type deduction, or on GPU using the reference meshletdec.slang implementation
  • Achieve compression rates of 5-7 bits per triangle for geometry and 9-12 bits per triangle for full meshlet representation

Frequently Asked Questions

What compression ratio does meshopt_encodeMeshlet achieve?

The codec achieves 5-7 bits per triangle for triangle data alone, and 9-12 bits per triangle when including vertex references. This assumes meshlets were optimized with meshopt_optimizeMeshletLevel before encoding, as locality optimization is required for maximum compression efficiency.

Can I decode meshlets directly in a mesh shader?

Yes. The demo/meshletdec.slang file provides a reference GPU decoder implementation compatible with HLSL and SLANG. The algorithm uses the same bit-packing layout as the CPU-side meshopt_decodeMeshlet function implemented in src/meshletcodec.cpp, ensuring bit-exact results between CPU and GPU decoding paths.

Why must meshlets be optimized before calling meshopt_encodeMeshlet?

The encoder relies on locality of reference to achieve its compression ratios. The meshopt_optimizeMeshletLevel function reorders vertices within a meshlet to maximize index locality, which the delta-encoding scheme in meshopt_encodeMeshlet exploits. Unoptimized meshlets will compress significantly worse, potentially approaching uncompressed sizes.

Is the meshopt_encodeMeshlet format stable across library versions?

Yes. The format is frameless, meaning it contains no version headers or flags. As documented in the repository README and implemented in src/meshletcodec.cpp, the encoded stream format is guaranteed compatible across all library versions, making it safe for long-term asset storage without format versioning concerns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →