Recommended max_vertices and max_triangles for NVIDIA vs AMD Mesh Shaders
For NVIDIA Turing and newer, use 64 max_vertices and 126 max_triangles per meshlet; for AMD RDNA2 and newer, 64/126 is safe, but symmetric limits like 64/64 or 128/128 may yield better hardware utilization.
Mesh shaders process geometry in small clusters called meshlets, and the max_vertices and max_triangles parameters define the capacity of these work groups processed by a single GPU work-group. The meshoptimizer library provides hardware-aware defaults in src/meshoptimizer.h and README.md that balance vertex reuse with GPU occupancy across different vendors.
NVIDIA Mesh Shader Limits
NVIDIA hardware imposes strict driver-level constraints that favor asymmetric limits. According to the meshoptimizer source code and documentation, Turing and newer GPUs perform best when max_vertices ≤ 64 and max_triangles ≤ 126. These values maximize triangle density while respecting the shader interface limits.
The 64/126 configuration is the primary recommendation in the library's documentation because it delivers the optimal balance between vertex reuse and parallel occupancy on NVIDIA architectures. Exceeding these limits may cause pipeline creation failures because the hardware output registers cannot accommodate larger primitive counts.
AMD Mesh Shader Limits
AMD GPUs exhibit different performance characteristics depending on the generation. The meshoptimizer documentation notes that older AMD hardware (pre-RDNA2) works best with symmetric limits where max_vertices = max_triangles, such as 64/64 or 128/128.
On RDNA2 and newer devices, the NVIDIA-compatible limits of 64/126 are also fully supported and safe. However, many AMD-focused implementations still recommend matched pairs to avoid under-utilizing the hardware's parallel execution units. When targeting AMD-only pipelines, experiment with symmetric configurations like 64/64 to determine if fill-rates improve for your specific mesh topology.
Implementing Meshlet Generation
The meshopt_buildMeshlets function declared in src/meshoptimizer.h consumes these limits as runtime parameters. First, calculate an upper bound using meshopt_buildMeshletsBound implemented in src/meshletutils.cpp, then allocate storage and build the meshlet data:
// NVIDIA-safe limits (also compatible with AMD RDNA2+)
const size_t max_vertices = 64;
const size_t max_triangles = 126;
float cone_weight = 0.0f; // 0 disables cone culling
// Calculate required meshlet capacity
size_t max_meshlets = meshopt_buildMeshletsBound(
indices.size(),
max_vertices,
max_triangles);
// Allocate output buffers
std::vector<meshopt_Meshlet> meshlets(max_meshlets);
std::vector<unsigned int> meshlet_vertices(indices.size());
std::vector<unsigned char> meshlet_triangles(indices.size());
// Generate meshlets from index buffer and vertex positions
size_t meshlet_count = meshopt_buildMeshlets(
meshlets.data(),
meshlet_vertices.data(),
meshlet_triangles.data(),
indices.data(),
indices.size(),
&vertices[0].x,
vertices.size(),
sizeof(Vertex),
max_vertices,
max_triangles,
cone_weight);
For AMD-specific optimization, replace max_triangles with a value matching max_vertices:
// AMD symmetric configuration (64/64) for pre-RDNA2 or fill-rate testing
const size_t max_vertices = 64;
const size_t max_triangles = 64; // Matched to vertices
Hardware Constraints and Validation
The absolute limits enforced by the library are defined in src/meshoptimizer.h, which declares max_vertices ≤ 256 and max_triangles ≤ 512 as hard bounds for the API. The implementation in src/meshletutils.cpp validates these constraints during the bounds calculation phase and will clamp or reject invalid configurations.
The demo application in demo/main.cpp provides a complete reference implementation, showing how to integrate meshlet generation with vertex optimization and index buffer remapping. The JavaScript wrapper in js/meshopt_clusterizer.js exposes the same parameters to WebGPU users.
Summary
- NVIDIA (Turing+): Use 64 vertices and 126 triangles per meshlet for optimal vertex reuse and occupancy.
- AMD (RDNA2+): 64/126 is safe, but consider 64/64 or 128/128 symmetric limits to maximize hardware utilization.
- Hardware limits: Absolute caps of 256 vertices and 512 triangles are enforced in
src/meshoptimizer.h. - API workflow: Call
meshopt_buildMeshletsBoundto size buffers, thenmeshopt_buildMeshletsto generate clusters with your chosen max_vertices and max_triangles values.
Frequently Asked Questions
What happens if I exceed 64 vertices or 126 triangles on NVIDIA?
The mesh shader hardware interface on NVIDIA Turing and newer GPUs has driver-enforced limits of 64 vertices and 126 triangles per meshlet. Exceeding these values may cause pipeline creation failures or undefined behavior in the mesh shader stage, as the shader output registers cannot accommodate larger primitive counts according to the implementation notes in the meshoptimizer README.
Can I use the same meshlet limits for all GPU vendors?
Yes, the 64/126 configuration works across both NVIDIA and AMD RDNA2 hardware. However, AMD GPUs prior to RDNA2 and some RDNA2 implementations may achieve higher throughput with symmetric limits (e.g., 64/64), matching vertices to triangles. Profile both configurations on your target hardware to determine optimal performance characteristics.
How does meshopt_buildMeshletsBound calculate the maximum meshlet count?
The function defined in src/meshletutils.cpp computes a conservative upper bound based on the total index count and your specified max_vertices and max_triangles limits. It accounts for the worst-case scenario where every meshlet is fully subdivided, ensuring you allocate sufficient storage for the meshlets, meshlet_vertices, and meshlet_triangles arrays before calling meshopt_buildMeshlets.
Should max_triangles always match max_vertices for AMD?
Not necessarily on RDNA2 and newer, where asymmetric limits like 64/126 are fully supported. However, many AMD optimization guides suggest matched limits (64/64 or 128/128) because older AMD GPUs perform best with symmetric configurations, and balanced meshlets can improve cache coherency across generations. Test your specific mesh to determine the optimal ratio.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →