How to Use meshopt_partitionClusters for Nanite-Style Hierarchical LOD
Use meshopt_partitionClusters to group mesh clusters into partitions of approximately equal size while preserving vertex connectivity, then build a DAG hierarchy where each partition becomes a LOD node that can be culled based on screen-space error.
The meshopt_partitionClusters function in the meshoptimizer library provides the core clustering mechanism needed to implement Nanite-style level-of-detail (LOD) systems. This function takes a flat list of mesh clusters and organizes them into hierarchical groups that maintain spatial locality and shared vertex connectivity. By combining this partitioner with spatial sorting and DAG construction, you can create a virtualized geometry pipeline similar to Unreal Engine 5's Nanite.
How meshopt_partitionClusters Works
The partitioner is declared in src/meshoptimizer.h (lines 842-452) and implemented in src/partition.cpp. It transforms a flat array of clusters into groups of approximately equal size while keeping clusters that share vertices together, which is essential for minimizing border artifacts in LOD transitions.
Input Data Requirements
The function requires three core inputs:
cluster_indices– A flattened array of all vertex indices used by each clustercluster_index_counts– An array specifying the per-cluster index counttarget_partition_size– The desired number of clusters per partition
The actual partition size may be up to +⅓ larger than the target (e.g., target 24 yields maximum 32). This flexibility allows the algorithm to optimize grouping while respecting connectivity constraints.
Optional Spatial Grouping
When you provide vertex positions via the vertex_positions parameter, the algorithm prioritizes keeping spatially close clusters together. If positions are NULL, the partitioner relies solely on vertex connectivity to group clusters.
Implementing the Nanite-Style Pipeline
The demo files demo/clusterlod.h and demo/nanite.cpp demonstrate a complete workflow that mirrors the Nanite pipeline. The following steps transform raw mesh clusters into a hierarchical LOD structure.
Step 1: Flatten Cluster Index Data
First, collect all cluster indices into a contiguous array and record the count per cluster:
std::vector<unsigned int> cluster_indices;
std::vector<unsigned int> cluster_counts(pending.size());
size_t total = 0;
for (size_t i = 0; i < pending.size(); ++i) {
const Cluster& c = clusters[pending[i]];
cluster_counts[i] = (unsigned)c.indices.size();
for (unsigned v : c.indices)
cluster_indices.push_back(remap[v]); // remap to position-only indices
total += c.indices.size();
}
This pattern appears in demo/clusterlod.h (lines 341-350).
Step 2: Partition the Clusters
Invoke meshopt_partitionClusters with the flattened data:
std::vector<unsigned int> cluster_part(pending.size());
size_t partition_count = meshopt_partitionClusters(
&cluster_part[0],
&cluster_indices[0], cluster_indices.size(),
&cluster_counts[0], cluster_counts.size(),
config.partition_spatial ? mesh.vertex_positions : NULL,
remap.size(),
mesh.vertex_positions_stride,
config.partition_size);
This call returns the number of partitions and fills cluster_part with a partition ID for each input cluster (see demo/clusterlod.h, lines 341-346).
Step 3: Spatially Sort Partitions (Optional)
Nanite sorts groups by their spatial centroid to improve cache locality when rendering cuts. Use meshopt_spatialSortRemap to achieve this:
std::vector<float> partition_point(partition_count * 3);
for (size_t i = 0; i < pending.size(); ++i)
memcpy(&partition_point[cluster_part[i] * 3],
clusters[pending[i]].bounds.center, sizeof(float) * 3);
std::vector<unsigned int> partition_remap(partition_count);
meshopt_spatialSortRemap(partition_remap.data(),
partition_point.data(),
partition_count,
sizeof(float) * 3);
See demo/clusterlod.h (lines 353-363) for the reference implementation.
Step 4: Distribute Clusters into Groups
Apply the spatial remap (if used) and organize clusters into their final partitions:
std::vector<std::vector<int>> partitions(partition_count);
for (size_t i = 0; i < pending.size(); ++i) {
unsigned pid = partition_remap.empty() ? cluster_part[i]
: partition_remap[cluster_part[i]];
partitions[pid].push_back(pending[i]);
}
This distribution logic appears in demo/clusterlod.h (lines 365-368).
Step 5: Construct the DAG Hierarchy
Each partition becomes a group in the LOD DAG. The clodBuild routine (used in demo/nanite.cpp) consumes these groups to create a hierarchy of clusters that can be cut based on screen-space error:
clodConfig cfg = clodDefaultConfig(/*max_triangles=*/128);
cfg.partition_spatial = true;
cfg.partition_sort = true;
cfg.partition_size = 16;
size_t totalClusters = clodBuild(cfg, mesh,
[&](clodGroup g, const clodCluster* cl, size_t count) -> int {
// Store groups and clusters for rendering
return int(current_group_id);
});
The Nanite demo evaluates boundsError per cluster (see demo/nanite.cpp, lines 27-32) to drive the screen-space error metric for culling.
Complete Implementation Example
Here is a consolidated workflow showing the full pipeline from cluster data to hierarchical groups:
// Flatten cluster data
std::vector<unsigned int> flatIndices;
std::vector<unsigned int> perClusterCount(pending.size());
for (size_t i = 0; i < pending.size(); ++i) {
const Cluster& c = clusters[pending[i]];
perClusterCount[i] = (unsigned)c.indices.size();
for (unsigned v : c.indices)
flatIndices.push_back(remap[v]);
}
// Partition clusters
std::vector<unsigned int> partId(pending.size());
size_t partCount = meshopt_partitionClusters(
partId.data(),
flatIndices.data(), flatIndices.size(),
perClusterCount.data(), perClusterCount.size(),
config.partition_spatial ? mesh.vertex_positions : nullptr,
remap.size(),
mesh.vertex_positions_stride,
config.partition_size);
// Optional: Sort partitions spatially
std::vector<unsigned int> partRemap;
if (config.partition_sort) {
std::vector<float> centroids(partCount * 3);
for (size_t i = 0; i < pending.size(); ++i)
memcpy(¢roids[partId[i] * 3],
clusters[pending[i]].bounds.center,
sizeof(float) * 3);
partRemap.resize(partCount);
meshopt_spatialSortRemap(partRemap.data(),
centroids.data(),
partCount,
sizeof(float) * 3);
}
// Build final groups
std::vector<std::vector<int>> groups(partCount);
for (size_t i = 0; i < pending.size(); ++i) {
unsigned pid = partRemap.empty() ? partId[i] : partRemap[partId[i]];
groups[pid].push_back(pending[i]);
}
Summary
meshopt_partitionClustersgroups clusters into partitions of approximatelytarget_partition_sizewhile preserving vertex connectivity- The partitioner accepts optional vertex positions to improve spatial locality, reducing cache misses during rendering
- Spatial sorting via
meshopt_spatialSortRemapmimics Nanite's centroid-based organization for optimal memory access patterns - Each partition becomes a node in a DAG hierarchy, consumed by
clodBuildto create a screen-space error-driven LOD system - Reference implementations in
demo/clusterlod.handdemo/nanite.cppprovide production-ready examples of the complete pipeline
Frequently Asked Questions
What is the maximum partition size generated by meshopt_partitionClusters?
The actual partition size can be up to +⅓ larger than your target_partition_size. For example, if you specify a target of 24 clusters per partition, the algorithm may generate partitions containing up to 32 clusters. This flexibility allows the partitioner to optimize grouping while respecting vertex connectivity constraints.
Why do I need to flatten cluster indices before partitioning?
meshopt_partitionClusters requires a contiguous array of vertex indices (cluster_indices) and a separate count array (cluster_index_counts) because it processes the input as a single stream rather than an array of arrays. This design minimizes memory overhead and allows the algorithm to efficiently traverse the entire mesh topology in a single pass.
How does spatial sorting improve Nanite-style rendering performance?
Spatial sorting using meshopt_spatialSortRemap organizes partitions by their spatial centroids, ensuring that clusters occupying nearby world space are stored contiguously in memory. This improves cache locality when rendering a DAG cut, as the GPU can access spatially coherent clusters with fewer memory fetches, directly mirroring the optimization described in the original Nanite paper.
Can I use meshopt_partitionClusters without vertex positions?
Yes. If you pass NULL for the vertex_positions parameter, the partitioner relies solely on vertex connectivity to group clusters. While this still produces valid partitions for hierarchical LOD, providing positions enables the algorithm to additionally optimize for spatial locality, which typically produces better results for frustum culling and occlusion testing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →