How Mesh Simplification Is Performed in TRELLIS.2: GPU-Based Decimation with nvdiffrast Post-Processing
Mesh simplification in TRELLIS.2 uses a two-stage pipeline: geometric decimation via the CUDA-based cumesh library, followed by differentiable texture baking using nvdiffrast to preserve high-fidelity PBR materials.
TRELLIS.2 implements a sophisticated mesh post-processing workflow that separates geometry reduction from texture generation. This design allows aggressive simplification for performance while maintaining visual fidelity through accurate attribute reprojection. The Microsoft research project, TRELLIS.2, combines these GPU-accelerated techniques to handle the large-scale meshes typical of neural volumetric reconstruction.
Geometric Simplification with CuMesh
The first stage reduces triangle count using Mesh.simplify, a wrapper around the proprietary cumesh.CuMesh CUDA extension.
Mesh Class Initialization and Simplification
Located in trellis2/representations/mesh/base.py, the Mesh class provides the primary interface:
from trellis2.representations.mesh.base import Mesh
# Create mesh from tensors
mesh = Mesh(vertices, faces) # Line 8-16
# Simplify to target vertex count
mesh.simplify(target=1_000_000) # Line 71-78
The simplification method performs four GPU-resident operations:
- Device transfer — moves vertex and face tensors to CUDA memory with
.cuda() - CuMesh initialization — creates
cumesh.CuMeshinstance and populates geometry viamesh.init - Edge-collapse decimation — executes
cumesh.CuMesh.simplify(target, ...)using internal GPU kernels - Result extraction — reads simplified geometry back through
mesh.read()
Because the entire operation stays on-device, TRELLIS.2 can process meshes with millions of triangles. The project's example.py demonstrates this with mesh.simplify(16777216) — a 16.7 million vertex target that respects nvdiffrast's internal limits while preserving sufficient geometric detail.
Post-Generation Processing with nvdiffrast
After geometric simplification comes the differentiable rasterization stage in o-voxel/o_voxel/postprocess.py. Here, nvdiffrast enables texture baking that compensates for geometric loss by sampling the original volumetric representation.
Rasterization Context Setup
import nvdiffrast.torch as dr
# Initialize CUDA rasterization context
ctx = dr.RasterizeCudaContext() # Line 30
Chunked Mesh Rasterization
Large meshes are processed in 100,000-face chunks to manage GPU memory:
rast_chunk, _ = dr.rasterize(
ctx,
uvs_rast, # UV coordinates in NDC space
out_faces[i:i+100_000], # Chunked face indices
resolution=[texture_size, texture_size],
) # Line 37-40
The rasterizer outputs:
- Per-pixel face identifiers
- Barycentric coordinates for interpolation
- Depth information for visibility testing
Position Interpolation for Volumetric Sampling
# Interpolate 3D positions per texel
pos = dr.interpolate(
out_vertices.unsqueeze(0),
rast,
out_faces
)[0][0] # Line 49
This critical step maps each texture pixel back to the original high-resolution mesh surface, not the simplified geometry. The recovered 3D positions then sample the sparse attribute volume using grid_sample_3d, extracting:
- Base color
- Metallic values
- Roughness coefficients
- Alpha transparency
Complete Simplification and Post-Processing Pipeline
Practical Implementation
The following example combines both stages as implemented in TRELLIS.2:
# Stage 1: Geometric simplification
from trellis2.representations.mesh.base import Mesh
import torch
verts = torch.randn(2_000_000, 3) # 2M vertices
faces = torch.randint(0, 2_000_000, (4_000_000, 3))
mesh = Mesh(verts, faces)
mesh.simplify(target=1_000_000, verbose=True)
# Stage 2: nvdiffrast texture baking
from o_voxel.o_voxel import postprocess
glb_mesh = postprocess.to_glb(
vertices=mesh.vertices,
faces=mesh.faces,
attr_volume=attr_volume, # (C, D, H, W) feature tensor
coords=coords, # (N, 3) voxel coordinates
attr_layout=attr_layout, # Channel layout definition
aabb=aabb, # Axis-aligned bounding box
voxel_size=voxel_size,
decimation_target=800_000, # Final vertex count
texture_size=2048,
remesh=False,
verbose=True,
)
glb_mesh.export("result.glb")
Key Source Files and Implementation Details
| File | Purpose | Critical Functions |
|---|---|---|
trellis2/representations/mesh/base.py |
Core mesh operations | Mesh.simplify (Lines 71-78), Mesh.uv_unwrap |
o-voxel/o_voxel/postprocess.py |
nvdiffrast integration | to_glb (Lines 14-331), rasterization, texture composition |
example.py |
End-to-end demonstration | Calls mesh.simplify(16777216) before post-processing |
trellis2/renderers/mesh_renderer.py |
Renderer integration | Shows nvdiffrast import patterns |
Design Rationale: Why the Two-Stage Approach?
The separation of concerns in TRELLIS.2's mesh simplification architecture delivers three concrete advantages:
- Performance —
cumeshperforms edge-collapse operations entirely in CUDA, processing millions of triangles orders of magnitude faster than CPU-based methods likepymeshlaboropen3d - Fidelity preservation — By rasterizing UV coordinates and interpolating positions on the original mesh, texture detail remains decoupled from geometric complexity
- Differentiability — The nvdiffrast pipeline maintains gradient flow, enabling future extensions such as inverse rendering or learned texture refinement
Summary
- Geometric simplification in TRELLIS.2 uses
Mesh.simplify, a GPU-accelerated wrapper around thecumesh.CuMeshlibrary - nvdiffrast post-processing handles differentiable rasterization and texture baking after decimation, not geometry simplification itself
- The UV unwrapping and position interpolation workflow ensures high-quality PBR textures despite aggressive vertex reduction
- All operations remain CUDA-resident, scaling to production meshes with 16+ million vertices
- The modular design preserves gradient information for potential optimization extensions
Frequently Asked Questions
Does nvdiffrast perform the actual mesh simplification in TRELLIS.2?
No. nvdiffrast provides differentiable rasterization for texture baking after simplification. The geometric decimation is handled by cumesh.CuMesh through the Mesh.simplify method in trellis2/representations/mesh/base.py. nvdiffrast's role begins after the mesh vertex count has been reduced.
Why does TRELLIS.2 use chunking during nvdiffrast rasterization?
The to_glb function processes meshes in 100,000-face chunks to manage GPU memory limits. Large reconstructed meshes from neural volumetric generation can exceed single-pass rasterization capacity. Chunking allows consistent 2048×2048 or larger texture generation regardless of input mesh complexity.
How does texture quality remain high after aggressive simplification?
TRELLIS.2 decouples geometry from appearance. While the mesh is reduced to meet performance targets, the dr.interpolate operation maps each texture pixel back to the original high-resolution surface positions. These positions sample the full volumetric attribute grid, preserving material detail independent of triangle count.
What is the CuMesh library and where does it come from?
cumesh is the proprietary CUDA mesh processing extension used internally by TRELLIS.2. It implements GPU-accelerated edge-collapse simplification. The library is wrapped by the Mesh class and initialized via cumesh.CuMesh within the simplify method.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →