How Mesh Simplification Is Performed in TRELLIS.2: GPU-Based Decimation with nvdiffrast Post-Processing

Mesh simplification in TRELLIS.2 uses a two-stage pipeline: geometric decimation via the CUDA-based cumesh library, followed by differentiable texture baking using nvdiffrast to preserve high-fidelity PBR materials.

TRELLIS.2 implements a sophisticated mesh post-processing workflow that separates geometry reduction from texture generation. This design allows aggressive simplification for performance while maintaining visual fidelity through accurate attribute reprojection. The Microsoft research project, TRELLIS.2, combines these GPU-accelerated techniques to handle the large-scale meshes typical of neural volumetric reconstruction.

Geometric Simplification with CuMesh

The first stage reduces triangle count using Mesh.simplify, a wrapper around the proprietary cumesh.CuMesh CUDA extension.

Mesh Class Initialization and Simplification

Located in trellis2/representations/mesh/base.py, the Mesh class provides the primary interface:

from trellis2.representations.mesh.base import Mesh

# Create mesh from tensors

mesh = Mesh(vertices, faces)           # Line 8-16

# Simplify to target vertex count

mesh.simplify(target=1_000_000)        # Line 71-78

The simplification method performs four GPU-resident operations:

  • Device transfer — moves vertex and face tensors to CUDA memory with .cuda()
  • CuMesh initialization — creates cumesh.CuMesh instance and populates geometry via mesh.init
  • Edge-collapse decimation — executes cumesh.CuMesh.simplify(target, ...) using internal GPU kernels
  • Result extraction — reads simplified geometry back through mesh.read()

Because the entire operation stays on-device, TRELLIS.2 can process meshes with millions of triangles. The project's example.py demonstrates this with mesh.simplify(16777216) — a 16.7 million vertex target that respects nvdiffrast's internal limits while preserving sufficient geometric detail.

Post-Generation Processing with nvdiffrast

After geometric simplification comes the differentiable rasterization stage in o-voxel/o_voxel/postprocess.py. Here, nvdiffrast enables texture baking that compensates for geometric loss by sampling the original volumetric representation.

Rasterization Context Setup

import nvdiffrast.torch as dr

# Initialize CUDA rasterization context

ctx = dr.RasterizeCudaContext()        # Line 30

Chunked Mesh Rasterization

Large meshes are processed in 100,000-face chunks to manage GPU memory:

rast_chunk, _ = dr.rasterize(
    ctx,
    uvs_rast,                          # UV coordinates in NDC space

    out_faces[i:i+100_000],            # Chunked face indices

    resolution=[texture_size, texture_size],
)                                      # Line 37-40

The rasterizer outputs:

  • Per-pixel face identifiers
  • Barycentric coordinates for interpolation
  • Depth information for visibility testing

Position Interpolation for Volumetric Sampling


# Interpolate 3D positions per texel

pos = dr.interpolate(
    out_vertices.unsqueeze(0), 
    rast, 
    out_faces
)[0][0]                                # Line 49

This critical step maps each texture pixel back to the original high-resolution mesh surface, not the simplified geometry. The recovered 3D positions then sample the sparse attribute volume using grid_sample_3d, extracting:

  • Base color
  • Metallic values
  • Roughness coefficients
  • Alpha transparency

Complete Simplification and Post-Processing Pipeline

Practical Implementation

The following example combines both stages as implemented in TRELLIS.2:


# Stage 1: Geometric simplification

from trellis2.representations.mesh.base import Mesh
import torch

verts = torch.randn(2_000_000, 3)      # 2M vertices

faces = torch.randint(0, 2_000_000, (4_000_000, 3))

mesh = Mesh(verts, faces)
mesh.simplify(target=1_000_000, verbose=True)

# Stage 2: nvdiffrast texture baking

from o_voxel.o_voxel import postprocess

glb_mesh = postprocess.to_glb(
    vertices=mesh.vertices,
    faces=mesh.faces,
    attr_volume=attr_volume,           # (C, D, H, W) feature tensor

    coords=coords,                     # (N, 3) voxel coordinates

    attr_layout=attr_layout,           # Channel layout definition

    aabb=aabb,                         # Axis-aligned bounding box

    voxel_size=voxel_size,
    decimation_target=800_000,         # Final vertex count

    texture_size=2048,
    remesh=False,
    verbose=True,
)

glb_mesh.export("result.glb")

Key Source Files and Implementation Details

File Purpose Critical Functions
trellis2/representations/mesh/base.py Core mesh operations Mesh.simplify (Lines 71-78), Mesh.uv_unwrap
o-voxel/o_voxel/postprocess.py nvdiffrast integration to_glb (Lines 14-331), rasterization, texture composition
example.py End-to-end demonstration Calls mesh.simplify(16777216) before post-processing
trellis2/renderers/mesh_renderer.py Renderer integration Shows nvdiffrast import patterns

Design Rationale: Why the Two-Stage Approach?

The separation of concerns in TRELLIS.2's mesh simplification architecture delivers three concrete advantages:

  • Performance — cumesh performs edge-collapse operations entirely in CUDA, processing millions of triangles orders of magnitude faster than CPU-based methods like pymeshlab or open3d
  • Fidelity preservation — By rasterizing UV coordinates and interpolating positions on the original mesh, texture detail remains decoupled from geometric complexity
  • Differentiability — The nvdiffrast pipeline maintains gradient flow, enabling future extensions such as inverse rendering or learned texture refinement

Summary

  • Geometric simplification in TRELLIS.2 uses Mesh.simplify, a GPU-accelerated wrapper around the cumesh.CuMesh library
  • nvdiffrast post-processing handles differentiable rasterization and texture baking after decimation, not geometry simplification itself
  • The UV unwrapping and position interpolation workflow ensures high-quality PBR textures despite aggressive vertex reduction
  • All operations remain CUDA-resident, scaling to production meshes with 16+ million vertices
  • The modular design preserves gradient information for potential optimization extensions

Frequently Asked Questions

Does nvdiffrast perform the actual mesh simplification in TRELLIS.2?

No. nvdiffrast provides differentiable rasterization for texture baking after simplification. The geometric decimation is handled by cumesh.CuMesh through the Mesh.simplify method in trellis2/representations/mesh/base.py. nvdiffrast's role begins after the mesh vertex count has been reduced.

Why does TRELLIS.2 use chunking during nvdiffrast rasterization?

The to_glb function processes meshes in 100,000-face chunks to manage GPU memory limits. Large reconstructed meshes from neural volumetric generation can exceed single-pass rasterization capacity. Chunking allows consistent 2048×2048 or larger texture generation regardless of input mesh complexity.

How does texture quality remain high after aggressive simplification?

TRELLIS.2 decouples geometry from appearance. While the mesh is reduced to meet performance targets, the dr.interpolate operation maps each texture pixel back to the original high-resolution surface positions. These positions sample the full volumetric attribute grid, preserving material detail independent of triangle count.

What is the CuMesh library and where does it come from?

cumesh is the proprietary CUDA mesh processing extension used internally by TRELLIS.2. It implements GPU-accelerated edge-collapse simplification. The library is wrapped by the Mesh class and initialized via cumesh.CuMesh within the simplify method.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →