# How Mesh Simplification Is Performed in TRELLIS.2: GPU-Based Decimation with nvdiffrast Post-Processing

> Discover how TRELLIS.2 simplifies meshes using GPU-accelerated decimation and nvdiffrast for high-fidelity texture baking. Learn about the two-stage pipeline for efficient 3D asset processing.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: how-to-guide
- Published: 2026-08-04

---

**Mesh simplification in TRELLIS.2 uses a two-stage pipeline: geometric decimation via the CUDA-based `cumesh` library, followed by differentiable texture baking using nvdiffrast to preserve high-fidelity PBR materials.**

TRELLIS.2 implements a sophisticated mesh post-processing workflow that separates geometry reduction from texture generation. This design allows aggressive simplification for performance while maintaining visual fidelity through accurate attribute reprojection. The Microsoft research project, [TRELLIS.2](https://github.com/microsoft/TRELLIS.2), combines these GPU-accelerated techniques to handle the large-scale meshes typical of neural volumetric reconstruction.

## Geometric Simplification with CuMesh

The first stage reduces triangle count using `Mesh.simplify`, a wrapper around the proprietary `cumesh.CuMesh` CUDA extension.

### Mesh Class Initialization and Simplification

Located in [`trellis2/representations/mesh/base.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/representations/mesh/base.py), the `Mesh` class provides the primary interface:

```python
from trellis2.representations.mesh.base import Mesh

# Create mesh from tensors

mesh = Mesh(vertices, faces)           # Line 8-16

# Simplify to target vertex count

mesh.simplify(target=1_000_000)        # Line 71-78

```

The simplification method performs four GPU-resident operations:

- **Device transfer** — moves vertex and face tensors to CUDA memory with `.cuda()`
- **CuMesh initialization** — creates `cumesh.CuMesh` instance and populates geometry via `mesh.init`
- **Edge-collapse decimation** — executes `cumesh.CuMesh.simplify(target, ...)` using internal GPU kernels
- **Result extraction** — reads simplified geometry back through `mesh.read()`

Because the entire operation stays on-device, TRELLIS.2 can process meshes with millions of triangles. The project's [`example.py`](https://github.com/microsoft/TRELLIS.2/blob/main/example.py) demonstrates this with `mesh.simplify(16777216)` — a 16.7 million vertex target that respects nvdiffrast's internal limits while preserving sufficient geometric detail.

## Post-Generation Processing with nvdiffrast

After geometric simplification comes the **differentiable rasterization stage** in [`o-voxel/o_voxel/postprocess.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o-voxel/o_voxel/postprocess.py). Here, nvdiffrast enables texture baking that compensates for geometric loss by sampling the original volumetric representation.

### Rasterization Context Setup

```python
import nvdiffrast.torch as dr

# Initialize CUDA rasterization context

ctx = dr.RasterizeCudaContext()        # Line 30

```

### Chunked Mesh Rasterization

Large meshes are processed in 100,000-face chunks to manage GPU memory:

```python
rast_chunk, _ = dr.rasterize(
    ctx,
    uvs_rast,                          # UV coordinates in NDC space

    out_faces[i:i+100_000],            # Chunked face indices

    resolution=[texture_size, texture_size],
)                                      # Line 37-40

```

The rasterizer outputs:
- Per-pixel face identifiers
- Barycentric coordinates for interpolation
- Depth information for visibility testing

### Position Interpolation for Volumetric Sampling

```python

# Interpolate 3D positions per texel

pos = dr.interpolate(
    out_vertices.unsqueeze(0), 
    rast, 
    out_faces
)[0][0]                                # Line 49

```

This critical step maps each texture pixel back to the **original high-resolution mesh surface**, not the simplified geometry. The recovered 3D positions then sample the sparse attribute volume using `grid_sample_3d`, extracting:
- **Base color**
- **Metallic values**
- **Roughness coefficients**
- **Alpha transparency**

## Complete Simplification and Post-Processing Pipeline

### Practical Implementation

The following example combines both stages as implemented in [TRELLIS.2](https://github.com/microsoft/TRELLIS.2):

```python

# Stage 1: Geometric simplification

from trellis2.representations.mesh.base import Mesh
import torch

verts = torch.randn(2_000_000, 3)      # 2M vertices

faces = torch.randint(0, 2_000_000, (4_000_000, 3))

mesh = Mesh(verts, faces)
mesh.simplify(target=1_000_000, verbose=True)

# Stage 2: nvdiffrast texture baking

from o_voxel.o_voxel import postprocess

glb_mesh = postprocess.to_glb(
    vertices=mesh.vertices,
    faces=mesh.faces,
    attr_volume=attr_volume,           # (C, D, H, W) feature tensor

    coords=coords,                     # (N, 3) voxel coordinates

    attr_layout=attr_layout,           # Channel layout definition

    aabb=aabb,                         # Axis-aligned bounding box

    voxel_size=voxel_size,
    decimation_target=800_000,         # Final vertex count

    texture_size=2048,
    remesh=False,
    verbose=True,
)

glb_mesh.export("result.glb")

```

## Key Source Files and Implementation Details

| File | Purpose | Critical Functions |
|------|---------|------------------|
| [`trellis2/representations/mesh/base.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/representations/mesh/base.py) | Core mesh operations | `Mesh.simplify` (Lines 71-78), `Mesh.uv_unwrap` |
| [`o-voxel/o_voxel/postprocess.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o-voxel/o_voxel/postprocess.py) | nvdiffrast integration | `to_glb` (Lines 14-331), rasterization, texture composition |
| [`example.py`](https://github.com/microsoft/TRELLIS.2/blob/main/example.py) | End-to-end demonstration | Calls `mesh.simplify(16777216)` before post-processing |
| [`trellis2/renderers/mesh_renderer.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/renderers/mesh_renderer.py) | Renderer integration | Shows nvdiffrast import patterns |

## Design Rationale: Why the Two-Stage Approach?

The separation of concerns in TRELLIS.2's mesh simplification architecture delivers three concrete advantages:

- **Performance** — `cumesh` performs edge-collapse operations entirely in CUDA, processing millions of triangles orders of magnitude faster than CPU-based methods like `pymeshlab` or `open3d`
- **Fidelity preservation** — By rasterizing UV coordinates and interpolating positions on the original mesh, texture detail remains decoupled from geometric complexity
- **Differentiability** — The nvdiffrast pipeline maintains gradient flow, enabling future extensions such as inverse rendering or learned texture refinement

## Summary

- **Geometric simplification** in TRELLIS.2 uses `Mesh.simplify`, a GPU-accelerated wrapper around the `cumesh.CuMesh` library
- **nvdiffrast post-processing** handles differentiable rasterization and texture baking after decimation, not geometry simplification itself
- The **UV unwrapping and position interpolation** workflow ensures high-quality PBR textures despite aggressive vertex reduction
- All operations remain **CUDA-resident**, scaling to production meshes with 16+ million vertices
- The modular design preserves **gradient information** for potential optimization extensions

## Frequently Asked Questions

### Does nvdiffrast perform the actual mesh simplification in TRELLIS.2?

No. nvdiffrast provides differentiable rasterization for texture baking after simplification. The geometric decimation is handled by `cumesh.CuMesh` through the `Mesh.simplify` method in [`trellis2/representations/mesh/base.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/representations/mesh/base.py). nvdiffrast's role begins after the mesh vertex count has been reduced.

### Why does TRELLIS.2 use chunking during nvdiffrast rasterization?

The `to_glb` function processes meshes in 100,000-face chunks to manage GPU memory limits. Large reconstructed meshes from neural volumetric generation can exceed single-pass rasterization capacity. Chunking allows consistent 2048×2048 or larger texture generation regardless of input mesh complexity.

### How does texture quality remain high after aggressive simplification?

TRELLIS.2 decouples geometry from appearance. While the mesh is reduced to meet performance targets, the `dr.interpolate` operation maps each texture pixel back to the original high-resolution surface positions. These positions sample the full volumetric attribute grid, preserving material detail independent of triangle count.

### What is the CuMesh library and where does it come from?

`cumesh` is the proprietary CUDA mesh processing extension used internally by TRELLIS.2. It implements GPU-accelerated edge-collapse simplification. The library is wrapped by the `Mesh` class and initialized via `cumesh.CuMesh` within the `simplify` method.