# How the Dual-Grid Representation Enables Instant Mesh-to-Latent Conversion in TRELLIS.2

> Discover how TRELLIS.2 achieves instant mesh-to-latent conversion using a dual-grid representation, eliminating runtime rasterization for efficient 3D processing. Explore the sparse, quantized approach.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: deep-dive
- Published: 2026-08-03

---

**The TRELLIS.2 framework converts raw 3D meshes into compressed latent vectors in a single forward pass by pre-computing a sparse, quantized dual-grid representation that eliminates runtime rasterization overhead.**

The TRELLIS.2 codebase addresses a critical bottleneck in 3D generative modeling: bridging irregular mesh topology with regular convolutional encoders. By leveraging a **dual-grid representation**—implemented across [`o_voxel/convert/flexible_dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/convert/flexible_dual_grid.py), [`data_toolkit/dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/dual_grid.py), and [`trellis2/models/sparse_structure_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_vae.py)—the system transforms polygonal meshes into sparse voxel grids that the **Sparse Structure VAE** can consume instantly without reconstruction.

## The Three-Stage Dual-Grid Pipeline

### Stage 1: Fast Voxelization via `mesh_to_flexible_dual_grid`

The conversion begins in [`o_voxel/convert/flexible_dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/convert/flexible_dual_grid.py), where the `mesh_to_flexible_dual_grid` function receives raw vertices and faces, computes an axis-aligned bounding box (AABB), and determines voxel resolution via `voxel_size` or `grid_size` parameters.

The function delegates to the optimized C++ backend `_C.mesh_to_flexible_dual_grid_cpu`, which performs three critical operations:

- **Voxel occupancy detection** – Rasterizes every mesh triangle into a sparse voxel grid.
- **Dual-vertex solving** – Solves a quadratic-energy-function (QEF) per voxel to compute a **dual vertex** that approximates the original surface, controlled by `face_weight`, `boundary_weight`, and `regularization_weight` parameters.
- **Intersection flag encoding** – Packs a 3-bit `intersected` flag per voxel indicating which faces the mesh penetrates.

The function returns three quantized tensors:

```python
voxel_indices,      # (N, 3) integer voxel coordinates

dual_vertices,      # (N, 3) uint8 dual-vertex positions (0-255)

intersected         # (N, 1) uint8 flag with 3 bits packed

```

Because the output is **sparse** (only occupied voxels stored) and **quantized** (uint8), it requires no additional preprocessing before entering PyTorch modules.

### Stage 2: Persistent Storage in `.vxz` Format

The [`data_toolkit/dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/dual_grid.py) script persists the sparse representation to disk using `o_voxel.io.write_vxz`. The `.vxz` format stores voxel indices alongside quantized dual vertices and intersection flags, enabling **instant reloading** via `read_vxz` without recomputing the expensive voxelization.

This persistent cache means the quadratic energy solving and rasterization happen once per mesh, not per training iteration.

### Stage 3: Direct Latent Encoding with Sparse Structure VAE

The `FlexiDualGridDataset` class in [`trellis2/datasets/flexi_dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/datasets/flexi_dual_grid.py) loads `.vxz` files and constructs a `SparseTensor` of shape `(num_active_voxels, C)`. Features combine normalized dual-vertex positions (`dual_vert.float() / 255.0`) and intersection flags.

This sparse tensor feeds directly into `SparseStructureEncoder` ([`trellis2/models/sparse_structure_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_vae.py)), which treats the grid as a 3D image and outputs latent parameters (`mean`, `logvar`) in one forward pass:

```python
latent_mean, latent_logvar = encoder(sparse_vox)

```

## Why the Conversion is "Instant"

The mesh-to-latent pipeline achieves instant conversion through three architectural decisions:

- **No mesh reconstruction** – The dual-grid encodes surface geometry via pre-computed dual vertices, eliminating the encoder's need to rasterize or remesh the object.
- **Sparse representation** – Only active voxels are processed, drastically reducing memory and computational load compared to dense voxel grids.
- **Quantized features** – uint8 storage minimizes GPU memory bandwidth, allowing the encoder to read geometry with minimal cache overhead.

## Practical Implementation

### Converting Meshes to Dual-Grid Format

The following pattern—derived from [`data_toolkit/dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/dual_grid.py)—voxelizes a mesh and persists it as `.vxz`:

```python
import torch, pickle, numpy as np, o_voxel
from pathlib import Path

def convert_and_save(mesh_pickle_path, out_dir, sha256, resolution=256):
    # Load mesh dump

    with open(mesh_pickle_path, 'rb') as f:
        dump = pickle.load(f)

    # Assemble vertices & faces

    verts, faces = [], []
    start = 0
    for obj in dump['objects']:
        if obj['vertices'].size == 0 or obj['faces'].size == 0:
            continue
        verts.append(obj['vertices'])
        faces.append(obj['faces'] + start)
        start += obj['vertices'].shape[0]

    vertices = torch.from_numpy(np.concatenate(verts, axis=0)).float()
    faces = torch.from_numpy(np.concatenate(faces, axis=0)).long()

    # Voxelize to dual grid

    voxel_idx, dual_vert, intersected = o_voxel.convert.mesh_to_flexible_dual_grid(
        vertices, faces,
        grid_size=resolution,
        aabb=[[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]],
        face_weight=1.0, 
        boundary_weight=0.2, 
        regularization_weight=1e-2
    )

    # Save as .vxz

    out_path = Path(out_dir) / f'dual_grid_{resolution}' / f'{sha256}.vxz'
    out_path.parent.mkdir(parents=True, exist_ok=True)
    
    intersected_packed = (intersected[:, 0:1] + 
                         2 * intersected[:, 1:2] + 
                         4 * intersected[:, 2:3]).type(torch.uint8)
    
    o_voxel.io.write_vxz(
        str(out_path), 
        voxel_idx, 
        {
            'vertices': (dual_vert * 255).type(torch.uint8),
            'intersected': intersected_packed
        }
    )

```

### Encoding Dual-Grids to Latent Vectors

Load the cached dual-grid and encode it using the Sparse Structure VAE:

```python
import torch
from trellis2.datasets.flexi_dual_grid import FlexiDualGridDataset
from trellis2.models.sparse_structure_vae import SparseStructureEncoder

# Initialize dataset

dataset = FlexiDualGridDataset(
    roots={
        'mesh_dump': '/path/to/mesh_dump',
        'dual_grid': '/path/to/dual_grid'
    },
    resolution=256, 
    max_active_voxels=500_000
)

# Load sample (already formatted as SparseTensor)

sample = dataset[0]
sparse_vox = sample['vertices']

# Build encoder

encoder = SparseStructureEncoder(
    in_channels=4,            # 3 dual-vertex + 1 intersected flag

    latent_channels=256,
    num_res_blocks=2,
    channels=[64, 128, 256],
    norm_type='layer',
    use_fp16=False
)

# Instant latent extraction

latent_mean, latent_logvar = encoder(sparse_vox)

```

## Summary

- The **dual-grid representation** pre-computes voxelized geometry using `mesh_to_flexible_dual_grid` in [`o_voxel/convert/flexible_dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/convert/flexible_dual_grid.py), solving QEFs to generate dual vertices and intersection flags.
- **`.vxz` files** store sparse, quantized grids via [`data_toolkit/dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/dual_grid.py), enabling zero-overhead reloading without re-voxelization.
- `FlexiDualGridDataset` converts stored grids into `SparseTensor` objects that `SparseStructureEncoder` processes in a single forward pass to produce latent codes.
- **uint8 quantization** and sparse structure minimize GPU bandwidth and memory, delivering true instant mesh-to-latent conversion.

## Frequently Asked Questions

### What makes the dual-grid representation "instant" compared to traditional voxelization?

Traditional pipelines rasterize meshes on-the-fly during encoding, which requires significant compute per forward pass. The TRELLIS.2 dual-grid pre-computes voxelization using the C++ backend `_C.mesh_to_flexible_dual_grid_cpu`, stores results as sparse `.vxz` files, and loads them as `SparseTensor` objects. This moves the expensive rasterization and QEF solving offline, leaving only a lightweight sparse convolution for the encoder.

### How does the `mesh_to_flexible_dual_grid` function handle surface approximation?

The function solves a **quadratic energy function (QEF)** per voxel to compute dual vertices that best approximate the original mesh surface. The energy minimization is controlled by three weights: `face_weight` (surface fidelity), `boundary_weight` (edge handling), and `regularization_weight` (numerical stability). This produces quantized `uint8` vertices that preserve geometric detail without requiring the original mesh topology at encoding time.

### Why are the dual-vertices and intersection flags stored as `uint8`?

Quantization to `uint8` reduces memory footprint by 75% compared to `float32` coordinates, allowing the `SparseStructureEncoder` to read geometry data with minimal GPU cache bandwidth. The `FlexiDualGridDataset` normalizes these values back to `[0, 1]` range (`dual_vert.float() / 255.0`) before feeding them to the encoder, maintaining numerical precision while maximizing throughput.

### Can I convert meshes to latents without saving intermediate `.vxz` files?

Yes, though not recommended for production training. You can call `o_voxel.convert.mesh_to_flexible_dual_grid` directly in your data loader to generate `voxel_indices`, `dual_vertices`, and `intersected` tensors on-the-fly, then construct a `SparseTensor` manually. However, this repeats the expensive QEF solving for every epoch, whereas the `.vxz` workflow amortizes this cost to a one-time preprocessing step.