How the Dual-Grid Representation Enables Instant Mesh-to-Latent Conversion in TRELLIS.2
The TRELLIS.2 framework converts raw 3D meshes into compressed latent vectors in a single forward pass by pre-computing a sparse, quantized dual-grid representation that eliminates runtime rasterization overhead.
The TRELLIS.2 codebase addresses a critical bottleneck in 3D generative modeling: bridging irregular mesh topology with regular convolutional encoders. By leveraging a dual-grid representation—implemented across o_voxel/convert/flexible_dual_grid.py, data_toolkit/dual_grid.py, and trellis2/models/sparse_structure_vae.py—the system transforms polygonal meshes into sparse voxel grids that the Sparse Structure VAE can consume instantly without reconstruction.
The Three-Stage Dual-Grid Pipeline
Stage 1: Fast Voxelization via mesh_to_flexible_dual_grid
The conversion begins in o_voxel/convert/flexible_dual_grid.py, where the mesh_to_flexible_dual_grid function receives raw vertices and faces, computes an axis-aligned bounding box (AABB), and determines voxel resolution via voxel_size or grid_size parameters.
The function delegates to the optimized C++ backend _C.mesh_to_flexible_dual_grid_cpu, which performs three critical operations:
- Voxel occupancy detection – Rasterizes every mesh triangle into a sparse voxel grid.
- Dual-vertex solving – Solves a quadratic-energy-function (QEF) per voxel to compute a dual vertex that approximates the original surface, controlled by
face_weight,boundary_weight, andregularization_weightparameters. - Intersection flag encoding – Packs a 3-bit
intersectedflag per voxel indicating which faces the mesh penetrates.
The function returns three quantized tensors:
voxel_indices, # (N, 3) integer voxel coordinates
dual_vertices, # (N, 3) uint8 dual-vertex positions (0-255)
intersected # (N, 1) uint8 flag with 3 bits packed
Because the output is sparse (only occupied voxels stored) and quantized (uint8), it requires no additional preprocessing before entering PyTorch modules.
Stage 2: Persistent Storage in .vxz Format
The data_toolkit/dual_grid.py script persists the sparse representation to disk using o_voxel.io.write_vxz. The .vxz format stores voxel indices alongside quantized dual vertices and intersection flags, enabling instant reloading via read_vxz without recomputing the expensive voxelization.
This persistent cache means the quadratic energy solving and rasterization happen once per mesh, not per training iteration.
Stage 3: Direct Latent Encoding with Sparse Structure VAE
The FlexiDualGridDataset class in trellis2/datasets/flexi_dual_grid.py loads .vxz files and constructs a SparseTensor of shape (num_active_voxels, C). Features combine normalized dual-vertex positions (dual_vert.float() / 255.0) and intersection flags.
This sparse tensor feeds directly into SparseStructureEncoder (trellis2/models/sparse_structure_vae.py), which treats the grid as a 3D image and outputs latent parameters (mean, logvar) in one forward pass:
latent_mean, latent_logvar = encoder(sparse_vox)
Why the Conversion is "Instant"
The mesh-to-latent pipeline achieves instant conversion through three architectural decisions:
- No mesh reconstruction – The dual-grid encodes surface geometry via pre-computed dual vertices, eliminating the encoder's need to rasterize or remesh the object.
- Sparse representation – Only active voxels are processed, drastically reducing memory and computational load compared to dense voxel grids.
- Quantized features – uint8 storage minimizes GPU memory bandwidth, allowing the encoder to read geometry with minimal cache overhead.
Practical Implementation
Converting Meshes to Dual-Grid Format
The following pattern—derived from data_toolkit/dual_grid.py—voxelizes a mesh and persists it as .vxz:
import torch, pickle, numpy as np, o_voxel
from pathlib import Path
def convert_and_save(mesh_pickle_path, out_dir, sha256, resolution=256):
# Load mesh dump
with open(mesh_pickle_path, 'rb') as f:
dump = pickle.load(f)
# Assemble vertices & faces
verts, faces = [], []
start = 0
for obj in dump['objects']:
if obj['vertices'].size == 0 or obj['faces'].size == 0:
continue
verts.append(obj['vertices'])
faces.append(obj['faces'] + start)
start += obj['vertices'].shape[0]
vertices = torch.from_numpy(np.concatenate(verts, axis=0)).float()
faces = torch.from_numpy(np.concatenate(faces, axis=0)).long()
# Voxelize to dual grid
voxel_idx, dual_vert, intersected = o_voxel.convert.mesh_to_flexible_dual_grid(
vertices, faces,
grid_size=resolution,
aabb=[[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]],
face_weight=1.0,
boundary_weight=0.2,
regularization_weight=1e-2
)
# Save as .vxz
out_path = Path(out_dir) / f'dual_grid_{resolution}' / f'{sha256}.vxz'
out_path.parent.mkdir(parents=True, exist_ok=True)
intersected_packed = (intersected[:, 0:1] +
2 * intersected[:, 1:2] +
4 * intersected[:, 2:3]).type(torch.uint8)
o_voxel.io.write_vxz(
str(out_path),
voxel_idx,
{
'vertices': (dual_vert * 255).type(torch.uint8),
'intersected': intersected_packed
}
)
Encoding Dual-Grids to Latent Vectors
Load the cached dual-grid and encode it using the Sparse Structure VAE:
import torch
from trellis2.datasets.flexi_dual_grid import FlexiDualGridDataset
from trellis2.models.sparse_structure_vae import SparseStructureEncoder
# Initialize dataset
dataset = FlexiDualGridDataset(
roots={
'mesh_dump': '/path/to/mesh_dump',
'dual_grid': '/path/to/dual_grid'
},
resolution=256,
max_active_voxels=500_000
)
# Load sample (already formatted as SparseTensor)
sample = dataset[0]
sparse_vox = sample['vertices']
# Build encoder
encoder = SparseStructureEncoder(
in_channels=4, # 3 dual-vertex + 1 intersected flag
latent_channels=256,
num_res_blocks=2,
channels=[64, 128, 256],
norm_type='layer',
use_fp16=False
)
# Instant latent extraction
latent_mean, latent_logvar = encoder(sparse_vox)
Summary
- The dual-grid representation pre-computes voxelized geometry using
mesh_to_flexible_dual_gridino_voxel/convert/flexible_dual_grid.py, solving QEFs to generate dual vertices and intersection flags. .vxzfiles store sparse, quantized grids viadata_toolkit/dual_grid.py, enabling zero-overhead reloading without re-voxelization.FlexiDualGridDatasetconverts stored grids intoSparseTensorobjects thatSparseStructureEncoderprocesses in a single forward pass to produce latent codes.- uint8 quantization and sparse structure minimize GPU bandwidth and memory, delivering true instant mesh-to-latent conversion.
Frequently Asked Questions
What makes the dual-grid representation "instant" compared to traditional voxelization?
Traditional pipelines rasterize meshes on-the-fly during encoding, which requires significant compute per forward pass. The TRELLIS.2 dual-grid pre-computes voxelization using the C++ backend _C.mesh_to_flexible_dual_grid_cpu, stores results as sparse .vxz files, and loads them as SparseTensor objects. This moves the expensive rasterization and QEF solving offline, leaving only a lightweight sparse convolution for the encoder.
How does the mesh_to_flexible_dual_grid function handle surface approximation?
The function solves a quadratic energy function (QEF) per voxel to compute dual vertices that best approximate the original mesh surface. The energy minimization is controlled by three weights: face_weight (surface fidelity), boundary_weight (edge handling), and regularization_weight (numerical stability). This produces quantized uint8 vertices that preserve geometric detail without requiring the original mesh topology at encoding time.
Why are the dual-vertices and intersection flags stored as uint8?
Quantization to uint8 reduces memory footprint by 75% compared to float32 coordinates, allowing the SparseStructureEncoder to read geometry data with minimal GPU cache bandwidth. The FlexiDualGridDataset normalizes these values back to [0, 1] range (dual_vert.float() / 255.0) before feeding them to the encoder, maintaining numerical precision while maximizing throughput.
Can I convert meshes to latents without saving intermediate .vxz files?
Yes, though not recommended for production training. You can call o_voxel.convert.mesh_to_flexible_dual_grid directly in your data loader to generate voxel_indices, dual_vertices, and intersected tensors on-the-fly, then construct a SparseTensor manually. However, this repeats the expensive QEF solving for every epoch, whereas the .vxz workflow amortizes this cost to a one-time preprocessing step.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →