# How O-Voxel Sparse Representation Works in TRELLIS.2 for 3D Generation

> Discover how O-Voxel sparse representation in TRELLIS.2 enables compact latents for 3D generation. Learn about its field-free format, dual-grid geometry, and PBR material attributes for large-scale diffusion models.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: deep-dive
- Published: 2026-08-04

---

**TRELLIS.2 uses a field-free sparse voxel format called O-Voxel that stores only occupied voxels together with dual-grid geometry and PBR material attributes, enabling compact latents for large-scale diffusion models while supporting arbitrary topology and interior cavities.**

TRELLIS.2 is Microsoft's open-source framework for high-quality 3D asset generation that replaces traditional signed-distance fields with an efficient sparse representation. The **O-Voxel sparse representation** captures geometry and physically-based rendering (PBR) materials in a compact latent space that bypasses the memory bottlenecks of dense volumetric grids. This architecture enables sub-second inference on complex meshes with full material support through specialized CUDA kernels and sparse convolution operations.

## What is O-Voxel Sparse Representation?

Unlike conventional neural radiance fields or signed-distance functions (SDFs) that require dense grid sampling, **O-Voxel** is a field-free format that explicitly stores only the occupied voxels of a mesh. According to the TRELLIS.2 source code, each occupied cell contains **dual vertices** calculated via Quadratic Error Function (QEF) minimization, which represent the geometric center of the cell's interior. The representation couples this sparse geometry with an **attribute volume** carrying per-voxel PBR channels—including base color, metallic, roughness, and opacity—creating a complete material-aware latent that can be processed by diffusion models.

## Encoding Meshes into O-Voxel Format

### Voxelization with Quadratic Error Functions

The conversion from polygon mesh to O-Voxel begins in [`o-voxel/o_voxel/convert/flexible_dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o-voxel/o_voxel/convert/flexible_dual_grid.py). The function `mesh_to_flexible_dual_grid` voxelizes the input mesh into a regular grid and solves a **Quadratic Error Function (QEF)** for each occupied cell to obtain dual vertices.

If the voxel size is omitted, the grid resolution is automatically inferred from the provided axis-aligned bounding box (AABB). This produces three core tensors:

- `coords`: Integer indices of occupied voxels
- `dual_vertices`: Float coordinates of the dual grid vertices
- `intersected`: Boolean flags indicating valid cells

### Sparse Storage Architecture

Rather than allocating dense memory, TRELLIS.2 stores only the non-empty cells. The **attribute volume** (`attr_volume`) is allocated on the same sparse grid as the coordinates, with channels for **base-color**, **metallic**, **roughness**, and **opacity**. This sparse structure—comprising coordinates, dual vertices, and attributes—is encoded by a **Structured-Latent VAE (SC-VAE)** into a compressed latent representation. Because empty space requires zero storage, the model can handle resolutions of 1024³ or 1536³ while consuming only a few hundred megabytes.

## Processing Sparse Voxels in the Diffusion Pipeline

### CUDA-Accelerated Attribute Rasterization

The diffusion flow model operates directly on sparse voxel tensors through hardware-accelerated rendering. In [`o-voxel/o_voxel/rasterize.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o-voxel/o_voxel/rasterize.py), the `VoxelRenderer.render` method projects sparse voxels into camera space using the CUDA kernel `_C.rasterize_voxels_cuda`. This produces color, depth, and alpha images that serve as **2-D conditioning signals** for the diffusion model, allowing the generation process to respect view-dependent material properties.

### Sparse Convolutions with FlexGEMM

Heavy computation is handled by the **FlexGEMM** library, which provides efficient sparse-convolution kernels that operate exclusively on active voxels. According to the TRELLIS.2 implementation, this keeps computational cost proportional to the number of occupied cells (O(Nₒ)) rather than the full grid volume (O(N³)), enabling real-time processing of high-resolution assets.

## Decoding O-Voxel Back to Production Assets

### Dual-Grid to Mesh Reconstruction

Reconstruction from the sparse representation uses `flexible_dual_grid_to_mesh` in [`flexible_dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/flexible_dual_grid.py). The algorithm builds a **hash-map** of occupied voxels to rapidly locate neighboring cells, then reconstructs quad faces that are split into triangles either via learned split weights or by minimizing the **dihedral angle** to ensure watertight surfaces. This approach naturally supports open surfaces, non-manifold geometry, and internal cavities that traditional implicit fields struggle to represent.

### GLB Export and Texture Baking

The final export stage occurs in [`o-voxel/o_voxel/postprocess.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o-voxel/o_voxel/postprocess.py) via the `to_glb` function. This pipeline cleans and optionally remeshes the geometry, unwraps UV coordinates, and **bakes the attribute volume** onto a 2-D texture using a differentiable rasterizer (`nvdiffrast`). The resulting GLB file encodes full PBR materials—base color with alpha, metallic, and roughness channels—that render correctly in standard 3D viewers.

## O-Voxel vs Traditional SDF: Performance Comparison

| Property | Traditional SDF | O-Voxel (TRELLIS.2) |
|----------|----------------|---------------------|
| **Memory** | Dense grid ⇒ O(N³) | Stores only occupied voxels ⇒ O(Nₒ) |
| **Topology** | Implicit iso-surface struggles with open/non-manifold geometry | Directly represents open surfaces, internal cavities, and non-manifold structures |
| **Detail** | Limited by voxel resolution; metal-roughness must be encoded into the field | Per-voxel attribute channels give high-fidelity PBR textures; decoupled geometry and material |
| **Speed** | Full-grid convolutions are costly | Sparse convolutions (FlexGEMM) and CUDA rasterisation give sub-second inference |

## Implementation Examples

### Converting a Mesh to O-Voxel

```python
import torch
from o_voxel.convert.flexible_dual_grid import mesh_to_flexible_dual_grid

# verts (N,3) and faces (M,3) are torch tensors of the input mesh

coords, dual_vertices, intersected = mesh_to_flexible_dual_grid(
    vertices=verts,
    faces=faces,
    voxel_size=0.01,               # voxel edge length in world units

    aabb=[[-0.5, -0.5, -0.5],
          [ 0.5,  0.5,  0.5]],
)

```

### Rendering Sparse Voxels for Diffusion Conditioning

```python
from o_voxel.rasterize import VoxelRenderer

# Initialize renderer with camera parameters

renderer = VoxelRenderer({"resolution": 512, "near": 0.1, "far": 5.0})

# attr_volume shape: (L, 4) containing base, metallic, roughness, alpha

rendered = renderer.render(
    position=coords.float(),
    attrs=attr_volume,
    voxel_size=0.01,
    extrinsics=cam_extr,
    intrinsics=cam_int,
)

# rendered.attr provides a (3, H, W) image for the diffusion pipeline

```

### Reconstructing Mesh and GLB Export

```python
from o_voxel.convert.flexible_dual_grid import flexible_dual_grid_to_mesh
from o_voxel.postprocess import to_glb

# Rebuild triangle mesh from sparse representation

mesh_verts, mesh_faces = flexible_dual_grid_to_mesh(
    coords=coords,
    dual_vertices=dual_vertices,
    intersected_flag=intersected,
    split_weight=None,               # automatic split selection

    aabb=[[-0.5, -0.5, -0.5],
          [ 0.5,  0.5,  0.5]],
    voxel_size=0.01,
)

# Export to GLB with baked PBR textures

glb = to_glb(
    vertices=mesh_verts,
    faces=mesh_faces,
    attr_volume=attr_volume,
    coords=coords,
    attr_layout=dict(
        base_color=slice(0, 1),
        metallic=slice(1, 2),
        roughness=slice(2, 3),
        alpha=slice(3, 4),
    ),
    aabb=[[-0.5, -0.5, -0.5],
          [ 0.5,  0.5,  0.5]],
    voxel_size=0.01,
    texture_size=4096,
    remesh=False,
)

glb.export("generated_asset.glb")

```

## Summary

- **O-Voxel** replaces dense SDF grids with a field-free sparse format that stores only occupied voxels and their dual-grid geometry.
- The representation encodes full **PBR material attributes** (base color, metallic, roughness, opacity) per voxel, enabling high-fidelity texture generation.
- **FlexGEMM** sparse convolutions and CUDA rasterization in [`rasterize.py`](https://github.com/microsoft/TRELLIS.2/blob/main/rasterize.py) ensure computational efficiency scales with surface area rather than volume.
- Conversion functions in [`flexible_dual_grid.py`](https://github.com/microsoft/TRELLIS.2/blob/main/flexible_dual_grid.py) handle both mesh-to-voxel encoding and watertight reconstruction, supporting arbitrary topology including non-manifold surfaces.
- The `to_glb` pipeline bakes sparse attributes into standard GLB files with UV-mapped textures for immediate use in production rendering.

## Frequently Asked Questions

### What makes O-Voxel "field-free" compared to SDF representations?

Traditional SDFs define surfaces implicitly as the zero-level set of a continuous distance field, requiring dense grid evaluation and struggling with open boundaries. **O-Voxel** explicitly stores only occupied voxels with their dual vertices and surface attributes, eliminating the need for implicit field evaluation and enabling direct representation of open surfaces and internal structures without contorted field formulations.

### How does O-Voxel handle non-manifold geometry and open surfaces?

The `flexible_dual_grid_to_mesh` function constructs a hash-map of occupied voxels to identify neighbors without assuming manifold connectivity. Because the representation does not rely on implicit iso-surfaces, it can represent **non-manifold edges**, **open surfaces**, and **internal cavities** naturally—geometries that would cause artifacts or failures in traditional SDF-based reconstruction methods.

### What are the memory requirements for O-Voxel at high resolutions like 1024³?

Because O-Voxel stores only occupied cells, memory consumption scales as **O(Nₒ)** (number of occupied voxels) rather than **O(N³)**. According to the TRELLIS.2 implementation, even at 1024³ or 1536³ resolutions, the sparse representation compresses to a few hundred megabytes when encoded through the SC-VAE, whereas a dense float32 grid at 1024³ would require approximately 4 GB for a single channel.

### Can O-Voxel represent full PBR material properties, not just geometry?

Yes. The **attribute volume** in O-Voxel stores per-voxel channels for **base color**, **metallic**, **roughness**, and **opacity** (alpha). During export, the `to_glb` function uses `nvdiffrast` to bake these attributes into standard 2-D PBR textures, producing GLB files that contain complete material definitions compatible with modern rendering engines and WebGL viewers.