# TRELLIS.2 Multi-Stage Cascade Pipeline Architecture: From 512³ to 1536³ Resolution

> Explore TRELLIS.2's multi-stage cascade pipeline architecture. Generate high-fidelity 3D assets from single images, scaling resolution from 512³ to 1536³ efficiently.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: architecture
- Published: 2026-08-03

---

**TRELLIS.2 generates high-fidelity 3D assets from a single image using a three-stage cascade pipeline—SI2 (Sparse-Structure Inference), IO24 (Intermediate Object Shape), and IS36 (Intermediate Surface Texture)—that progressively scales voxel resolution from 512³ through 1024³ to 1536³ while maintaining computational efficiency.**

Microsoft’s TRELLIS.2 repository implements a sophisticated **multi-stage cascade pipeline architecture** that transforms a single input image into detailed 3D meshes with physically based rendering (PBR) materials. The system employs a coarse-to-fine approach, internally abbreviated as **SI2 → IO24 → IS36**, to efficiently scale from sparse occupancy grids up to high-resolution voxel representations without exhausting GPU memory budgets.

## The Three-Stage Cascade Architecture (SI2 → IO24 → IS36)

The pipeline processes an input image through three distinct stages, each responsible for a specific level of geometric and textural detail.

### Stage 1: Sparse-Structure Inference (SI2)

The **SI2** stage establishes the foundational geometry by predicting a coarse, field-free **O-Voxel** occupancy grid. In [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) (lines 88-105), the image conditioning (`cond_512`) feeds into the `sparse_structure_flow_model`, which transforms random noise into a latent `z_s`. The `sparse_structure_decoder` then decodes this into a binary voxel mask, extracting the coordinates of occupied voxels that define the object’s overall shape.

### Stage 2: Intermediate Object Shape (IO24)

The **IO24** stage generates a **structured latent (SLat)** encoding detailed shape information through a two-step cascade:

- **Low-resolution pass (512³):** The pipeline uses `shape_slat_flow_model_512` with the initial conditioning to produce a low-resolution shape SLat.
- **High-resolution refinement (1024³):** The `shape_slat_decoder.upsample` method up-samples the low-res SLat, which is then refined by `shape_slat_flow_model_1024` using high-resolution conditioning (`cond_1024`) and the up-sampled coordinates.

This cascade, implemented in lines 77-89 of the main pipeline file, allows the model to generate high-fidelity geometry while keeping the initial token count tractable.

### Stage 3: Intermediate Surface Texture (IS36)

The final **IS36** stage produces a **texture-latent (SLat)** carrying per-voxel PBR attributes including base color, metallic, roughness, and alpha. The high-resolution shape SLat is normalized and concatenated with `cond_1024`, then processed by `tex_slat_flow_model_1024` and `tex_slat_decoder` (lines 91-99). This stage scales the representation to **1536³ resolution**, outputting voxel-wise material attributes that are decoded and combined with the mesh geometry into a `MeshWithVoxel` object ready for rendering.

## How the Resolution Cascade Works (512 → 1024 → 1536)

The **512 → 1024 → 1536 progression** operates as a memory-efficient scaling mechanism. The pipeline begins at 512³ resolution to establish coarse geometry, up-samples to 1024³ for detailed shape refinement, and finally generates texture attributes at 1536³. When the token budget (`max_num_tokens`) would be exceeded at higher resolutions, the pipeline automatically falls back to lower resolutions (steps 35-40 in [`trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2_image_to_3d.py)), ensuring robust execution across varying hardware constraints.

## Key Implementation Files

The cascade architecture is implemented across several critical modules:

- **[`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py)**: Orchestrates the full SI2 → IO24 → IS36 cascade, including sampling, up-sampling, and decoding logic.
- **[`trellis2/modules/sparse.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/modules/sparse.py)**: Defines the `SparseTensor` class used throughout all three stages for efficient voxel representation.
- **`trellis2/modules/transformer/`**: Contains the transformer-based flow models (`shape_slat_flow_model_512`, `shape_slat_flow_model_1024`, `tex_slat_flow_model_1024`) that power each stage.
- **`o-voxel/`**: Core library handling the **O-Voxel** representation and mesh-to-voxel conversions used by the SI2 stage.

## Practical Usage Examples

### Basic 1024-Cascade Inference

```python
import torch
from PIL import Image
from trellis2.pipelines import Trellis2ImageTo3DPipeline
from trellis2.renderers import EnvMap
import o_voxel

# Load the pretrained cascade pipeline (SI2 → IO24 → IS36)

pipeline = Trellis2ImageTo3DPipeline.from_pretrained("microsoft/TRELLIS.2-4B")
pipeline.cuda()

# Run inference

prompt = Image.open("input.png")
mesh = pipeline.run(prompt)[0]  # Returns MeshWithVoxel

# Export to GLB with PBR materials

glb = o_voxel.postprocess.to_glb(
    vertices=mesh.vertices,
    faces=mesh.faces,
    attr_volume=mesh.attrs,
    coords=mesh.coords,
    attr_layout=mesh.layout,
    voxel_size=mesh.voxel_size,
    aabb=[[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]],
    decimation_target=1_000_000,
    texture_size=4096
)
glb.export("output.glb")

```

### Running the Full 1536-Cascade

```python

# Configure pipeline for maximum resolution

pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    "microsoft/TRELLIS.2-4B",
    config_file="pipeline.json"
)
pipeline.default_pipeline_type = "1536_cascade"  # SI2 → IO24 → IS36 at 1536³

mesh = pipeline.run(prompt)[0]

```

### Inspecting Intermediate Latents

```python

# Return intermediate SLat representations for debugging

out_mesh, (shape_slat, tex_slat, resolution) = pipeline.run(
    prompt, 
    return_latent=True, 
    pipeline_type="1024_cascade"
)

print(f"Shape SLat shape: {shape_slat.feats.shape}")   # IO24 output

print(f"Texture SLat shape: {tex_slat.feats.shape}")  # IS36 output

print(f"Final resolution: {resolution}")

```

## Summary

- **TRELLIS.2** uses a **three-stage cascade** (SI2 → IO24 → IS36) to generate 3D assets from single images without overwhelming GPU memory.
- **SI2** establishes coarse geometry using **O-Voxel** occupancy grids at the base resolution.
- **IO24** performs a two-step shape refinement, processing **512³** then **1024³** resolutions via structured latents (SLat) with up-sampling between stages.
- **IS36** generates high-resolution **1536³** PBR texture attributes conditioned on the final shape SLat.
- The architecture automatically manages token budgets by adjusting resolution when necessary, as implemented in the main pipeline class.

## Frequently Asked Questions

### What does the 512 → 1024 → 1536 progression represent in TRELLIS.2?

The progression represents the voxel resolution scaling across the cascade stages. **512³** is the initial low-resolution shape generation in IO24, **1024³** is the refined high-resolution shape, and **1536³** is the final texture resolution in IS36. This staged approach allows the model to handle high-resolution outputs efficiently by processing coarse geometry before expensive high-resolution texturing.

### How does TRELLIS.2 handle memory limitations during the cascade?

The pipeline monitors the `max_num_tokens` parameter during inference. If the voxel count at the target resolution would exceed this budget, the system automatically lowers the resolution (as seen in steps 35-40 of [`trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2_image_to_3d.py)). This ensures the **multi-stage cascade** remains robust across different GPU memory configurations without manual intervention.

### What is the difference between the O-Voxel representation and the structured latent (SLat)?

**O-Voxel** is the sparse occupancy grid used in the SI2 stage to represent coarse binary geometry (occupied vs. empty space). **SLat (Structured Latent)** is a continuous, high-dimensional feature representation used in IO24 and IS36 that encodes fine geometric details and material properties. While O-Voxels define where geometry exists, SLats encode what the geometry looks like and how it is textured.

### Can I export intermediate outputs from specific cascade stages?

Yes. By setting `return_latent=True` in the `pipeline.run()` method, you can extract the intermediate **shape SLat** (from IO24) and **texture SLat** (from IS36) tensors. This allows researchers to inspect the latent representations or manipulate them before final mesh decoding, facilitating debugging and custom post-processing workflows.