TRELLIS.2 Multi-Stage Cascade Pipeline Architecture: From 512³ to 1536³ Resolution
TRELLIS.2 generates high-fidelity 3D assets from a single image using a three-stage cascade pipeline—SI2 (Sparse-Structure Inference), IO24 (Intermediate Object Shape), and IS36 (Intermediate Surface Texture)—that progressively scales voxel resolution from 512³ through 1024³ to 1536³ while maintaining computational efficiency.
Microsoft’s TRELLIS.2 repository implements a sophisticated multi-stage cascade pipeline architecture that transforms a single input image into detailed 3D meshes with physically based rendering (PBR) materials. The system employs a coarse-to-fine approach, internally abbreviated as SI2 → IO24 → IS36, to efficiently scale from sparse occupancy grids up to high-resolution voxel representations without exhausting GPU memory budgets.
The Three-Stage Cascade Architecture (SI2 → IO24 → IS36)
The pipeline processes an input image through three distinct stages, each responsible for a specific level of geometric and textural detail.
Stage 1: Sparse-Structure Inference (SI2)
The SI2 stage establishes the foundational geometry by predicting a coarse, field-free O-Voxel occupancy grid. In trellis2/pipelines/trellis2_image_to_3d.py (lines 88-105), the image conditioning (cond_512) feeds into the sparse_structure_flow_model, which transforms random noise into a latent z_s. The sparse_structure_decoder then decodes this into a binary voxel mask, extracting the coordinates of occupied voxels that define the object’s overall shape.
Stage 2: Intermediate Object Shape (IO24)
The IO24 stage generates a structured latent (SLat) encoding detailed shape information through a two-step cascade:
- Low-resolution pass (512³): The pipeline uses
shape_slat_flow_model_512with the initial conditioning to produce a low-resolution shape SLat. - High-resolution refinement (1024³): The
shape_slat_decoder.upsamplemethod up-samples the low-res SLat, which is then refined byshape_slat_flow_model_1024using high-resolution conditioning (cond_1024) and the up-sampled coordinates.
This cascade, implemented in lines 77-89 of the main pipeline file, allows the model to generate high-fidelity geometry while keeping the initial token count tractable.
Stage 3: Intermediate Surface Texture (IS36)
The final IS36 stage produces a texture-latent (SLat) carrying per-voxel PBR attributes including base color, metallic, roughness, and alpha. The high-resolution shape SLat is normalized and concatenated with cond_1024, then processed by tex_slat_flow_model_1024 and tex_slat_decoder (lines 91-99). This stage scales the representation to 1536³ resolution, outputting voxel-wise material attributes that are decoded and combined with the mesh geometry into a MeshWithVoxel object ready for rendering.
How the Resolution Cascade Works (512 → 1024 → 1536)
The 512 → 1024 → 1536 progression operates as a memory-efficient scaling mechanism. The pipeline begins at 512³ resolution to establish coarse geometry, up-samples to 1024³ for detailed shape refinement, and finally generates texture attributes at 1536³. When the token budget (max_num_tokens) would be exceeded at higher resolutions, the pipeline automatically falls back to lower resolutions (steps 35-40 in trellis2_image_to_3d.py), ensuring robust execution across varying hardware constraints.
Key Implementation Files
The cascade architecture is implemented across several critical modules:
trellis2/pipelines/trellis2_image_to_3d.py: Orchestrates the full SI2 → IO24 → IS36 cascade, including sampling, up-sampling, and decoding logic.trellis2/modules/sparse.py: Defines theSparseTensorclass used throughout all three stages for efficient voxel representation.trellis2/modules/transformer/: Contains the transformer-based flow models (shape_slat_flow_model_512,shape_slat_flow_model_1024,tex_slat_flow_model_1024) that power each stage.o-voxel/: Core library handling the O-Voxel representation and mesh-to-voxel conversions used by the SI2 stage.
Practical Usage Examples
Basic 1024-Cascade Inference
import torch
from PIL import Image
from trellis2.pipelines import Trellis2ImageTo3DPipeline
from trellis2.renderers import EnvMap
import o_voxel
# Load the pretrained cascade pipeline (SI2 → IO24 → IS36)
pipeline = Trellis2ImageTo3DPipeline.from_pretrained("microsoft/TRELLIS.2-4B")
pipeline.cuda()
# Run inference
prompt = Image.open("input.png")
mesh = pipeline.run(prompt)[0] # Returns MeshWithVoxel
# Export to GLB with PBR materials
glb = o_voxel.postprocess.to_glb(
vertices=mesh.vertices,
faces=mesh.faces,
attr_volume=mesh.attrs,
coords=mesh.coords,
attr_layout=mesh.layout,
voxel_size=mesh.voxel_size,
aabb=[[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]],
decimation_target=1_000_000,
texture_size=4096
)
glb.export("output.glb")
Running the Full 1536-Cascade
# Configure pipeline for maximum resolution
pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
"microsoft/TRELLIS.2-4B",
config_file="pipeline.json"
)
pipeline.default_pipeline_type = "1536_cascade" # SI2 → IO24 → IS36 at 1536³
mesh = pipeline.run(prompt)[0]
Inspecting Intermediate Latents
# Return intermediate SLat representations for debugging
out_mesh, (shape_slat, tex_slat, resolution) = pipeline.run(
prompt,
return_latent=True,
pipeline_type="1024_cascade"
)
print(f"Shape SLat shape: {shape_slat.feats.shape}") # IO24 output
print(f"Texture SLat shape: {tex_slat.feats.shape}") # IS36 output
print(f"Final resolution: {resolution}")
Summary
- TRELLIS.2 uses a three-stage cascade (SI2 → IO24 → IS36) to generate 3D assets from single images without overwhelming GPU memory.
- SI2 establishes coarse geometry using O-Voxel occupancy grids at the base resolution.
- IO24 performs a two-step shape refinement, processing 512³ then 1024³ resolutions via structured latents (SLat) with up-sampling between stages.
- IS36 generates high-resolution 1536³ PBR texture attributes conditioned on the final shape SLat.
- The architecture automatically manages token budgets by adjusting resolution when necessary, as implemented in the main pipeline class.
Frequently Asked Questions
What does the 512 → 1024 → 1536 progression represent in TRELLIS.2?
The progression represents the voxel resolution scaling across the cascade stages. 512³ is the initial low-resolution shape generation in IO24, 1024³ is the refined high-resolution shape, and 1536³ is the final texture resolution in IS36. This staged approach allows the model to handle high-resolution outputs efficiently by processing coarse geometry before expensive high-resolution texturing.
How does TRELLIS.2 handle memory limitations during the cascade?
The pipeline monitors the max_num_tokens parameter during inference. If the voxel count at the target resolution would exceed this budget, the system automatically lowers the resolution (as seen in steps 35-40 of trellis2_image_to_3d.py). This ensures the multi-stage cascade remains robust across different GPU memory configurations without manual intervention.
What is the difference between the O-Voxel representation and the structured latent (SLat)?
O-Voxel is the sparse occupancy grid used in the SI2 stage to represent coarse binary geometry (occupied vs. empty space). SLat (Structured Latent) is a continuous, high-dimensional feature representation used in IO24 and IS36 that encodes fine geometric details and material properties. While O-Voxels define where geometry exists, SLats encode what the geometry looks like and how it is textured.
Can I export intermediate outputs from specific cascade stages?
Yes. By setting return_latent=True in the pipeline.run() method, you can extract the intermediate shape SLat (from IO24) and texture SLat (from IS36) tensors. This allows researchers to inspect the latent representations or manipulate them before final mesh decoding, facilitating debugging and custom post-processing workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →