How the TRELLIS.2 Cascade Pipeline Upsamples Resolution (512→1024→1536) for 3D Generation
The TRELLIS.2 cascade pipeline progressively upsamples structured latents through three stages—SI2, IO24, and IS36—using a low-resolution flow model, a learned decoder, and a high-resolution flow model with token pruning to fit GPU memory limits.
Microsoft's TRELLIS.2 implements an efficient cascade pipeline for resolution upsampling that enables high-quality 3D generation at resolutions up to 1536×1536 while respecting hardware constraints. This article breaks down the SI2 → IO24 → IS36 pipeline implemented in trellis2/pipelines/trellis2_image_to_3d.py and explains how each stage transforms sparse structured latents.
Understanding the Three-Stage Cascade Architecture
The cascade resolution upsampling follows a deliberate progression where each abbreviation encodes both the operation and the dimensional scaling:
| Stage | Abbreviation Meaning | Resolution | Core Operation |
|---|---|---|---|
| SI2 | Sparse Input 2 | 256 (configurable) | Low-resolution flow sampling |
| IO24 | Intermediate Output 24 | 1024 (4× upsample) | Decoder-based spatial expansion |
| IS36 | Interpolated Structured 36 | 1024–1536 (target) | High-resolution flow refinement |
This staged approach lets TRELLIS.2 allocate computational budget efficiently: coarse structure emerges at low resolution, then detail fills in progressively.
Stage 1: SI2 (Sparse Input 2) — Low-Resolution Foundation
The cascade begins with low-resolution structured latent sampling using a dedicated lightweight flow model.
In trellis2/pipelines/trellis2_image_to_3d.py, the sample_shape_slat_cascade method first generates the SI2 latent:
noise = SparseTensor(
feats=torch.randn(..., flow_model_lr.in_channels),
coords=coords
)
slat = self.shape_slat_sampler.sample(
flow_model_lr,
noise,
**lr_cond,
**sampler_params
).samples
slat = slat * std + mean # de-normalize
The flow_model_lr (defined in trellis2/models/sparse_structure_flow.py) operates on a sparse grid at lr_resolution (typically 256). This produces a coarse but complete structural scaffold with manageable token count.
Stage 2: IO24 — Decoder-Based 4× Upsampling
The shape-splat decoder expands the SI2 latent through learned upsampling, bridging resolution gaps without dense computation.
The critical call occurs in sample_shape_slat_cascade:
hr_coords = self.models['shape_slat_decoder'].upsample(slat, upsample_times=4)
This operation:
- Expands the coordinate grid by 4× in each spatial dimension (16× in voxel volume)
- Uses the decoder's
upsamplemethod (implementation intrellis2/modules/spatial.pyor the decoder class) - Produces
hr_coordsathr_resolution(typically 1024 initially)
The "24" in IO24 reflects that this intermediate representation occupies a middle ground—higher resolution than input, not yet fully refined.
Stage 3: IS36 — Token-Pruned High-Resolution Sampling
The final stage introduces adaptive token budget management to enable resolutions up to 1536 without exceeding GPU memory.
Token Count Pruning
Before high-resolution sampling, TRELLIS.2 ensures the token count stays below max_num_tokens:
while True:
quant_coords = torch.cat([
hr_coords[:, :1],
((hr_coords[:, 1:] + 0.5) / lr_resolution * (hr_resolution // 16)).int(),
], dim=1)
coords = quant_coords.unique(dim=0)
if coords.shape[0] < max_num_tokens or hr_resolution == 1024:
break
hr_resolution -= 128
This loop progressively reduces hr_resolution (1024 → 896 → 768 → ...) until the unique coordinate count fits the budget. The quantization maps high-resolution coordinates back to a normalized grid for deduplication.
Final Flow Sampling
With pruned coordinates, the full flow model completes the cascade:
noise = SparseTensor(
feats=torch.randn(..., flow_model.in_channels),
coords=coords
)
slat = self.shape_slat_sampler.sample(
flow_model,
noise,
**cond,
**sampler_params
).samples
slat = slat * std + mean # de-normalize
The flow_model (higher capacity than flow_model_lr) refines detail at the target resolution, producing the final IS36 structured latent.
Complete Cascade Pipeline Usage
Configure TRELLIS.2 for cascade operation using the 1024_cascade or 1536_cascade pipeline types:
from trellis2.pipelines.trellis2_image_to_3d import TRELLIS2ImageTo3DPipeline
# Initialize with cascade configuration
pipeline = TRELLIS2ImageTo3DPipeline(
default_pipeline_type='1536_cascade', # or '1024_cascade'
low_vram=False,
)
# Prepare conditioning for both stages
lr_cond = {
'image_embeds': image_features,
'depth_cond': depth_low_res,
}
cond = {
'image_embeds': image_features,
'depth_cond': depth_high_res,
}
# Execute full cascade upsampling
structured_latent, effective_resolution = pipeline.sample_shape_slat_cascade(
lr_cond=lr_cond,
cond=cond,
flow_model_lr=pipeline.models['shape_slat_flow_lr'],
flow_model=pipeline.models['shape_slat_flow'],
lr_resolution=256,
resolution=1536, # Target IS36 resolution
coords=initial_sparse_coords,
sampler_params={'num_steps': 50, 'cfg_scale': 4.0},
)
print(f"Generated latent: {structured_latent.shape}")
print(f"Effective resolution: {effective_resolution}") # May be < 1536 if pruned
Key Source Files and Implementation Details
The cascade resolution upsampling spans three critical files:
trellis2/pipelines/trellis2_image_to_3d.py— Containssample_shape_slat_cascadeorchestrating the full SI2→IO24→IS36 pipelinetrellis2/models/sparse_structure_flow.py— DefinesSparseStructureFlowModelused for bothflow_model_lrandflow_modeltrellis2/modules/spatial.py— Houses decoder upsampling operations that expand coordinate grids
According to the TRELLIS.2 source code, the cascade design deliberately separates concerns: flow_model_lr handles coarse structure efficiently, the decoder learns spatial interpolation patterns, and flow_model allocates capacity to fine detail where it matters most.
Why This Architecture Matters
The cascade pipeline for resolution upsampling solves a fundamental tension in 3D generation:
- Quality requires high spatial resolution (1024+)
- Feasibility requires bounded token counts (memory limits)
By decomposing generation into SI2→IO24→IS36, TRELLIS.2 achieves both: the initial flow model operates sparsely at low cost, the decoder densifies representation through learned upsampling, and token pruning ensures the final stage remains tractable. The 1536_cascade configuration pushes this to 2.25× the pixels of the base 1024 mode while maintaining reasonable inference times.
Summary
- SI2 stage: Low-resolution flow sampling at 256×256 using
flow_model_lrestablishes coarse structure efficiently - IO24 stage: Decoder
upsamplemethod expands coordinates 4× (16× voxel space) to intermediate resolution - IS36 stage: Token pruning loop adjusts
hr_resolutiondownward untilmax_num_tokensconstraint satisfied - Resolution adaptation: Final effective resolution may be reduced from target (1536→1408→1280...) based on hardware limits
- Pipeline types:
1024_cascadeand1536_cascadeactivate this multi-stage generation path in TRELLIS.2
Frequently Asked Questions
What do SI2, IO24, and IS36 stand for in TRELLIS.2?
SI2 means "Sparse Input 2" (low-resolution latent sampling), IO24 means "Intermediate Output 24" (decoder upsampling stage), and IS36 means "Interpolated Structured 36" (final high-resolution sampling). The numbers encode dimensional relationships: 2 refers to the initial sparse grid, 24 marks the 4× expanded intermediate, and 36 indicates the target structured resolution.
Why does the cascade pipeline use two separate flow models?
TRELLIS.2 uses flow_model_lr for coarse structure and flow_model for detail refinement because their computational profiles differ. The low-resolution model can afford more sampling steps or parameters per token, while the high-resolution model focuses capacity on the pruned coordinate set where detail matters. This separation prevents wasting computation on empty or uniform spatial regions.
How does token pruning affect final output quality?
The while-loop that reduces hr_resolution by 128 until coordinates fit max_num_tokens acts as a quality safeguard rather than a degradation. According to the source in trellis2_image_to_3d.py, pruning threshold triggers at 1024 minimum, so 1024-cascade runs typically use full resolution. For 1536-cascade, the 1408 or 1280 effective resolution still exceeds standard 1024 outputs while respecting GPU memory constraints.
Can I disable token pruning for maximum quality?
The sample_shape_slat_cascade method hardcodes the pruning logic with hr_resolution == 1024 as the floor condition. To override, you would need to modify the source in trellis2/pipelines/trellis2_image_to_3d.py—though this risks out-of-memory errors on standard hardware. The low_vram=False pipeline flag already optimizes memory usage to minimize pruning frequency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →