# How the TRELLIS.2 Cascade Pipeline Upsamples Resolution (512→1024→1536) for 3D Generation

> Discover how the TRELLIS.2 cascade pipeline upsamples resolution from 512 to 1536 for 3D generation. Learn about its three stages SI2, IO24, and IS36 and memory-efficient techniques.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: deep-dive
- Published: 2026-08-04

---

**The TRELLIS.2 cascade pipeline progressively upsamples structured latents through three stages—SI2, IO24, and IS36—using a low-resolution flow model, a learned decoder, and a high-resolution flow model with token pruning to fit GPU memory limits.**

Microsoft's TRELLIS.2 implements an efficient **cascade pipeline for resolution upsampling** that enables high-quality 3D generation at resolutions up to 1536×1536 while respecting hardware constraints. This article breaks down the SI2 → IO24 → IS36 pipeline implemented in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) and explains how each stage transforms sparse structured latents.

## Understanding the Three-Stage Cascade Architecture

The **cascade resolution upsampling** follows a deliberate progression where each abbreviation encodes both the operation and the dimensional scaling:

| Stage | Abbreviation Meaning | Resolution | Core Operation |
|-------|---------------------|------------|----------------|
| **SI2** | Sparse Input 2 | 256 (configurable) | Low-resolution flow sampling |
| **IO24** | Intermediate Output 24 | 1024 (4× upsample) | Decoder-based spatial expansion |
| **IS36** | Interpolated Structured 36 | 1024–1536 (target) | High-resolution flow refinement |

This staged approach lets TRELLIS.2 allocate computational budget efficiently: coarse structure emerges at low resolution, then detail fills in progressively.

## Stage 1: SI2 (Sparse Input 2) — Low-Resolution Foundation

The cascade begins with **low-resolution structured latent sampling** using a dedicated lightweight flow model.

In [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), the `sample_shape_slat_cascade` method first generates the SI2 latent:

```python
noise = SparseTensor(
    feats=torch.randn(..., flow_model_lr.in_channels),
    coords=coords
)
slat = self.shape_slat_sampler.sample(
    flow_model_lr, 
    noise, 
    **lr_cond, 
    **sampler_params
).samples
slat = slat * std + mean  # de-normalize

```

The `flow_model_lr` (defined in [`trellis2/models/sparse_structure_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_flow.py)) operates on a sparse grid at `lr_resolution` (typically 256). This produces a coarse but complete structural scaffold with manageable token count.

## Stage 2: IO24 — Decoder-Based 4× Upsampling

The **shape-splat decoder** expands the SI2 latent through learned upsampling, bridging resolution gaps without dense computation.

The critical call occurs in `sample_shape_slat_cascade`:

```python
hr_coords = self.models['shape_slat_decoder'].upsample(slat, upsample_times=4)

```

This operation:
- Expands the coordinate grid by **4× in each spatial dimension** (16× in voxel volume)
- Uses the decoder's `upsample` method (implementation in [`trellis2/modules/spatial.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/modules/spatial.py) or the decoder class)
- Produces `hr_coords` at `hr_resolution` (typically 1024 initially)

The "24" in IO24 reflects that this intermediate representation occupies a middle ground—higher resolution than input, not yet fully refined.

## Stage 3: IS36 — Token-Pruned High-Resolution Sampling

The final stage introduces **adaptive token budget management** to enable resolutions up to 1536 without exceeding GPU memory.

### Token Count Pruning

Before high-resolution sampling, TRELLIS.2 ensures the token count stays below `max_num_tokens`:

```python
while True:
    quant_coords = torch.cat([
        hr_coords[:, :1],
        ((hr_coords[:, 1:] + 0.5) / lr_resolution * (hr_resolution // 16)).int(),
    ], dim=1)
    coords = quant_coords.unique(dim=0)
    
    if coords.shape[0] < max_num_tokens or hr_resolution == 1024:
        break
    hr_resolution -= 128

```

This loop progressively reduces `hr_resolution` (1024 → 896 → 768 → ...) until the unique coordinate count fits the budget. The quantization maps high-resolution coordinates back to a normalized grid for deduplication.

### Final Flow Sampling

With pruned coordinates, the full flow model completes the cascade:

```python
noise = SparseTensor(
    feats=torch.randn(..., flow_model.in_channels),
    coords=coords
)
slat = self.shape_slat_sampler.sample(
    flow_model,
    noise,
    **cond,
    **sampler_params
).samples
slat = slat * std + mean  # de-normalize

```

The `flow_model` (higher capacity than `flow_model_lr`) refines detail at the target resolution, producing the final **IS36** structured latent.

## Complete Cascade Pipeline Usage

Configure TRELLIS.2 for cascade operation using the `1024_cascade` or `1536_cascade` pipeline types:

```python
from trellis2.pipelines.trellis2_image_to_3d import TRELLIS2ImageTo3DPipeline

# Initialize with cascade configuration

pipeline = TRELLIS2ImageTo3DPipeline(
    default_pipeline_type='1536_cascade',  # or '1024_cascade'

    low_vram=False,
)

# Prepare conditioning for both stages

lr_cond = {
    'image_embeds': image_features,
    'depth_cond': depth_low_res,
}
cond = {
    'image_embeds': image_features,
    'depth_cond': depth_high_res,
}

# Execute full cascade upsampling

structured_latent, effective_resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=pipeline.models['shape_slat_flow_lr'],
    flow_model=pipeline.models['shape_slat_flow'],
    lr_resolution=256,
    resolution=1536,  # Target IS36 resolution

    coords=initial_sparse_coords,
    sampler_params={'num_steps': 50, 'cfg_scale': 4.0},
)

print(f"Generated latent: {structured_latent.shape}")
print(f"Effective resolution: {effective_resolution}")  # May be < 1536 if pruned

```

## Key Source Files and Implementation Details

The **cascade resolution upsampling** spans three critical files:

- **[`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py)** — Contains `sample_shape_slat_cascade` orchestrating the full SI2→IO24→IS36 pipeline
- **[`trellis2/models/sparse_structure_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_flow.py)** — Defines `SparseStructureFlowModel` used for both `flow_model_lr` and `flow_model`
- **[`trellis2/modules/spatial.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/modules/spatial.py)** — Houses decoder upsampling operations that expand coordinate grids

According to the TRELLIS.2 source code, the cascade design deliberately separates concerns: `flow_model_lr` handles coarse structure efficiently, the decoder learns spatial interpolation patterns, and `flow_model` allocates capacity to fine detail where it matters most.

## Why This Architecture Matters

The **cascade pipeline for resolution upsampling** solves a fundamental tension in 3D generation:

- **Quality** requires high spatial resolution (1024+)
- **Feasibility** requires bounded token counts (memory limits)

By decomposing generation into SI2→IO24→IS36, TRELLIS.2 achieves both: the initial flow model operates sparsely at low cost, the decoder densifies representation through learned upsampling, and token pruning ensures the final stage remains tractable. The `1536_cascade` configuration pushes this to 2.25× the pixels of the base 1024 mode while maintaining reasonable inference times.

## Summary

- **SI2 stage**: Low-resolution flow sampling at 256×256 using `flow_model_lr` establishes coarse structure efficiently
- **IO24 stage**: Decoder `upsample` method expands coordinates 4× (16× voxel space) to intermediate resolution
- **IS36 stage**: Token pruning loop adjusts `hr_resolution` downward until `max_num_tokens` constraint satisfied
- **Resolution adaptation**: Final effective resolution may be reduced from target (1536→1408→1280...) based on hardware limits
- **Pipeline types**: `1024_cascade` and `1536_cascade` activate this multi-stage generation path in TRELLIS.2

## Frequently Asked Questions

### What do SI2, IO24, and IS36 stand for in TRELLIS.2?

**SI2** means "Sparse Input 2" (low-resolution latent sampling), **IO24** means "Intermediate Output 24" (decoder upsampling stage), and **IS36** means "Interpolated Structured 36" (final high-resolution sampling). The numbers encode dimensional relationships: 2 refers to the initial sparse grid, 24 marks the 4× expanded intermediate, and 36 indicates the target structured resolution.

### Why does the cascade pipeline use two separate flow models?

TRELLIS.2 uses `flow_model_lr` for coarse structure and `flow_model` for detail refinement because their computational profiles differ. The low-resolution model can afford more sampling steps or parameters per token, while the high-resolution model focuses capacity on the pruned coordinate set where detail matters. This separation prevents wasting computation on empty or uniform spatial regions.

### How does token pruning affect final output quality?

The while-loop that reduces `hr_resolution` by 128 until coordinates fit `max_num_tokens` acts as a quality safeguard rather than a degradation. According to the source in [`trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2_image_to_3d.py), pruning threshold triggers at 1024 minimum, so 1024-cascade runs typically use full resolution. For 1536-cascade, the 1408 or 1280 effective resolution still exceeds standard 1024 outputs while respecting GPU memory constraints.

### Can I disable token pruning for maximum quality?

The `sample_shape_slat_cascade` method hardcodes the pruning logic with `hr_resolution == 1024` as the floor condition. To override, you would need to modify the source in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py)—though this risks out-of-memory errors on standard hardware. The `low_vram=False` pipeline flag already optimizes memory usage to minimize pruning frequency.