# How TRELLIS.2 Resolution Scaling Works in the Cascade Pipeline

> Discover how TRELLIS.2 resolution scaling creates high-quality 3D outputs. Learn about its cascade pipeline from 32x32 latents to 1024x1024 pixels.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: deep-dive
- Published: 2026-08-03

---

**TRELLIS.2 uses a coarse-to-fine cascade pipeline that first generates a 32×32 shape latent at base resolution, then upsamples it to 64×64 or 96×96 to produce final 3D outputs at 1024×1024 or 1536×1536 pixels.**

The microsoft/TRELLIS.2 repository implements an advanced image-to-3D generation system that employs resolution scaling through a specialized cascade architecture. Unlike single-stage pipelines that process high-resolution inputs directly, TRELLIS.2 uses a progressive refinement strategy to generate detailed meshes while maintaining computational efficiency. Understanding how this cascade resolution scaling works is essential for optimizing output quality and managing GPU memory.

## Pipeline Type Selection and Cascade Activation

The cascade behavior is controlled by the `pipeline_type` parameter in `Trellis2ImageTo3DPipeline`. According to the source code in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) at line 514, the system accepts four distinct pipeline modes: `'512'`, `'1024'`, `'1024_cascade'`, and `'1536_cascade'`.

When the pipeline type string ends with `_cascade`, the code enters a specialized branch at lines 525-529. This branch activates the two-stage generation process that separates coarse shape prediction from high-resolution detail refinement.

### Supported Resolution Modes

- **512**: Standard single-stage pipeline rendering at 512×512 resolution
- **1024**: Single-stage pipeline rendering at 1024×1024 resolution  
- **1024_cascade**: Two-stage cascade producing 1024×1024 output via upsampling
- **1536_cascade**: Two-stage cascade producing 1536×1536 output via upsampling

## The Two-Stage Cascade Architecture

The cascade pipeline implements a **coarse-to-fine strategy** that decouples global shape estimation from fine detail generation. This architecture centers on the `sample_shape_slat_cascade` method and the resolution handling logic in the main pipeline class.

### Stage 1: Coarse SLAT Generation at 32×32

Regardless of the final output resolution, the cascade pipeline always begins by generating a coarse **Shape-and-Texture Latent (SLAT)** at 32×32 resolution. At line 541 of [`trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2_image_to_3d.py), a dictionary maps pipeline types to their base SLAT resolutions, with cascade modes consistently using **32** as the initial resolution.

The method `sample_shape_slat_cascade` at line 277 generates this coarse tensor by processing the input image through the shape VAE. This 32×32 latent encodes the global geometry and topology of the object without committing to fine surface details, allowing the model to establish accurate overall proportions while minimizing memory consumption.

### Stage 2: Upsampling to Target Resolution

After obtaining the coarse SLAT, the pipeline upsamples it to match the target output resolution. The upsampling logic at lines 566-579 transforms the 32×32 latent into either:

- **64×64** for 1024×1024 pixel outputs
- **96×96** for 1536×1536 pixel outputs

This upsampled SLAT is then combined with higher-resolution texture latents from the texture VAE and processed through the 3D decoder to produce the final high-detail mesh. The decoder can focus exclusively on adding fine geometric details because the coarse structure is already established.

## Implementation Details in the Source Code

The resolution scaling mechanism is implemented across several key methods in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py).

To initialize a cascade pipeline:

```python
from trellis2.pipelines.trellis2_image_to_3d import Trellis2ImageTo3DPipeline

# 1024-pixel cascade pipeline

pipeline = Trellis2ImageTo3DPipeline(
    default_pipeline_type='1024_cascade',
)

# 1536-pixel cascade pipeline  

pipeline = Trellis2ImageTo3DPipeline(
    default_pipeline_type='1536_cascade',
)

```

Running the pipeline processes the input through both cascade stages:

```python

# Process image through coarse-to-fine stages

output = pipeline.run(image_path='input.png')

# Output mesh is rasterized at the target resolution (1024² or 1536²)

```

The `sample_shape_slat_cascade` method handles both the initial sampling at 32×32 and the subsequent refinement at the upsampled resolution, returning a latent grid that the decoder transforms into the final high-resolution mesh.

## Why Cascade Improves Detail and Efficiency

The cascade approach provides three distinct advantages over single-stage high-resolution generation:

- **Coarse-to-Fine Strategy**: By establishing global shape at 32×32 resolution first, the model avoids noise accumulation in fine details and ensures topological correctness before surface refinement begins.

- **Progressive Refinement**: Upsampling the coarse SLAT and fusing it with high-resolution texture latents allows the decoder to allocate computational resources specifically to detail generation rather than global structure.

- **Memory Efficiency**: Processing the coarse SLAT at 32×32 consumes significantly less GPU memory than operating directly at 64×64 or 96×96 resolutions throughout the entire pipeline, enabling high-resolution outputs on consumer hardware.

## Summary

- TRELLIS.2 cascade resolution scaling uses a two-stage pipeline activated by `'1024_cascade'` or `'1536_cascade'` pipeline types in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py).
- The system first generates a coarse 32×32 SLAT using `sample_shape_slat_cascade` at line 277, regardless of the final output size.
- Coarse latents are upsampled to 64×64 or 96×96 at lines 566-579 to match the target output resolution of 1024×1024 or 1536×1536 pixels.
- This architecture balances computational efficiency with high-resolution output quality by separating global shape estimation from fine detail generation.

## Frequently Asked Questions

### What resolutions does TRELLIS.2 cascade support?

The cascade pipeline supports two high-resolution output modes: 1024×1024 pixels (activated via `'1024_cascade'`) and 1536×1536 pixels (activated via `'1536_cascade'`). Both modes use the same 32×32 coarse SLAT resolution but upsample to different final latent sizes (64×64 and 96×96 respectively) before decoding.

### How does the cascade pipeline differ from single-stage generation?

Single-stage pipelines process the input at the target resolution throughout the entire forward pass, while the cascade pipeline first establishes geometry at 32×32 resolution before upsampling. This coarse-to-fine approach reduces memory usage by approximately 75% during the initial shape estimation phase and improves geometric accuracy compared to direct high-resolution generation.

### Where is the upsampling logic implemented?

The upsampling logic resides in the `sample_shape_slat_cascade` method within [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), specifically around lines 566-579. This code handles the transformation from the coarse 32×32 SLAT to the target resolution latent grid, combining the upsampled shape information with high-resolution texture features.

### Can I use the cascade pipeline with custom input sizes?

The cascade pipeline is designed to process input images at the base 512×512 resolution regardless of the output resolution setting. The resolution scaling applies to the internal SLAT representations and final mesh output, not the input image size, which remains fixed at 512×512 for the coarse generation stage defined at line 541.