# What Parameter Controls max_num_tokens and Output Resolution in TRELLIS 2?

> Discover how the max_num_tokens parameter in TRELLIS 2 controls output resolution by limiting 3-D tokens generated in the shape-latent cascade.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: how-to-guide
- Published: 2026-08-04

---

**The `max_num_tokens` parameter in TRELLIS 2 limits the number of 3‑D tokens generated during the shape‑latent cascade, forcing the pipeline to reduce output resolution when this threshold is exceeded.**

In Microsoft’s **TRELLIS 2** image-to-3D pipeline, geometric detail is governed by how many discrete 3‑D tokens the model can sample. The `max_num_tokens` parameter acts as a hard ceiling on token generation, with direct consequences for the final mesh resolution. Understanding this control lets you balance quality against computational cost without modifying model weights.

## Where max_num_tokens Is Defined and Enforced

The resolution-gating logic resides in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), specifically within the **`sample_shape_slat_cascade`** method (lines 287‑390). This method implements the two-stage upsampling workflow:

1. **Low-resolution sampling** — generates an initial latent representation (`lr_cond`)
2. **High-resolution upsampling** — expands to target resolution, then quantises coordinates into tokens

After upsampling, the pipeline computes `num_tokens = coords.shape[0]`. If this exceeds `max_num_tokens`, resolution reduction triggers automatically.

## How Resolution Reduction Works

The enforcement algorithm follows a simple decrement loop:

```python

# Conceptual flow from trellis2/pipelines/trellis2_image_to_3d.py

while num_tokens > max_num_tokens and hr_resolution > 1024:
    hr_resolution -= 128
    # re-quantise coordinates at new resolution

    num_tokens = new_coords.shape[0]

```

When the token count finally falls below `max_num_tokens` (or the hard floor of **1024 voxels** is reached), the pipeline proceeds with that resolution. If clamped, you’ll see:

```text
Due to the limited number of tokens, the resolution is reduced to <hr_resolution>.

```

Larger `max_num_tokens` → finer voxel grids; smaller values → coarser, downsampled geometry.

## Practical Code Examples

### Default Behavior (49,152 tokens)

```python
from trellis2.pipelines.trellis2_image_to_3d import Trellis2ImageTo3D

pipeline = Trellis2ImageTo3D(models=..., datasets=..., ...)

slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
)
print(f"Output resolution: {resolution}")  # May be <256 if token count exceeds 49,152

```

### Increased Token Budget for Higher Detail

```python
slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
    max_num_tokens=100_000,  # Higher ceiling → higher possible output resolution

)
print(f"Output resolution with higher token budget: {resolution}")

```

### Reduced Token Budget for Faster Inference

```python
slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
    max_num_tokens=30_000,  # Fewer tokens → lower final resolution

)
print(f"Output resolution with tighter token budget: {resolution}")

```

## Key Source Files and Their Roles

- **[`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py)** — Contains `sample_shape_slat_cascade()` where `max_num_tokens` is enforced and resolution adjustment occurs

- **[`trellis2/models/sparse_structure_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_flow.py)** — Defines the flow models that generate the latent representations being tokenised

- **[`trellis2/models/sc_vaes/sparse_unet_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sc_vaes/sparse_unet_vae.py)** — Provides the decoder that renders final geometry at the selected resolution

## Summary

- **`max_num_tokens`** caps the number of 3‑D coordinate tokens generated during the shape‑latent cascade in TRELLIS 2
- **Excess tokens trigger automatic resolution reduction** in 128‑voxel decrements until the count falls below threshold or hits 1024 voxels
- **Default value is 49,152**; raising it permits higher-resolution outputs while lowering it reduces computational load and geometric detail
- **Implementation lives in `sample_shape_slat_cascade`** within [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py)

## Frequently Asked Questions

### What happens if I set max_num_tokens extremely high?

The pipeline will attempt to preserve your requested resolution, but you may encounter **GPU memory exhaustion** or **out-of-memory errors** before reaching theoretical limits. The 1024‑voxel floor prevents infinite resolution, not memory constraints.

### Does max_num_tokens affect the low-resolution stage?

No. The parameter governs only the **high-resolution upsampling phase** (`resolution` argument). The low-resolution conditioning (`lr_resolution=64` typically) proceeds unchanged regardless of token budget.

### Can I disable resolution reduction entirely?

Not directly. Setting `max_num_tokens` to a very large value (e.g., `10_000_000`) effectively disables the reduction loop for practical purposes, though the 1024‑voxel minimum remains enforced.

### Why does resolution change in 128‑voxel steps specifically?

The **128‑voxel decrement** matches the spatial block alignment used in the sparse convolution architecture of TRELLIS 2’s VAE decoder. This granularity ensures tensor dimensions remain compatible with the underlying [`sparse_unet_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/sparse_unet_vae.py) implementation.