# How max_num_tokens Controls Output Resolution in TRELLIS.2

> Learn how the max_num_tokens parameter in TRELLIS.2 controls output resolution. Understand its impact on 3-D token generation and automatic downscaling.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: deep-dive
- Published: 2026-08-03

---

**The `max_num_tokens` parameter enforces a hard ceiling on 3-D token generation during the shape-latent cascade, automatically downscaling the output resolution by 128-voxel increments whenever the sampled coordinate count exceeds the limit.**

In the **microsoft/TRELLIS.2** open-source repository, the `max_num_tokens` parameter serves as a critical safeguard within the image-to-3D pipeline. This value directly bounds the density of sampled voxel coordinates during the cascade upsampling stage, determining whether the final geometry renders at the requested resolution or a reduced, computationally cheaper alternative.

## The Shape-Latent Cascade Process

The resolution control mechanism activates during the **shape-latent cascade** stage of generation. The pipeline first samples a low-resolution latent representation (`lr_cond`), then upsamples this latent to a higher target resolution (`hr_resolution`). Following upsampling, the system quantizes the high-resolution coordinates and counts the resulting tokens (`num_tokens = coords.shape[0]`).

This token count represents the discrete 3-D points that define the output geometry. Higher resolution targets naturally produce more tokens, creating a direct relationship between voxel density and memory consumption.

## The max_num_tokens Resolution Cap

When the quantized token count exceeds the user-specified `max_num_tokens` threshold, the pipeline triggers an automatic resolution reduction loop. Implemented in the **`sample_shape_slat_cascade`** method of [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) (lines 287-390), the logic follows this sequence:

1. Quantize high-resolution coordinates and count tokens.
2. If `num_tokens > max_num_tokens`, reduce `hr_resolution` by **128 voxels**.
3. Re-quantize at the new lower resolution.
4. Repeat until the token count falls below the threshold **or** the resolution reaches the hard floor of **1024 voxels**.

If the resolution hits the 1024-voxel limit while tokens still exceed the budget, the system clamps the resolution and prints a runtime warning:

```text
Due to the limited number of tokens, the resolution is reduced to <hr_resolution>.

```

The default `max_num_tokens` value is **49,152**, balancing detail and resource usage for typical consumer hardware.

## Supporting Components

Three primary modules interact to enforce this resolution-token relationship:

- **[`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py)**: Contains the `sample_shape_slat_cascade` method that implements the reduction loop and threshold enforcement.
- **[`trellis2/models/sparse_structure_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_flow.py)**: Defines the flow model responsible for latent sampling that generates the initial token distributions.
- **[`trellis2/models/sc_vaes/sparse_unet_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sc_vaes/sparse_unet_vae.py)**: Provides the decoder that ultimately renders geometry at the final (potentially reduced) resolution.

## Practical Configuration Examples

The following examples demonstrate how to adjust `max_num_tokens` to control output fidelity when calling the pipeline programmatically.

### Default Token Budget (49,152)

```python
from trellis2.pipelines.trellis2_image_to_3d import Trellis2ImageTo3D

pipeline = Trellis2ImageTo3D(models=..., datasets=..., ...)

# Using default max_num_tokens=49,152

slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
)
print(f"Output resolution: {resolution}")  # May be reduced if tokens > 49,152

```

### High-Detail Generation

Increase the token ceiling to permit finer geometric detail without resolution reduction:

```python
slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
    max_num_tokens=100_000,  # Higher ceiling preserves full resolution

)
print(f"Output resolution with higher token budget: {resolution}")

```

### Aggressive Downsampling

Force lower resolution outputs to reduce memory footprint and inference time:

```python
slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
    max_num_tokens=30_000,  # Tighter budget forces down-sampling

)
print(f"Output resolution with tighter token budget: {resolution}")

```

## Summary

- **`max_num_tokens`** sets a hard upper bound on the number of 3-D tokens generated during the shape-latent cascade.
- Exceeding the limit triggers automatic resolution reduction in **128-voxel decrements** until the token count complies or the resolution hits **1024 voxels**.
- The default value of **49,152** provides a balanced baseline for standard hardware.
- Higher values preserve fine geometric detail but require more VRAM and compute; lower values produce coarser outputs faster.

## Frequently Asked Questions

### What happens when the token count exceeds max_num_tokens?

When the quantized coordinate count surpasses the `max_num_tokens` threshold, TRELLIS.2 automatically reduces the target resolution by 128 voxels and re-quantizes. This loop continues iteratively until the token count falls below the limit or the resolution reaches the hard floor of 1024 voxels.

### What is the default max_num_tokens value in TRELLIS.2?

The default value is **49,152** tokens. This default was chosen to balance geometric detail with memory efficiency on typical consumer GPUs, though you should adjust this based on your available VRAM and quality requirements.

### How does max_num_tokens affect generation quality and performance?

**Larger values** permit higher-resolution outputs with finer surface detail but increase VRAM consumption and inference latency. **Smaller values** force earlier resolution reduction, producing coarser geometry but enabling faster generation and lower memory usage.

### Can I disable the max_num_tokens limit entirely?

No, you cannot disable the limit entirely, but you can effectively neutralize it by setting `max_num_tokens` to an arbitrarily high integer (e.g., `1_000_000`) that exceeds any plausible token count for your target resolution. Note that doing so may cause out-of-memory errors if the hardware cannot accommodate the full resolution density.