How max_num_tokens Controls Output Resolution in TRELLIS.2

The max_num_tokens parameter enforces a hard ceiling on 3-D token generation during the shape-latent cascade, automatically downscaling the output resolution by 128-voxel increments whenever the sampled coordinate count exceeds the limit.

In the microsoft/TRELLIS.2 open-source repository, the max_num_tokens parameter serves as a critical safeguard within the image-to-3D pipeline. This value directly bounds the density of sampled voxel coordinates during the cascade upsampling stage, determining whether the final geometry renders at the requested resolution or a reduced, computationally cheaper alternative.

The Shape-Latent Cascade Process

The resolution control mechanism activates during the shape-latent cascade stage of generation. The pipeline first samples a low-resolution latent representation (lr_cond), then upsamples this latent to a higher target resolution (hr_resolution). Following upsampling, the system quantizes the high-resolution coordinates and counts the resulting tokens (num_tokens = coords.shape[0]).

This token count represents the discrete 3-D points that define the output geometry. Higher resolution targets naturally produce more tokens, creating a direct relationship between voxel density and memory consumption.

The max_num_tokens Resolution Cap

When the quantized token count exceeds the user-specified max_num_tokens threshold, the pipeline triggers an automatic resolution reduction loop. Implemented in the sample_shape_slat_cascade method of trellis2/pipelines/trellis2_image_to_3d.py (lines 287-390), the logic follows this sequence:

  1. Quantize high-resolution coordinates and count tokens.
  2. If num_tokens > max_num_tokens, reduce hr_resolution by 128 voxels.
  3. Re-quantize at the new lower resolution.
  4. Repeat until the token count falls below the threshold or the resolution reaches the hard floor of 1024 voxels.

If the resolution hits the 1024-voxel limit while tokens still exceed the budget, the system clamps the resolution and prints a runtime warning:

Due to the limited number of tokens, the resolution is reduced to <hr_resolution>.

The default max_num_tokens value is 49,152, balancing detail and resource usage for typical consumer hardware.

Supporting Components

Three primary modules interact to enforce this resolution-token relationship:

Practical Configuration Examples

The following examples demonstrate how to adjust max_num_tokens to control output fidelity when calling the pipeline programmatically.

Default Token Budget (49,152)

from trellis2.pipelines.trellis2_image_to_3d import Trellis2ImageTo3D

pipeline = Trellis2ImageTo3D(models=..., datasets=..., ...)

# Using default max_num_tokens=49,152

slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
)
print(f"Output resolution: {resolution}")  # May be reduced if tokens > 49,152

High-Detail Generation

Increase the token ceiling to permit finer geometric detail without resolution reduction:

slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
    max_num_tokens=100_000,  # Higher ceiling preserves full resolution

)
print(f"Output resolution with higher token budget: {resolution}")

Aggressive Downsampling

Force lower resolution outputs to reduce memory footprint and inference time:

slat, resolution = pipeline.sample_shape_slat_cascade(
    lr_cond=lr_cond,
    cond=cond,
    flow_model_lr=flow_lr,
    flow_model=flow_hr,
    lr_resolution=64,
    resolution=256,
    coords=coords,
    max_num_tokens=30_000,  # Tighter budget forces down-sampling

)
print(f"Output resolution with tighter token budget: {resolution}")

Summary

  • max_num_tokens sets a hard upper bound on the number of 3-D tokens generated during the shape-latent cascade.
  • Exceeding the limit triggers automatic resolution reduction in 128-voxel decrements until the token count complies or the resolution hits 1024 voxels.
  • The default value of 49,152 provides a balanced baseline for standard hardware.
  • Higher values preserve fine geometric detail but require more VRAM and compute; lower values produce coarser outputs faster.

Frequently Asked Questions

What happens when the token count exceeds max_num_tokens?

When the quantized coordinate count surpasses the max_num_tokens threshold, TRELLIS.2 automatically reduces the target resolution by 128 voxels and re-quantizes. This loop continues iteratively until the token count falls below the limit or the resolution reaches the hard floor of 1024 voxels.

What is the default max_num_tokens value in TRELLIS.2?

The default value is 49,152 tokens. This default was chosen to balance geometric detail with memory efficiency on typical consumer GPUs, though you should adjust this based on your available VRAM and quality requirements.

How does max_num_tokens affect generation quality and performance?

Larger values permit higher-resolution outputs with finer surface detail but increase VRAM consumption and inference latency. Smaller values force earlier resolution reduction, producing coarser geometry but enabling faster generation and lower memory usage.

Can I disable the max_num_tokens limit entirely?

No, you cannot disable the limit entirely, but you can effectively neutralize it by setting max_num_tokens to an arbitrarily high integer (e.g., 1_000_000) that exceeds any plausible token count for your target resolution. Note that doing so may cause out-of-memory errors if the hardware cannot accommodate the full resolution density.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →