How max_num_tokens Controls Output Resolution in TRELLIS.2
The max_num_tokens parameter enforces a hard ceiling on 3-D token generation during the shape-latent cascade, automatically downscaling the output resolution by 128-voxel increments whenever the sampled coordinate count exceeds the limit.
In the microsoft/TRELLIS.2 open-source repository, the max_num_tokens parameter serves as a critical safeguard within the image-to-3D pipeline. This value directly bounds the density of sampled voxel coordinates during the cascade upsampling stage, determining whether the final geometry renders at the requested resolution or a reduced, computationally cheaper alternative.
The Shape-Latent Cascade Process
The resolution control mechanism activates during the shape-latent cascade stage of generation. The pipeline first samples a low-resolution latent representation (lr_cond), then upsamples this latent to a higher target resolution (hr_resolution). Following upsampling, the system quantizes the high-resolution coordinates and counts the resulting tokens (num_tokens = coords.shape[0]).
This token count represents the discrete 3-D points that define the output geometry. Higher resolution targets naturally produce more tokens, creating a direct relationship between voxel density and memory consumption.
The max_num_tokens Resolution Cap
When the quantized token count exceeds the user-specified max_num_tokens threshold, the pipeline triggers an automatic resolution reduction loop. Implemented in the sample_shape_slat_cascade method of trellis2/pipelines/trellis2_image_to_3d.py (lines 287-390), the logic follows this sequence:
- Quantize high-resolution coordinates and count tokens.
- If
num_tokens > max_num_tokens, reducehr_resolutionby 128 voxels. - Re-quantize at the new lower resolution.
- Repeat until the token count falls below the threshold or the resolution reaches the hard floor of 1024 voxels.
If the resolution hits the 1024-voxel limit while tokens still exceed the budget, the system clamps the resolution and prints a runtime warning:
Due to the limited number of tokens, the resolution is reduced to <hr_resolution>.
The default max_num_tokens value is 49,152, balancing detail and resource usage for typical consumer hardware.
Supporting Components
Three primary modules interact to enforce this resolution-token relationship:
trellis2/pipelines/trellis2_image_to_3d.py: Contains thesample_shape_slat_cascademethod that implements the reduction loop and threshold enforcement.trellis2/models/sparse_structure_flow.py: Defines the flow model responsible for latent sampling that generates the initial token distributions.trellis2/models/sc_vaes/sparse_unet_vae.py: Provides the decoder that ultimately renders geometry at the final (potentially reduced) resolution.
Practical Configuration Examples
The following examples demonstrate how to adjust max_num_tokens to control output fidelity when calling the pipeline programmatically.
Default Token Budget (49,152)
from trellis2.pipelines.trellis2_image_to_3d import Trellis2ImageTo3D
pipeline = Trellis2ImageTo3D(models=..., datasets=..., ...)
# Using default max_num_tokens=49,152
slat, resolution = pipeline.sample_shape_slat_cascade(
lr_cond=lr_cond,
cond=cond,
flow_model_lr=flow_lr,
flow_model=flow_hr,
lr_resolution=64,
resolution=256,
coords=coords,
)
print(f"Output resolution: {resolution}") # May be reduced if tokens > 49,152
High-Detail Generation
Increase the token ceiling to permit finer geometric detail without resolution reduction:
slat, resolution = pipeline.sample_shape_slat_cascade(
lr_cond=lr_cond,
cond=cond,
flow_model_lr=flow_lr,
flow_model=flow_hr,
lr_resolution=64,
resolution=256,
coords=coords,
max_num_tokens=100_000, # Higher ceiling preserves full resolution
)
print(f"Output resolution with higher token budget: {resolution}")
Aggressive Downsampling
Force lower resolution outputs to reduce memory footprint and inference time:
slat, resolution = pipeline.sample_shape_slat_cascade(
lr_cond=lr_cond,
cond=cond,
flow_model_lr=flow_lr,
flow_model=flow_hr,
lr_resolution=64,
resolution=256,
coords=coords,
max_num_tokens=30_000, # Tighter budget forces down-sampling
)
print(f"Output resolution with tighter token budget: {resolution}")
Summary
max_num_tokenssets a hard upper bound on the number of 3-D tokens generated during the shape-latent cascade.- Exceeding the limit triggers automatic resolution reduction in 128-voxel decrements until the token count complies or the resolution hits 1024 voxels.
- The default value of 49,152 provides a balanced baseline for standard hardware.
- Higher values preserve fine geometric detail but require more VRAM and compute; lower values produce coarser outputs faster.
Frequently Asked Questions
What happens when the token count exceeds max_num_tokens?
When the quantized coordinate count surpasses the max_num_tokens threshold, TRELLIS.2 automatically reduces the target resolution by 128 voxels and re-quantizes. This loop continues iteratively until the token count falls below the limit or the resolution reaches the hard floor of 1024 voxels.
What is the default max_num_tokens value in TRELLIS.2?
The default value is 49,152 tokens. This default was chosen to balance geometric detail with memory efficiency on typical consumer GPUs, though you should adjust this based on your available VRAM and quality requirements.
How does max_num_tokens affect generation quality and performance?
Larger values permit higher-resolution outputs with finer surface detail but increase VRAM consumption and inference latency. Smaller values force earlier resolution reduction, producing coarser geometry but enabling faster generation and lower memory usage.
Can I disable the max_num_tokens limit entirely?
No, you cannot disable the limit entirely, but you can effectively neutralize it by setting max_num_tokens to an arbitrarily high integer (e.g., 1_000_000) that exceeds any plausible token count for your target resolution. Note that doing so may cause out-of-memory errors if the hardware cannot accommodate the full resolution density.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →