What Parameter Controls max_num_tokens and Output Resolution in TRELLIS 2?
The max_num_tokens parameter in TRELLIS 2 limits the number of 3‑D tokens generated during the shape‑latent cascade, forcing the pipeline to reduce output resolution when this threshold is exceeded.
In Microsoft’s TRELLIS 2 image-to-3D pipeline, geometric detail is governed by how many discrete 3‑D tokens the model can sample. The max_num_tokens parameter acts as a hard ceiling on token generation, with direct consequences for the final mesh resolution. Understanding this control lets you balance quality against computational cost without modifying model weights.
Where max_num_tokens Is Defined and Enforced
The resolution-gating logic resides in trellis2/pipelines/trellis2_image_to_3d.py, specifically within the sample_shape_slat_cascade method (lines 287‑390). This method implements the two-stage upsampling workflow:
- Low-resolution sampling — generates an initial latent representation (
lr_cond) - High-resolution upsampling — expands to target resolution, then quantises coordinates into tokens
After upsampling, the pipeline computes num_tokens = coords.shape[0]. If this exceeds max_num_tokens, resolution reduction triggers automatically.
How Resolution Reduction Works
The enforcement algorithm follows a simple decrement loop:
# Conceptual flow from trellis2/pipelines/trellis2_image_to_3d.py
while num_tokens > max_num_tokens and hr_resolution > 1024:
hr_resolution -= 128
# re-quantise coordinates at new resolution
num_tokens = new_coords.shape[0]
When the token count finally falls below max_num_tokens (or the hard floor of 1024 voxels is reached), the pipeline proceeds with that resolution. If clamped, you’ll see:
Due to the limited number of tokens, the resolution is reduced to <hr_resolution>.
Larger max_num_tokens → finer voxel grids; smaller values → coarser, downsampled geometry.
Practical Code Examples
Default Behavior (49,152 tokens)
from trellis2.pipelines.trellis2_image_to_3d import Trellis2ImageTo3D
pipeline = Trellis2ImageTo3D(models=..., datasets=..., ...)
slat, resolution = pipeline.sample_shape_slat_cascade(
lr_cond=lr_cond,
cond=cond,
flow_model_lr=flow_lr,
flow_model=flow_hr,
lr_resolution=64,
resolution=256,
coords=coords,
)
print(f"Output resolution: {resolution}") # May be <256 if token count exceeds 49,152
Increased Token Budget for Higher Detail
slat, resolution = pipeline.sample_shape_slat_cascade(
lr_cond=lr_cond,
cond=cond,
flow_model_lr=flow_lr,
flow_model=flow_hr,
lr_resolution=64,
resolution=256,
coords=coords,
max_num_tokens=100_000, # Higher ceiling → higher possible output resolution
)
print(f"Output resolution with higher token budget: {resolution}")
Reduced Token Budget for Faster Inference
slat, resolution = pipeline.sample_shape_slat_cascade(
lr_cond=lr_cond,
cond=cond,
flow_model_lr=flow_lr,
flow_model=flow_hr,
lr_resolution=64,
resolution=256,
coords=coords,
max_num_tokens=30_000, # Fewer tokens → lower final resolution
)
print(f"Output resolution with tighter token budget: {resolution}")
Key Source Files and Their Roles
-
trellis2/pipelines/trellis2_image_to_3d.py— Containssample_shape_slat_cascade()wheremax_num_tokensis enforced and resolution adjustment occurs -
trellis2/models/sparse_structure_flow.py— Defines the flow models that generate the latent representations being tokenised -
trellis2/models/sc_vaes/sparse_unet_vae.py— Provides the decoder that renders final geometry at the selected resolution
Summary
max_num_tokenscaps the number of 3‑D coordinate tokens generated during the shape‑latent cascade in TRELLIS 2- Excess tokens trigger automatic resolution reduction in 128‑voxel decrements until the count falls below threshold or hits 1024 voxels
- Default value is 49,152; raising it permits higher-resolution outputs while lowering it reduces computational load and geometric detail
- Implementation lives in
sample_shape_slat_cascadewithintrellis2/pipelines/trellis2_image_to_3d.py
Frequently Asked Questions
What happens if I set max_num_tokens extremely high?
The pipeline will attempt to preserve your requested resolution, but you may encounter GPU memory exhaustion or out-of-memory errors before reaching theoretical limits. The 1024‑voxel floor prevents infinite resolution, not memory constraints.
Does max_num_tokens affect the low-resolution stage?
No. The parameter governs only the high-resolution upsampling phase (resolution argument). The low-resolution conditioning (lr_resolution=64 typically) proceeds unchanged regardless of token budget.
Can I disable resolution reduction entirely?
Not directly. Setting max_num_tokens to a very large value (e.g., 10_000_000) effectively disables the reduction loop for practical purposes, though the 1024‑voxel minimum remains enforced.
Why does resolution change in 128‑voxel steps specifically?
The 128‑voxel decrement matches the spatial block alignment used in the sparse convolution architecture of TRELLIS 2’s VAE decoder. This granularity ensures tensor dimensions remain compatible with the underlying sparse_unet_vae.py implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →