How to Debug Generation Failures and Tune Sampler Parameters in TRELLIS 2
Most generation failures in TRELLIS 2 result from invalid sampler_params ranges, stale ResumableSampler state, or shape mismatches between models and conditioning tensors.
TRELLIS 2 generates 3D assets by sampling latent spaces through a hierarchy of sampler objects defined in the microsoft/TRELLIS.2 repository. When a generation run crashes, produces NaN values, or yields empty meshes, the root cause almost always traces back to one of three areas: sampler configuration, data loader state management, or model-sampler compatibility.
Verify Sampler Parameter Ranges
All pipelines expose a *_sampler_params dictionary that is merged with per-call overrides before sampling. In trellis2/pipelines/trellis2_texturing.py (lines 246–254), the texture latent sampler parameters are merged as follows:
sampler_params = {**self.tex_slat_sampler_params, **sampler_params}
slat = self.tex_slat_sampler.sample(model, **sampler_params)
Key parameters control the diffusion process:
- sigma_min: Minimum noise level. Safe range is 0.01 to 0.1; values near 0 cause NaN tensors.
- sigma_max: Maximum noise level. Must be greater than sigma_min, typically 10 to 100; excessive values produce noisy geometry.
- num_steps: Number of diffusion steps. Range 10 to 100 balances quality and speed; values above 200 slow generation significantly.
- guidance_scale: Strength of classifier-free guidance. Keep between 1.0 and 5.0; values above 5.0 often create artifacts.
If generation fails with NaN values, lower sigma_max or increase num_steps. For overly blurred outputs, raise sigma_min or reduce guidance_scale.
Reset Data Sampler State Between Inference Runs
During training, TRELLIS 2 uses ResumableSampler or BalancedResumableSampler to manage dataset epochs. When reusing a model for inference, failing to reset the sampler causes it to start at a non-zero index, resulting in empty batches or repeated latent vectors.
In trellis2/utils/data_utils.py (lines 53–60), the sampler state updates after each epoch:
if isinstance(data_loader.sampler, ResumableSampler):
data_loader.sampler.idx += data_loader.batch_size
data_loader.sampler.epoch += 1
data_loader.sampler.idx = 0
Before inference, manually reset these fields or instantiate a fresh sampler:
if hasattr(data_loader.sampler, 'reset'):
data_loader.sampler.reset()
else:
data_loader.sampler.idx = 0
data_loader.sampler.epoch = 0
Ensure Model-Sampler Compatibility
All samplers derive from the abstract base class in trellis2/pipelines/samplers/base.py:
class Sampler(ABC):
@abstractmethod
def sample(self, model, **kwargs):
"""Sample from a model."""
pass
Concrete implementations like FlowEulerSampler and FlowEulerCfgSampler (defined in trellis2/pipelines/samplers/flow_euler.py) expect the model to accept specific noise schedules and optional conditioning tensors. When swapping between VAE models—such as replacing a standard VAE with a sparse-structure variant—verify that the sampler invocation matches the model's expected signature (e.g., model(z, **cond)).
Shape mismatches manifest as RuntimeError: Expected ... but got ... exceptions. Validate compatibility by running a single forward pass with dummy inputs before the full sampling loop.
Common Generation Failure Modes and Fixes
| Failure | Likely Cause | Quick Fix |
|---|---|---|
| CUDA out of memory | num_steps multiplied by batch size exceeds GPU memory. |
Reduce num_steps or batch size; enable gradient checkpointing in trellis2/utils/grad_clip_utils.py. |
| All zeros / empty meshes | ResumableSampler.idx not reset; sampler starts at end of dataset. |
Call sampler.reset() or manually set idx and epoch to 0. |
| Extremely noisy geometry | sigma_max exceeds 100 or guidance_scale above 5.0. |
Lower sigma_max to 30–50 range; reduce guidance_scale to 2.0–3.0. |
| NaNs in latent tensor | sigma_min set to 0 or below 0.01. |
Set sigma_min ≥ 0.01. |
| Slow generation | num_steps set to 200 or higher. |
Decrease to 30–50 steps while monitoring visual quality. |
Step-by-Step Debugging Workflow
Follow this systematic checklist to isolate generation failures:
-
Print final sampler parameters before the sampling call:
print("Merged sampler params:", sampler_params) -
Execute a single diffusion step to isolate where NaNs or crashes occur. Most samplers expose an internal
stepmethod for debugging. -
Inspect data loader indices after each inference call:
print(f"Sampler state: idx={data_loader.sampler.idx}, epoch={data_loader.sampler.epoch}") -
Validate model output shapes using dummy inputs matching the expected conditioning dimensions.
-
Enable verbose logging by setting
verbose=Truein the sampler constructor to trace the noise schedule through each timestep.
Practical Example – Configuring FlowEulerSampler
The following example demonstrates conservative parameter tuning for the texturing pipeline using FlowEulerSampler:
from trellis2.pipelines import samplers
# Instantiate with stable defaults
tex_sampler = samplers.FlowEulerSampler(
sigma_min=0.05,
sigma_max=30.0,
num_steps=30,
guidance_scale=2.0,
verbose=True
)
# Merge with user overrides
user_params = {"num_steps": 20, "guidance_scale": 1.5}
sampler_params = {**tex_sampler.__dict__, **user_params}
# Generate texture latent
texture_latent = tex_sampler.sample(model, **sampler_params)
This configuration minimizes NaN risks while maintaining generation quality. For production use, wrap the sampling call in a try-except block to catch RuntimeError exceptions and log the final parameter state for post-mortem analysis.
Summary
- Validate ranges: Keep
sigma_min≥ 0.01,sigma_maxbetween 10–100, andguidance_scalebelow 5.0 to avoid numerical instabilities. - Reset state: Always reset
ResumableSamplerindices or create fresh samplers between inference runs to prevent empty batch errors. - Check compatibility: Ensure model signatures match the abstract
Sampler.sample()interface defined intrellis2/pipelines/samplers/base.py. - Diagnose systematically: Print parameters, run single steps, and enable verbose logging to isolate failure points.
Frequently Asked Questions
Why does my TRELLIS 2 generation produce NaN values in the latent tensor?
NaN values typically indicate that sigma_min is set to 0 or an extremely small value below 0.01, causing numerical instability in the Euler ODE solver. According to the FlowEulerSampler implementation in trellis2/pipelines/samplers/flow_euler.py, you should set sigma_min to at least 0.01 and ensure sigma_max does not exceed 100.
How do I fix empty mesh outputs when running inference multiple times?
Empty outputs occur when the ResumableSampler state is not reset between runs, causing the sampler to start at the end of the dataset. Reset the sampler by setting data_loader.sampler.idx = 0 and data_loader.sampler.epoch = 0 as shown in trellis2/utils/data_utils.py, or instantiate a new sampler object before each generation run.
What causes CUDA out-of-memory errors during high-step generation?
The product of num_steps and batch size determines memory consumption. Reduce num_steps from 100 to 30–50, or enable gradient checkpointing via the utilities in trellis2/utils/grad_clip_utils.py to trade computation for memory efficiency.
Why is my generated geometry extremely noisy despite correct prompts?
Excessive noise usually results from sigma_max values above 100 or guidance_scale settings above 5.0. Lower sigma_max to the 30–50 range and reduce guidance_scale to 2.0–3.0 according to the parameter safe ranges defined in the pipeline configuration files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →