What Causes Out-of-Memory Errors in TRELLIS.2 High-Resolution Generation
Out-of-memory errors in TRELLIS.2 high-resolution generation occur because the framework allocates intermediate tensors that scale quadratically with image resolution and cubically with voxel resolution, quickly exhausting GPU memory during rasterization and latent encoding.
When working with microsoft/TRELLIS.2, pushing the rendering resolution to 1024 px or higher often triggers CUDA out-of-memory exceptions. According to the TRELLIS.2 source code, this happens not from the model weights themselves, but from temporary buffers created during the rendering and encoding pipeline. Understanding exactly where these allocations occur allows you to configure the renderer to stay within your GPU's memory budget.
Rasterization Buffer Allocation in MeshRenderer
The primary memory bottleneck during 2D generation stems from the MeshRenderer.render method in trellis2/renderers/mesh_renderer.py. At lines 87‑108, the renderer allocates multiple full-resolution floating-point buffers using torch.zeros to store per-pixel intermediate results:
- Mask buffer:
resolution × resolutionfloats - Depth buffer: Same dimensions
- Normal maps: 3 channels × resolution²
- Attribute maps: Base color, metallic, roughness, and alpha channels
At 1024 px resolution, the mask alone consumes approximately 4 MiB (1 million floats), while the complete set of buffers can easily exceed several hundred megabytes. The PBRMeshRenderer class in trellis2/renderers/pbr_mesh_renderer.py creates additional attribute buffers, further multiplying the footprint.
The Cubic Cost of Dense Voxel Grids
For 3D generation, memory consumption follows cubic scaling in data_toolkit/encode_ss_latent.py. At lines 118‑120, the code allocates a dense 4‑D tensor:
torch.zeros(1, res, res, res)
This means:
- 256³ grid: ~64 MiB
- 512³ grid: ~512 MiB
A single 512³ voxel grid can consume half a gigabyte of GPU memory before any model weights or additional buffers are loaded. The voxelize_pbr.py script compounds this issue by looping over multiple resolutions simultaneously, stacking per-resolution buffers that accumulate memory pressure.
Supersampling and Chunked Rendering Impact
Two configuration parameters dramatically alter memory requirements:
SSAA (Supersample Anti-Aliasing): When ssaa is set greater than 1, the rasterizer internally multiplies the resolution by this factor. A 1024 px render with ssaa=2 actually allocates buffers at 2048 px, quadrupling the memory requirement.
Chunk Size: By default, chunk_size=None forces the renderer to process the entire image in one pass. Setting a specific tile size (e.g., 256 or 512) splits the workload into smaller sub-regions, keeping peak memory usage flat regardless of final output resolution.
Practical Fixes to Avoid OOM Errors
Enable Chunked Rendering
The most effective immediate fix is to specify a chunk_size in your rendering options. This forces MeshRenderer to process the image in tiles rather than allocating buffers for the full resolution:
from trellis2.renderers.mesh_renderer import MeshRenderer
renderer = MeshRenderer(
rendering_options={
"resolution": 1024,
"near": 0.1,
"far": 10.0,
"ssaa": 1,
"chunk_size": 256, # Process in 256×256 tiles
"antialias": True,
},
device="cuda"
)
output = renderer.render(mesh, extrinsics, intrinsics, ["mask", "normal", "depth"])
Reduce Resolution and SSAA
If chunked rendering is insufficient, lower the base resolution and keep supersampling disabled:
renderer = MeshRenderer(
rendering_options={
"resolution": 512, # Reduce from 1024 to 512
"ssaa": 1, # Disable supersampling
"chunk_size": 256,
},
device="cuda"
)
Optimize Voxel Encoding Resolution
When encoding sparse structure latent representations, avoid dense high-resolution grids. The encode_ss_latent.py script accepts a resolution parameter that directly controls the cubic allocation:
import argparse
from data_toolkit.encode_ss_latent import main as encode_ss_latent
opt = argparse.Namespace()
opt.resolution = 256 # Keep at 256 or lower to avoid 512³ allocations
opt.enc_pretrained = "path/to/pretrained/model"
# ... additional options ...
encode_ss_latent(opt) # Allocates (1,256,256,256) ≈ 64 MiB
Free Intermediate Tensors
After rendering completes, explicitly clear the CUDA cache if your script continues with other memory-intensive operations:
import torch
torch.cuda.empty_cache()
Summary
- Quadratic scaling:
MeshRendererallocates multipleresolution × resolutionbuffers for masks, depth, and attributes, causing memory to grow with the square of image size. - Cubic scaling:
encode_ss_latent.pycreates dense 3‑D tensors of sizeres³, making high-resolution voxel grids prohibitively expensive. - SSAA multiplier: Values greater than 1 internally increase resolution before downsampling, silently quadrupling buffer sizes.
- Chunked rendering: Setting
chunk_sizeto 256 or 512 forces tile-based processing, capping peak memory usage regardless of final output dimensions. - Key files:
trellis2/renderers/mesh_renderer.py(lines 87‑108),trellis2/renderers/pbr_mesh_renderer.py, anddata_toolkit/encode_ss_latent.py(lines 118‑120) contain the critical allocation sites.
Frequently Asked Questions
Why does TRELLIS.2 run out of memory at 1024 px but not at 512 px?
At 1024 px, the rasterization buffers in mesh_renderer.py require four times the memory of 512 px due to quadratic scaling. A full set of attribute maps (mask, depth, normal, base-color, metallic, roughness) at 1024 px can exceed 500 MiB just for intermediate storage, whereas at 512 px the same buffers use roughly 125 MiB. If your GPU has 8 GiB or less, this difference often crosses the available memory threshold.
What is the chunk_size parameter in TRELLIS.2 renderers?
The chunk_size parameter controls tile-based rendering in MeshRenderer.render. When set to an integer value (e.g., 256), the renderer splits the image into tiles of that size and processes them sequentially, reusing the same smaller buffer for each tile. When None (default), the renderer allocates buffers for the entire image at once, which causes out-of-memory errors at high resolutions.
Does supersampling affect memory usage in TRELLIS.2?
Yes. The ssaa (supersample anti-aliasing) parameter multiplies the internal rendering resolution by its value. Setting ssaa=2 on a 1024 px render actually allocates buffers at 2048 px, increasing memory requirements by 4×. To minimize memory usage, keep ssaa=1 unless you specifically need anti-aliasing for high-quality outputs.
How can I encode high-resolution 3D latents without running out of memory?
Avoid dense voxel grids in encode_ss_latent.py. The script allocates a torch.zeros(1, res, res, res) tensor that scales cubically. Use sparse voxel representations supported by TRELLIS.2 instead of dense tensors, or limit resolution to 256 or lower. At 256³, the dense tensor consumes only ~64 MiB, while 512³ requires ~512 MiB, which quickly exhausts available GPU RAM when combined with model weights.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →