TRELLIS.2 GPU Memory Requirements by Output Resolution: Complete Hardware Guide
TRELLIS.2 requires a minimum of 24 GB GPU VRAM for 512³ output, approximately 48 GB for 1024³, and roughly 80 GB for 1536³, with memory usage scaling cubically due to dense voxel attribute storage.
TRELLIS.2, Microsoft's open-source 3D generation model, produces high-quality textured meshes through its novel O-Voxel representation. Understanding GPU memory requirements by resolution is critical for deployment planning, as the per-voxel geometry and material attributes create substantial VRAM demands that scale with the cube of linear resolution.
Supported Resolutions and Memory Requirements
TRELLIS.2 offers three discrete output resolutions, each with distinct hardware demands. The repository's README.md documents these specifications explicitly.
512³ Resolution: Baseline Configuration
Minimum GPU Memory: ≥ 24 GB
This is the entry-level resolution supported by TRELLIS.2. Inference completes in approximately 3 seconds total—2 seconds for initial generation plus 1 second for post-processing. The 24 GB requirement represents the absolute hardware floor for running the model.
According to the Hardware section of the repository, this baseline ensures sufficient headroom for the O-Voxel decoder and attribute buffers.
1024³ Resolution: Mid-Range Production
Approximate GPU Memory: ≈ 48 GB
Stepping up to 1024³ increases inference time to roughly 17 seconds (10 s + 7 s). While the voxel count increases 8× from 512³ to 1024³, shared attribute tables and streaming optimizations reduce practical memory growth to approximately 2×.
This configuration fits comfortably on NVIDIA A100 (40 GB) and H100 (80 GB) GPUs. Users with 40 GB cards may need memory optimization techniques (see below).
1536³ Resolution: Maximum Quality
Approximate GPU Memory: ≈ 80 GB
The highest-resolution mode demands the full capacity of an H100-80GB GPU, with inference taking approximately 60 seconds (35 s + 25 s). The authors explicitly verified this configuration on H100 hardware, indicating that 1536³ consumes nearly all available VRAM on current flagship GPUs.
Why GPU Memory Scales Cubically
TRELLIS.2's O-Voxel representation stores dense per-voxel attributes including:
- Base color (RGB)
- Roughness
- Metallic
- Opacity
- Spatial coordinates
The total voxel count grows as N³ with linear resolution. While the sparse-voxel VAE encoder performs 16× downsampling, the decoder expands latent representations back to full target resolution during inference. This requires the GPU to maintain the complete dense representation in memory.
The postprocess.py module implements chunked rasterization to limit peak memory during mesh export, but the primary inference path still holds full-resolution buffers.
Hardware Verification and Testing
The TRELLIS.2 authors validated their memory specifications on NVIDIA H100 GPUs with 80 GB VRAM. This testing environment confirms:
- 512³ and 1024³ modes run reliably within available memory
- 1536³ operates at the hardware limit with minimal headroom
Users attempting 1536³ generation on lower-capacity GPUs (A100-40GB or consumer RTX cards) will encounter out-of-memory errors during the decode phase.
Memory Optimization for Constrained GPUs
For GPUs below recommended specifications, TRELLIS.2 supports flexible memory allocation through PyTorch configuration.
Environment Variable Configuration
Set PYTORCH_CUDA_ALLOC_CONF before importing torch:
import os
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
import torch
from trellis2.pipelines import Trellis2ImageTo3DPipeline
# Initialize pipeline at reduced resolution
pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
"microsoft/TRELLIS.2-4B",
resolution="512" # Use lowest resolution for 24 GB GPUs
)
pipeline.cuda()
This technique, shown in example.py, enables more flexible memory block allocation and may allow 1024³ generation on 40 GB GPUs that would otherwise fail.
Resolution Selection in Pipeline
The resolution parameter directly controls voxel grid allocation in trellis2_image_to_3d.py:
from PIL import Image
from trellis2.pipelines import Trellis2ImageTo3DPipeline
# Available: "512", "1024", "1536"
RESOLUTION = "1024"
pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
"microsoft/TRELLIS.2-4B",
resolution=RESOLUTION
)
pipeline.cuda()
img = Image.open("input.png")
mesh = pipeline.run(img)[0]
Practical Deployment Recommendations
| Use Case | Recommended Resolution | Minimum GPU | Expected Performance |
|---|---|---|---|
| Rapid prototyping, limited VRAM | 512³ | RTX 3090 (24 GB) | ~3 seconds |
| Production quality, mid-range servers | 1024³ | A100-40GB or above | ~17 seconds |
| Maximum fidelity, research | 1536³ | H100-80GB exclusively | ~60 seconds |
Summary
- 512³ requires ≥ 24 GB VRAM — the baseline for TRELLIS.2 operation
- 1024³ requires ≈ 48 GB VRAM — roughly 2× the baseline memory
- 1536³ requires ≈ 80 GB VRAM — full H100-80GB capacity
- Memory scaling follows cubic growth due to dense O-Voxel attribute storage
- Set
PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"for constrained GPUs - Resolution selection occurs at pipeline initialization via the
resolutionparameter
Frequently Asked Questions
What happens if I try to run 1536³ on a 24 GB GPU?
The inference will fail with a CUDA out-of-memory error during the VAE decode phase. The 1536³ mode requires approximately 80 GB VRAM because the dense voxel buffer alone exceeds 24 GB before accounting for model weights and intermediate activations. Stick to 512³ for 24 GB cards or use cloud H100 instances for maximum resolution.
Can I use multiple GPUs to split the memory load?
TRELLIS.2 does not implement native multi-GPU parallelism for single inference runs. The trellis2_image_to_3d.py pipeline loads the complete model onto a single CUDA device. Workarounds would require manual model parallelism implementation not provided in the open-source release.
Does batch size affect GPU memory requirements?
TRELLIS.2 processes single images per inference call, so batch size is fixed at 1. The O-Voxel representation and post-processing pipeline are designed for sequential generation. Memory consumption is dominated by the output resolution voxel buffer rather than batch dimensions, making per-image VRAM requirements consistent.
Why does 1024³ only double memory instead of increasing 8×?
The sparse-voxel VAE and shared attribute tables reduce effective memory growth. While raw voxel count increases 8×, the implementation uses 16× downsampled latent representations and efficient attribute storage that compress the in-flight memory footprint. The remaining growth factor of ~2× reflects the decoded dense buffer and intermediate feature maps required at higher resolution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →