TRELLIS.2 GPU Memory Requirements by Output Resolution: Complete Hardware Guide

TRELLIS.2 requires a minimum of 24 GB GPU VRAM for 512³ output, approximately 48 GB for 1024³, and roughly 80 GB for 1536³, with memory usage scaling cubically due to dense voxel attribute storage.

TRELLIS.2, Microsoft's open-source 3D generation model, produces high-quality textured meshes through its novel O-Voxel representation. Understanding GPU memory requirements by resolution is critical for deployment planning, as the per-voxel geometry and material attributes create substantial VRAM demands that scale with the cube of linear resolution.

Supported Resolutions and Memory Requirements

TRELLIS.2 offers three discrete output resolutions, each with distinct hardware demands. The repository's README.md documents these specifications explicitly.

512³ Resolution: Baseline Configuration

Minimum GPU Memory: ≥ 24 GB

This is the entry-level resolution supported by TRELLIS.2. Inference completes in approximately 3 seconds total—2 seconds for initial generation plus 1 second for post-processing. The 24 GB requirement represents the absolute hardware floor for running the model.

According to the Hardware section of the repository, this baseline ensures sufficient headroom for the O-Voxel decoder and attribute buffers.

1024³ Resolution: Mid-Range Production

Approximate GPU Memory: ≈ 48 GB

Stepping up to 1024³ increases inference time to roughly 17 seconds (10 s + 7 s). While the voxel count increases 8× from 512³ to 1024³, shared attribute tables and streaming optimizations reduce practical memory growth to approximately 2×.

This configuration fits comfortably on NVIDIA A100 (40 GB) and H100 (80 GB) GPUs. Users with 40 GB cards may need memory optimization techniques (see below).

1536³ Resolution: Maximum Quality

Approximate GPU Memory: ≈ 80 GB

The highest-resolution mode demands the full capacity of an H100-80GB GPU, with inference taking approximately 60 seconds (35 s + 25 s). The authors explicitly verified this configuration on H100 hardware, indicating that 1536³ consumes nearly all available VRAM on current flagship GPUs.

Why GPU Memory Scales Cubically

TRELLIS.2's O-Voxel representation stores dense per-voxel attributes including:

  • Base color (RGB)
  • Roughness
  • Metallic
  • Opacity
  • Spatial coordinates

The total voxel count grows as N³ with linear resolution. While the sparse-voxel VAE encoder performs 16× downsampling, the decoder expands latent representations back to full target resolution during inference. This requires the GPU to maintain the complete dense representation in memory.

The postprocess.py module implements chunked rasterization to limit peak memory during mesh export, but the primary inference path still holds full-resolution buffers.

Hardware Verification and Testing

The TRELLIS.2 authors validated their memory specifications on NVIDIA H100 GPUs with 80 GB VRAM. This testing environment confirms:

  • 512³ and 1024³ modes run reliably within available memory
  • 1536³ operates at the hardware limit with minimal headroom

Users attempting 1536³ generation on lower-capacity GPUs (A100-40GB or consumer RTX cards) will encounter out-of-memory errors during the decode phase.

Memory Optimization for Constrained GPUs

For GPUs below recommended specifications, TRELLIS.2 supports flexible memory allocation through PyTorch configuration.

Environment Variable Configuration

Set PYTORCH_CUDA_ALLOC_CONF before importing torch:

import os
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"

import torch
from trellis2.pipelines import Trellis2ImageTo3DPipeline

# Initialize pipeline at reduced resolution

pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    "microsoft/TRELLIS.2-4B",
    resolution="512"  # Use lowest resolution for 24 GB GPUs

)
pipeline.cuda()

This technique, shown in example.py, enables more flexible memory block allocation and may allow 1024³ generation on 40 GB GPUs that would otherwise fail.

Resolution Selection in Pipeline

The resolution parameter directly controls voxel grid allocation in trellis2_image_to_3d.py:

from PIL import Image
from trellis2.pipelines import Trellis2ImageTo3DPipeline

# Available: "512", "1024", "1536"

RESOLUTION = "1024"

pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    "microsoft/TRELLIS.2-4B",
    resolution=RESOLUTION
)
pipeline.cuda()

img = Image.open("input.png")
mesh = pipeline.run(img)[0]

Practical Deployment Recommendations

Use Case Recommended Resolution Minimum GPU Expected Performance
Rapid prototyping, limited VRAM 512³ RTX 3090 (24 GB) ~3 seconds
Production quality, mid-range servers 1024³ A100-40GB or above ~17 seconds
Maximum fidelity, research 1536³ H100-80GB exclusively ~60 seconds

Summary

  • 512³ requires ≥ 24 GB VRAM — the baseline for TRELLIS.2 operation
  • 1024³ requires ≈ 48 GB VRAM — roughly 2× the baseline memory
  • 1536³ requires ≈ 80 GB VRAM — full H100-80GB capacity
  • Memory scaling follows cubic growth due to dense O-Voxel attribute storage
  • Set PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True" for constrained GPUs
  • Resolution selection occurs at pipeline initialization via the resolution parameter

Frequently Asked Questions

What happens if I try to run 1536³ on a 24 GB GPU?

The inference will fail with a CUDA out-of-memory error during the VAE decode phase. The 1536³ mode requires approximately 80 GB VRAM because the dense voxel buffer alone exceeds 24 GB before accounting for model weights and intermediate activations. Stick to 512³ for 24 GB cards or use cloud H100 instances for maximum resolution.

Can I use multiple GPUs to split the memory load?

TRELLIS.2 does not implement native multi-GPU parallelism for single inference runs. The trellis2_image_to_3d.py pipeline loads the complete model onto a single CUDA device. Workarounds would require manual model parallelism implementation not provided in the open-source release.

Does batch size affect GPU memory requirements?

TRELLIS.2 processes single images per inference call, so batch size is fixed at 1. The O-Voxel representation and post-processing pipeline are designed for sequential generation. Memory consumption is dominated by the output resolution voxel buffer rather than batch dimensions, making per-image VRAM requirements consistent.

Why does 1024³ only double memory instead of increasing 8×?

The sparse-voxel VAE and shared attribute tables reduce effective memory growth. While raw voxel count increases 8×, the implementation uses 16× downsampled latent representations and efficient attribute storage that compress the in-flight memory footprint. The remaining growth factor of ~2× reflects the decoded dense buffer and intermediate feature maps required at higher resolution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →