# TRELLIS.2 GPU Memory Requirements by Output Resolution: Complete Hardware Guide

> Discover TRELLIS.2 GPU memory requirements for 512, 1024, and 1536 resolutions. Understand VRAM needs for your hardware and optimize performance with our complete guide.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: hardware-guide
- Published: 2026-08-04

---

**TRELLIS.2 requires a minimum of 24 GB GPU VRAM for 512³ output, approximately 48 GB for 1024³, and roughly 80 GB for 1536³**, with memory usage scaling cubically due to dense voxel attribute storage.

TRELLIS.2, Microsoft's open-source 3D generation model, produces high-quality textured meshes through its novel O-Voxel representation. Understanding GPU memory requirements by resolution is critical for deployment planning, as the per-voxel geometry and material attributes create substantial VRAM demands that scale with the cube of linear resolution.

## Supported Resolutions and Memory Requirements

TRELLIS.2 offers three discrete output resolutions, each with distinct hardware demands. The repository's [README.md](https://github.com/microsoft/TRELLIS.2/blob/main/README.md) documents these specifications explicitly.

### 512³ Resolution: Baseline Configuration

**Minimum GPU Memory: ≥ 24 GB**

This is the entry-level resolution supported by TRELLIS.2. Inference completes in approximately 3 seconds total—2 seconds for initial generation plus 1 second for post-processing. The 24 GB requirement represents the absolute hardware floor for running the model.

According to the [Hardware section](https://github.com/microsoft/TRELLIS.2/blob/main/README.md#L60-L61) of the repository, this baseline ensures sufficient headroom for the O-Voxel decoder and attribute buffers.

### 1024³ Resolution: Mid-Range Production

**Approximate GPU Memory: ≈ 48 GB**

Stepping up to 1024³ increases inference time to roughly 17 seconds (10 s + 7 s). While the voxel count increases 8× from 512³ to 1024³, shared attribute tables and streaming optimizations reduce practical memory growth to approximately 2×.

This configuration fits comfortably on **NVIDIA A100 (40 GB)** and **H100 (80 GB)** GPUs. Users with 40 GB cards may need memory optimization techniques (see below).

### 1536³ Resolution: Maximum Quality

**Approximate GPU Memory: ≈ 80 GB**

The highest-resolution mode demands the full capacity of an **H100-80GB GPU**, with inference taking approximately 60 seconds (35 s + 25 s). The authors explicitly [verified this configuration on H100 hardware](https://github.com/microsoft/TRELLIS.2/blob/main/README.md#L27-L29), indicating that 1536³ consumes nearly all available VRAM on current flagship GPUs.

## Why GPU Memory Scales Cubically

TRELLIS.2's **O-Voxel representation** stores dense per-voxel attributes including:

- Base color (RGB)
- Roughness
- Metallic
- Opacity
- Spatial coordinates

The total voxel count grows as N³ with linear resolution. While the sparse-voxel VAE encoder performs 16× downsampling, the decoder expands latent representations back to full target resolution during inference. This requires the GPU to maintain the complete dense representation in memory.

The [postprocess.py](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/postprocess.py) module implements chunked rasterization to limit peak memory during mesh export, but the primary inference path still holds full-resolution buffers.

## Hardware Verification and Testing

The TRELLIS.2 authors validated their memory specifications on **NVIDIA H100 GPUs with 80 GB VRAM**. This testing environment confirms:

- 512³ and 1024³ modes run reliably within available memory
- 1536³ operates at the hardware limit with minimal headroom

Users attempting 1536³ generation on lower-capacity GPUs (A100-40GB or consumer RTX cards) will encounter out-of-memory errors during the decode phase.

## Memory Optimization for Constrained GPUs

For GPUs below recommended specifications, TRELLIS.2 supports flexible memory allocation through PyTorch configuration.

### Environment Variable Configuration

Set `PYTORCH_CUDA_ALLOC_CONF` before importing torch:

```python
import os
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"

import torch
from trellis2.pipelines import Trellis2ImageTo3DPipeline

# Initialize pipeline at reduced resolution

pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    "microsoft/TRELLIS.2-4B",
    resolution="512"  # Use lowest resolution for 24 GB GPUs

)
pipeline.cuda()

```

This technique, [shown in example.py](https://github.com/microsoft/TRELLIS.2/blob/main/example.py), enables more flexible memory block allocation and may allow 1024³ generation on 40 GB GPUs that would otherwise fail.

### Resolution Selection in Pipeline

The resolution parameter directly controls voxel grid allocation in [trellis2_image_to_3d.py](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py):

```python
from PIL import Image
from trellis2.pipelines import Trellis2ImageTo3DPipeline

# Available: "512", "1024", "1536"

RESOLUTION = "1024"

pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    "microsoft/TRELLIS.2-4B",
    resolution=RESOLUTION
)
pipeline.cuda()

img = Image.open("input.png")
mesh = pipeline.run(img)[0]

```

## Practical Deployment Recommendations

| Use Case | Recommended Resolution | Minimum GPU | Expected Performance |
|----------|------------------------|-------------|----------------------|
| Rapid prototyping, limited VRAM | 512³ | RTX 3090 (24 GB) | ~3 seconds |
| Production quality, mid-range servers | 1024³ | A100-40GB or above | ~17 seconds |
| Maximum fidelity, research | 1536³ | H100-80GB exclusively | ~60 seconds |

## Summary

- **512³ requires ≥ 24 GB VRAM** — the baseline for TRELLIS.2 operation
- **1024³ requires ≈ 48 GB VRAM** — roughly 2× the baseline memory
- **1536³ requires ≈ 80 GB VRAM** — full H100-80GB capacity
- Memory scaling follows cubic growth due to dense O-Voxel attribute storage
- Set `PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"` for constrained GPUs
- Resolution selection occurs at pipeline initialization via the `resolution` parameter

## Frequently Asked Questions

### What happens if I try to run 1536³ on a 24 GB GPU?

**The inference will fail with a CUDA out-of-memory error during the VAE decode phase.** The 1536³ mode requires approximately 80 GB VRAM because the dense voxel buffer alone exceeds 24 GB before accounting for model weights and intermediate activations. Stick to 512³ for 24 GB cards or use cloud H100 instances for maximum resolution.

### Can I use multiple GPUs to split the memory load?

**TRELLIS.2 does not implement native multi-GPU parallelism for single inference runs.** The [trellis2_image_to_3d.py](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) pipeline loads the complete model onto a single CUDA device. Workarounds would require manual model parallelism implementation not provided in the open-source release.

### Does batch size affect GPU memory requirements?

**TRELLIS.2 processes single images per inference call, so batch size is fixed at 1.** The O-Voxel representation and post-processing pipeline are designed for sequential generation. Memory consumption is dominated by the output resolution voxel buffer rather than batch dimensions, making per-image VRAM requirements consistent.

### Why does 1024³ only double memory instead of increasing 8×?

**The sparse-voxel VAE and shared attribute tables reduce effective memory growth.** While raw voxel count increases 8×, the implementation uses 16× downsampled latent representations and efficient attribute storage that compress the in-flight memory footprint. The remaining growth factor of ~2× reflects the decoded dense buffer and intermediate feature maps required at higher resolution.