# How to Enable Low‑VRAM Mode and Configure PYTORCH_CUDA_ALLOC_CONF for Memory Optimization in TRELLIS.2

> Optimize TRELLIS.2 memory usage by enabling low VRAM mode and configuring PYTORCH_CUDA_ALLOC_CONF. Reduce VRAM consumption up to 70% by loading model weights on CPU.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: how-to-guide
- Published: 2026-08-04

---

**Set `PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"` before importing torch, and instantiate `Trellis2ImageTo3DPipeline` with `low_vram=True` to keep heavy model weights on CPU and only move them to GPU during inference, reducing peak VRAM usage by up to 70%.**

Running large 3D generation models like TRELLIS.2 on consumer GPUs requires aggressive memory management. This guide explains how to enable low_vram mode and configure PYTORCH_CUDA_ALLOC_CONF to minimize CUDA memory allocation and prevent out-of-memory errors during sparse structure and texture generation.

## Understanding Low‑VRAM Mode in TRELLIS.2

TRELLIS.2 implements a **low‑VRAM mode** that defers loading heavy model components onto the GPU until they are actively needed for inference. This keeps most of the model architecture in system RAM, dramatically lowering peak GPU allocation while maintaining generation quality.

### The low_vram Constructor Flag

According to the source code in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) (lines 29‑55), the pipeline accepts a `low_vram` boolean parameter in its constructor:

```python
def __init__(…, low_vram: bool = True, …):
    …
    self.low_vram = low_vram               # enables low‑VRAM mode

```

When `low_vram=True` (the default), the pipeline keeps the sparse structure flow models, shape SLAT flow models, and texture SLAT flow models on the CPU. Only lightweight components like the image conditioning encoder permanently reside on the GPU.

### Conditional Device Transfer in the to() Method

The `to()` method in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) (lines 19‑26) demonstrates how the pipeline conditionally moves models based on this flag:

```python
def to(self, device: torch.device) -> None:
    self._device = device
    if not self.low_vram:                 # skip when low‑VRAM is on

        super().to(device)
        self.image_cond_model.to(device)
        if self.rembg_model is not None:
            self.rembg_model.to(device)

```

When `low_vram` is enabled, this method skips moving the heavy background removal and image conditioning models to the GPU during initialization.

## How Low‑VRAM Mode Swaps Models During Inference

The memory optimization works by temporarily moving specific models to the GPU only for their required computation steps, then immediately returning them to CPU RAM. This on‑demand loading occurs at three critical points in the pipeline.

### Background Removal Model Swapping

In `preprocess_image()` ([`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), lines 44‑50), the background removal model moves to the device only for the removal step:

```python
if self.low_vram:
    self.rembg_model.to(self.device)
output = self.rembg_model(input)
if self.low_vram:
    self.rembg_model.cpu()

```

### Image Conditioning Model Swapping

Similarly, in `get_cond()` ([`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), lines 75‑80), the image conditioning model transfers temporarily:

```python
if self.low_vram:
    self.image_cond_model.to(self.device)
cond = self.image_cond_model(image)
if self.low_vram:
    self.image_cond_model.cpu()

```

### Flow Model Sampling Loop Swapping

During the actual 3D generation in methods like `sample_sparse_structure` ([`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), lines 241‑255), the heavy flow models move to the GPU immediately before the sampling loop and return to CPU afterward:

```python
if self.low_vram:
    flow_model.to(self.device)          # move just before sampling

… # sampling steps

if self.low_vram:
    flow_model.cpu()                    # free GPU memory afterwards

```

The [`trellis2/pipelines/trellis2_texturing.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_texturing.py) file implements identical logic for texture generation, ensuring consistent memory management across both pipelines.

## Configuring PYTORCH_CUDA_ALLOC_CONF

While low‑VRAM mode manages model placement, **`PYTORCH_CUDA_ALLOC_CONF`** optimizes how PyTorch's CUDA memory allocator manages freed tensors. This environment variable must be set **before** importing torch, as the allocator initializes at import time.

### Setting Expandable Segments

The repository examples in [`example.py`](https://github.com/microsoft/TRELLIS.2/blob/main/example.py), [`app.py`](https://github.com/microsoft/TRELLIS.2/blob/main/app.py), and [`example_texturing.py`](https://github.com/microsoft/TRELLIS.2/blob/main/example_texturing.py) (lines 2‑5) set the allocator to use expandable segments:

```python
import os

# Enable expandable‑segment allocator (saves memory by reusing freed blocks)

os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"

```

This configuration prevents the allocator from holding onto large fixed‑size memory pools after tensors are freed. Instead, the allocator can shrink its pool when memory is released, reducing the overall GPU footprint—especially effective when combined with low‑VRAM mode's frequent allocations and deallocations.

### Alternative Allocator Configurations

For specific hardware constraints, you can tune additional parameters before importing torch:

```python
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:128,garbage_collection_threshold:0.6"
import torch

```

## Complete Production Implementation

Combine both optimizations in your entry script. The following example demonstrates the correct initialization order and pipeline configuration:

```python

# --------------------------------------------------------------

# 0️⃣  Set allocator options BEFORE importing torch

# --------------------------------------------------------------

import os
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"

# --------------------------------------------------------------

# 1️⃣  Imports (torch is now loaded with the custom allocator)

# --------------------------------------------------------------

import torch
from PIL import Image
from trellis2.pipelines import Trellis2ImageTo3DPipeline

# --------------------------------------------------------------

# 2️⃣  Build the pipeline with low‑VRAM enabled (default)

# --------------------------------------------------------------

pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    "microsoft/TRELLIS.2-4B",
    # low_vram=True is the default; you can omit it or set explicitly

    low_vram=True
)
pipeline.to(torch.device("cuda"))          # move only the lightweight parts

# --------------------------------------------------------------

# 3️⃣  Run inference (heavy flow models will be swapped in/out automatically)

# --------------------------------------------------------------

input_image = Image.open("assets/example_image/T.png")
mesh = pipeline.run(input_image)[0]        # returns a MeshWithVoxel object

```

All heavy components (`sparse_structure_flow_model`, `shape_slat_flow_model_*`, `tex_slat_flow_model_*`) remain on CPU until the sampling code temporarily moves them to the GPU, keeping peak GPU usage well under the full model size.

## When to Disable Low‑VRAM Mode

**Disable low‑VRAM mode** by passing `low_vram=False` when you have ample GPU memory (≥ 24 GB) and require maximum inference speed. When disabled, all models reside permanently on the GPU, eliminating the overhead of CPU‑to‑GPU transfers during each sampling step.

```python
pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    "microsoft/TRELLIS.2-4B",
    low_vram=False  # Load all models to GPU immediately

)
pipeline.to('cuda')

```

## Summary

- **Set `PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"`** before any `import torch` statement to enable efficient memory reuse in the CUDA allocator.
- **Enable `low_vram=True`** (the default) when constructing `Trellis2ImageTo3DPipeline` to keep heavy flow models on CPU and only transfer them to GPU during active inference.
- **Heavy model components** (sparse structure, shape SLAT, and texture SLAT flow models) temporarily move to the GPU inside sampling loops in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) and [`trellis2/pipelines/trellis2_texturing.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_texturing.py), then return to CPU immediately after computation.
- **Use `low_vram=False`** only on GPUs with 24 GB or more VRAM when inference latency is more critical than memory conservation.

## Frequently Asked Questions

### Does low_vram mode affect generation quality?

No, low_vram mode does not affect generation quality. It only changes the physical location of model weights (CPU versus GPU) and the timing of tensor transfers. The same checkpoint weights and inference algorithms apply, ensuring identical mesh outputs regardless of memory management strategy.

### Why must PYTORCH_CUDA_ALLOC_CONF be set before importing torch?

PyTorch initializes its CUDA memory allocator immediately upon importing the torch module. Once initialized, the allocator configuration becomes read‑only. Setting the environment variable after `import torch` has no effect because the allocator has already created its memory pools using default parameters.

### Can I use low_vram mode on a GPU with 8GB of VRAM?

Yes, low_vram mode is specifically designed for consumer‑grade GPUs with limited VRAM. By keeping the 4B parameter flow models on CPU and only moving active batches to GPU during sampling steps, TRELLIS.2 can generate 3D assets on GPUs with as little as 8‑12 GB of memory, though generation time will increase due to PCIe transfer overhead.

### What happens if I set low_vram=False on a GPU with insufficient memory?

If you disable low_vram mode on a GPU without enough VRAM to hold the full model (approximately 16‑20 GB for the 4B checkpoint), PyTorch will raise a `torch.cuda.OutOfMemoryError` during the first forward pass when the pipeline attempts to simultaneously load the sparse structure, shape, and texture flow models onto the device.