How to Enable Low‑VRAM Mode and Configure PYTORCH_CUDA_ALLOC_CONF for Memory Optimization in TRELLIS.2
Set PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True" before importing torch, and instantiate Trellis2ImageTo3DPipeline with low_vram=True to keep heavy model weights on CPU and only move them to GPU during inference, reducing peak VRAM usage by up to 70%.
Running large 3D generation models like TRELLIS.2 on consumer GPUs requires aggressive memory management. This guide explains how to enable low_vram mode and configure PYTORCH_CUDA_ALLOC_CONF to minimize CUDA memory allocation and prevent out-of-memory errors during sparse structure and texture generation.
Understanding Low‑VRAM Mode in TRELLIS.2
TRELLIS.2 implements a low‑VRAM mode that defers loading heavy model components onto the GPU until they are actively needed for inference. This keeps most of the model architecture in system RAM, dramatically lowering peak GPU allocation while maintaining generation quality.
The low_vram Constructor Flag
According to the source code in trellis2/pipelines/trellis2_image_to_3d.py (lines 29‑55), the pipeline accepts a low_vram boolean parameter in its constructor:
def __init__(…, low_vram: bool = True, …):
…
self.low_vram = low_vram # enables low‑VRAM mode
When low_vram=True (the default), the pipeline keeps the sparse structure flow models, shape SLAT flow models, and texture SLAT flow models on the CPU. Only lightweight components like the image conditioning encoder permanently reside on the GPU.
Conditional Device Transfer in the to() Method
The to() method in trellis2/pipelines/trellis2_image_to_3d.py (lines 19‑26) demonstrates how the pipeline conditionally moves models based on this flag:
def to(self, device: torch.device) -> None:
self._device = device
if not self.low_vram: # skip when low‑VRAM is on
super().to(device)
self.image_cond_model.to(device)
if self.rembg_model is not None:
self.rembg_model.to(device)
When low_vram is enabled, this method skips moving the heavy background removal and image conditioning models to the GPU during initialization.
How Low‑VRAM Mode Swaps Models During Inference
The memory optimization works by temporarily moving specific models to the GPU only for their required computation steps, then immediately returning them to CPU RAM. This on‑demand loading occurs at three critical points in the pipeline.
Background Removal Model Swapping
In preprocess_image() (trellis2/pipelines/trellis2_image_to_3d.py, lines 44‑50), the background removal model moves to the device only for the removal step:
if self.low_vram:
self.rembg_model.to(self.device)
output = self.rembg_model(input)
if self.low_vram:
self.rembg_model.cpu()
Image Conditioning Model Swapping
Similarly, in get_cond() (trellis2/pipelines/trellis2_image_to_3d.py, lines 75‑80), the image conditioning model transfers temporarily:
if self.low_vram:
self.image_cond_model.to(self.device)
cond = self.image_cond_model(image)
if self.low_vram:
self.image_cond_model.cpu()
Flow Model Sampling Loop Swapping
During the actual 3D generation in methods like sample_sparse_structure (trellis2/pipelines/trellis2_image_to_3d.py, lines 241‑255), the heavy flow models move to the GPU immediately before the sampling loop and return to CPU afterward:
if self.low_vram:
flow_model.to(self.device) # move just before sampling
… # sampling steps
if self.low_vram:
flow_model.cpu() # free GPU memory afterwards
The trellis2/pipelines/trellis2_texturing.py file implements identical logic for texture generation, ensuring consistent memory management across both pipelines.
Configuring PYTORCH_CUDA_ALLOC_CONF
While low‑VRAM mode manages model placement, PYTORCH_CUDA_ALLOC_CONF optimizes how PyTorch's CUDA memory allocator manages freed tensors. This environment variable must be set before importing torch, as the allocator initializes at import time.
Setting Expandable Segments
The repository examples in example.py, app.py, and example_texturing.py (lines 2‑5) set the allocator to use expandable segments:
import os
# Enable expandable‑segment allocator (saves memory by reusing freed blocks)
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
This configuration prevents the allocator from holding onto large fixed‑size memory pools after tensors are freed. Instead, the allocator can shrink its pool when memory is released, reducing the overall GPU footprint—especially effective when combined with low‑VRAM mode's frequent allocations and deallocations.
Alternative Allocator Configurations
For specific hardware constraints, you can tune additional parameters before importing torch:
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:128,garbage_collection_threshold:0.6"
import torch
Complete Production Implementation
Combine both optimizations in your entry script. The following example demonstrates the correct initialization order and pipeline configuration:
# --------------------------------------------------------------
# 0️⃣ Set allocator options BEFORE importing torch
# --------------------------------------------------------------
import os
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
# --------------------------------------------------------------
# 1️⃣ Imports (torch is now loaded with the custom allocator)
# --------------------------------------------------------------
import torch
from PIL import Image
from trellis2.pipelines import Trellis2ImageTo3DPipeline
# --------------------------------------------------------------
# 2️⃣ Build the pipeline with low‑VRAM enabled (default)
# --------------------------------------------------------------
pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
"microsoft/TRELLIS.2-4B",
# low_vram=True is the default; you can omit it or set explicitly
low_vram=True
)
pipeline.to(torch.device("cuda")) # move only the lightweight parts
# --------------------------------------------------------------
# 3️⃣ Run inference (heavy flow models will be swapped in/out automatically)
# --------------------------------------------------------------
input_image = Image.open("assets/example_image/T.png")
mesh = pipeline.run(input_image)[0] # returns a MeshWithVoxel object
All heavy components (sparse_structure_flow_model, shape_slat_flow_model_*, tex_slat_flow_model_*) remain on CPU until the sampling code temporarily moves them to the GPU, keeping peak GPU usage well under the full model size.
When to Disable Low‑VRAM Mode
Disable low‑VRAM mode by passing low_vram=False when you have ample GPU memory (≥ 24 GB) and require maximum inference speed. When disabled, all models reside permanently on the GPU, eliminating the overhead of CPU‑to‑GPU transfers during each sampling step.
pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
"microsoft/TRELLIS.2-4B",
low_vram=False # Load all models to GPU immediately
)
pipeline.to('cuda')
Summary
- Set
PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"before anyimport torchstatement to enable efficient memory reuse in the CUDA allocator. - Enable
low_vram=True(the default) when constructingTrellis2ImageTo3DPipelineto keep heavy flow models on CPU and only transfer them to GPU during active inference. - Heavy model components (sparse structure, shape SLAT, and texture SLAT flow models) temporarily move to the GPU inside sampling loops in
trellis2/pipelines/trellis2_image_to_3d.pyandtrellis2/pipelines/trellis2_texturing.py, then return to CPU immediately after computation. - Use
low_vram=Falseonly on GPUs with 24 GB or more VRAM when inference latency is more critical than memory conservation.
Frequently Asked Questions
Does low_vram mode affect generation quality?
No, low_vram mode does not affect generation quality. It only changes the physical location of model weights (CPU versus GPU) and the timing of tensor transfers. The same checkpoint weights and inference algorithms apply, ensuring identical mesh outputs regardless of memory management strategy.
Why must PYTORCH_CUDA_ALLOC_CONF be set before importing torch?
PyTorch initializes its CUDA memory allocator immediately upon importing the torch module. Once initialized, the allocator configuration becomes read‑only. Setting the environment variable after import torch has no effect because the allocator has already created its memory pools using default parameters.
Can I use low_vram mode on a GPU with 8GB of VRAM?
Yes, low_vram mode is specifically designed for consumer‑grade GPUs with limited VRAM. By keeping the 4B parameter flow models on CPU and only moving active batches to GPU during sampling steps, TRELLIS.2 can generate 3D assets on GPUs with as little as 8‑12 GB of memory, though generation time will increase due to PCIe transfer overhead.
What happens if I set low_vram=False on a GPU with insufficient memory?
If you disable low_vram mode on a GPU without enough VRAM to hold the full model (approximately 16‑20 GB for the 4B checkpoint), PyTorch will raise a torch.cuda.OutOfMemoryError during the first forward pass when the pipeline attempts to simultaneously load the sparse structure, shape, and texture flow models onto the device.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →