How TRELLIS.2 Low VRAM Mode Solves GPU Memory Constraints
TRELLIS.2 low VRAM mode minimizes GPU memory consumption by keeping models on the CPU and only moving them to the GPU during active forward passes, enabling high-quality 3D generation on consumer hardware with as little as 8GB of VRAM.
The microsoft/TRELLIS.2 repository implements a sophisticated memory management strategy that allows users to run large-scale 3D texturing and reconstruction pipelines on GPUs with limited memory. By strategically orchestrating device placement across the inference pipeline, TRELLIS.2 low VRAM mode ensures that only the currently active model resides in GPU memory at any given time.
How TRELLIS.2 Low VRAM Mode Works
The implementation centers on three architectural principles that work together to reduce the peak VRAM footprint to roughly the size of a single model plus temporary activations.
Deferred GPU Allocation
Instead of loading the entire pipeline onto the GPU at initialization, TRELLIS.2 low VRAM mode defers allocation until the moment of computation. In trellis2/pipelines/trellis2_texturing.py, the to() method checks self.low_vram before deciding how to handle device placement. When the flag is False, the entire pipeline moves to the target device upfront; when True, models remain on the CPU until explicitly needed for a forward pass.
This logic appears at lines 100-105, where the method selectively moves sub-modules based on the low_vram state:
# Conceptual implementation from trellis2_texturing.py L100-L105
if not self.low_vram:
# Standard mode: move everything to GPU
super().to(device)
else:
# Low VRAM mode: keep on CPU, track target device for later use
self.device = device
Immediate CPU Fallback
After each heavy computation completes, the pipeline immediately returns models to the CPU to free VRAM. This pattern repeats across critical pipeline stages including encode_shape_slat, sample_tex_slat, and decode_tex_slat.
For example, at lines 172-173, the shape encoder follows this workflow:
# From trellis2/pipelines/trellis2_texturing.py L172-L173
if self.low_vram:
self.shape_encoder.to(self.device)
# ... perform encoding ...
self.shape_encoder.cpu()
The texture sampler (lines 145-150) and texture decoder (lines 160-164) implement identical patterns, ensuring that VRAM is released before the next model loads.
Selective Device Usage for Auxiliary Models
Even auxiliary components like the background removal model (rembg) participate in the memory-saving strategy. During preprocess_image at lines 140-145, the pipeline temporarily moves self.rembg_model to the GPU only for the duration of the background removal operation:
# From trellis2/pipelines/trellis2_texturing.py L140-L145
if self.low_vram:
self.rembg_model.to(self.device)
# ... remove background ...
self.rembg_model.cpu()
Source Code Architecture
The TRELLIS.2 low VRAM implementation spans multiple files within the repository, with the primary logic concentrated in the pipeline definitions.
The file trellis2/pipelines/trellis2_texturing.py contains the core Trellis2TexturingPipeline class and handles the low_vram flag for model device placement, shape encoding, texture sampling, and background removal. The companion file trellis2/pipelines/trellis2_image_to_3d.py provides Trellis2ImageTo3DPipeline, which mirrors the same low-VRAM logic for image-to-3D reconstruction tasks.
Supporting utilities in trellis2/utils/general_utils.py provide device handling functions that facilitate the CPU-GPU transfers, while trellis2/modules/sparse/__init__.py defines the SparseTensor abstraction used throughout these memory-efficient workflows.
Enabling and Configuring Low VRAM Mode
By default, TRELLIS.2 pipelines initialize with low_vram=True, making them accessible to users with limited GPU memory. You can explicitly control this behavior when loading pretrained checkpoints.
Default Low VRAM Usage
To run the texturing pipeline with default memory-efficient settings:
from trellis2.pipelines.trellis2_texturing import Trellis2TexturingPipeline
import torch
from PIL import Image
import trimesh
# Load pipeline (low_vram=True by default in config)
pipeline = Trellis2TexturingPipeline.from_pretrained("path/to/checkpoint")
pipeline.to(torch.device("cuda:0"))
# Prepare inputs
mesh = trimesh.load("assets/example.obj")
image = Image.open("assets/example.jpg").convert("RGB")
# Run inference with per-step CPU fallback
textured_mesh = pipeline.run(mesh, image)
textured_mesh.export("output_textured.glb")
In this configuration, the pipeline internally toggles GPU usage for each sub-model, moving the shape encoder to CUDA only during encode_shape_slat, the texture decoder only during decode_tex_slat, and so on.
Disabling Low VRAM for Maximum Speed
If your hardware has sufficient VRAM to hold all models simultaneously, disabling low VRAM mode eliminates the CPU-GPU transfer overhead:
from trellis2.pipelines.trellis2_image_to_3d import Trellis2ImageTo3DPipeline
# Load and explicitly disable low VRAM mode
pipeline = Trellis2ImageTo3DPipeline.from_pretrained("path/to/checkpoint")
pipeline.low_vram = False
pipeline.to(torch.device("cuda:0"))
# All models stay resident on GPU for faster inference
mesh = pipeline.run(Image.open("photo.jpg"))
mesh.export("reconstructed.glb")
Performance and Memory Trade-offs
While TRELLIS.2 low VRAM mode significantly reduces memory requirements, it introduces a small runtime overhead due to model transfer costs. Each model.to(device) and model.cpu() operation incurs a copy cost between system RAM and VRAM. However, this overhead is typically negligible compared to the computation time of the forward passes themselves, and it enables execution on hardware that would otherwise fail with out-of-memory errors.
The granular control provided by the low_vram flag ensures consistent memory-saving behavior across the entire pipeline while maintaining a device-agnostic API. Whether low_vram is enabled or disabled, the pipeline accepts standard torch.device arguments and manages internal state accordingly.
Summary
- TRELLIS.2 low VRAM mode implements deferred GPU allocation, keeping models on CPU until immediately before use and returning them afterward.
- The
Trellis2TexturingPipeline.to()method intrellis2/pipelines/trellis2_texturing.py(lines 100-105) checks thelow_vramflag to determine allocation strategy. - Critical operations like
encode_shape_slat,sample_tex_slat, anddecode_tex_slatwrap their forward passes with conditional device transfers to minimize peak VRAM usage. - Auxiliary models including the background remover follow the same selective device placement pattern.
- Users can toggle the
low_vramboolean attribute on any pipeline instance to optimize for either memory efficiency (default) or inference speed.
Frequently Asked Questions
How much VRAM does TRELLIS.2 require with low VRAM mode enabled?
With TRELLIS.2 low VRAM mode activated, the pipeline requires approximately the memory footprint of a single model plus temporary activations for one forward pass, typically allowing operation on consumer GPUs with 8GB of VRAM. Without this mode, the full pipeline suite would require significantly more memory to keep all models resident simultaneously.
Does enabling low VRAM mode affect output quality?
No, TRELLIS.2 low VRAM mode preserves identical inference quality to the standard mode. The mechanism only changes device placement strategy—moving models between CPU and GPU—without modifying model weights, precision, or sampling algorithms. The reconstructed meshes and textures remain identical regardless of the low_vram setting.
Which pipeline classes support the low VRAM mode?
Both Trellis2TexturingPipeline from trellis2/pipelines/trellis2_texturing.py and Trellis2ImageTo3DPipeline from trellis2/pipelines/trellis2_image_to_3d.py fully implement the low VRAM functionality. Each propagates the low_vram flag to their respective sub-modules, including shape encoders, texture decoders, and auxiliary preprocessing models like rembg.
What is the performance impact of low VRAM mode?
Low VRAM mode introduces minor latency overhead due to CPU-GPU memory transfers before and after each forward pass. According to the source implementation in lines 145-150 and 160-164 of the texturing pipeline, each model transfer adds a copy cost, but this is typically outweighed by the benefit of being able to run the pipeline at all on memory-constrained hardware. For production deployments with ample VRAM, setting low_vram=False eliminates this overhead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →