How to Use `Trellis2ImageTo3DPipeline.from_pretrained()` for Image-to-3D Inference
Call Trellis2ImageTo3DPipeline.from_pretrained("microsoft/TRELLIS.2-4B") to load a pretrained checkpoint from Hugging Face or a local directory, move it to GPU with .cuda(), then execute pipeline.run(image) to generate a fully textured 3D mesh from a single 2D image.
The Trellis2ImageTo3DPipeline is the high-level inference interface in Microsoft's TRELLIS.2 repository that orchestrates the complete image-to-3D generation workflow. This pipeline inherits from the generic Pipeline base class defined in trellis2/pipelines/base.py and handles everything from background removal and sparse structure sampling to textured mesh decoding. Mastering the from_pretrained initialization method allows you to leverage pretrained flow models for high-fidelity 3D asset generation with minimal code.
Understanding the Pipeline Architecture
The Trellis2ImageTo3DPipeline serves as the concrete implementation of the abstract Pipeline class, specifically designed for single-image 3D reconstruction. When you invoke from_pretrained, the system executes a structured loading sequence that reconstructs the entire inference graph from a JSON configuration.
Configuration Loading and Validation
The from_pretrained class method begins by locating and parsing the pipeline.json configuration file (or a user-specified config_file) from the checkpoint directory. According to the implementation in trellis2/pipelines/base.py (lines 26‑38), this JSON defines an args dictionary containing model paths, sampler hyperparameters, and conditioning settings. The method validates that all required components are specified before attempting instantiation.
Selective Model Instantiation
After loading the configuration, the pipeline filters the args['models'] entries against the subclass attribute model_names_to_load defined in trellis2/pipelines/trellis2_image_to_3d.py (lines 31‑40). Only the models listed in this attribute—typically the image condition model, sparse structure sampler, and flow transformers—are passed to trellis2.models.from_pretrained for weight restoration. This selective loading prevents unnecessary memory consumption by omitting training-only components.
Sampler Wiring and Device Placement
Once base models are loaded, from_pretrained extracts sampler configurations, background removal parameters, and conditioning model settings from the parsed args (see lines 92‑107 of trellis2_image_to_3d.py). The pipeline initializes with self._device = 'cpu' for safety, requiring an explicit call to pipeline.cuda() or pipeline.to(device) to enable GPU acceleration. This design ensures low-VRAM compatibility by keeping weights on CPU until explicitly moved.
The Inference Execution Workflow
When you call pipeline.run(image), the pipeline executes a six-stage deterministic process that transforms a PIL Image into a MeshWithVoxel object containing geometry, texture voxels, and PBR material attributes.
Image Preprocessing and Condition Extraction
First, the preprocess_image method (defined in trellis2/pipelines/trellis2_image_to_3d.py, lines 27‑62) removes the background using rembg and crops the object to a centered square region. The processed image is then forwarded through the image condition model via get_cond (lines 64‑85) to produce a latent conditioning tensor that guides subsequent generation steps.
Sparse Structure and Latent Sampling
The pipeline generates coarse geometry using sample_sparse_structure (lines 88‑35), which produces a 3D voxel mask defining the object's structural boundaries. Depending on the selected pipeline_type (512, 1024, 1024_cascade, or 1536_cascade), the system samples the shape latent (shape_slat) either directly or through a cascade of low- and high-resolution flow models (lines 46‑74). Higher resolution modes provide finer geometric detail but require more VRAM.
Texture Generation and Mesh Decoding
After establishing geometry, sample_tex_slat (lines 92‑32) conditions on both the shape latent and the input image to generate texture information. Finally, decode_latent (lines 56‑86) transforms both the shape and texture latents into a list of MeshWithVoxel objects, each containing vertices, faces, attribute volumes, and coordinate mappings ready for rendering or export.
Implementation Guide
Minimal End-to-End Inference Script
The following complete example demonstrates loading a checkpoint, running inference, rendering a preview video, and exporting to GLB format:
import os
os.environ['OPENCV_IO_ENABLE_EXR'] = '1'
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
import cv2
import imageio
from PIL import Image
import torch
from trellis2.pipelines import Trellis2ImageTo3DPipeline
from trellis2.utils import render_utils
from trellis2.renderers import EnvMap
import o_voxel
# Prepare environment map for PBR rendering (optional)
envmap = EnvMap(
torch.tensor(
cv2.cvtColor(cv2.imread('assets/hdri/forest.exr',
cv2.IMREAD_UNCHANGED), cv2.COLOR_BGR2RGB),
dtype=torch.float32, device='cuda')
)
# Load pipeline from Hugging Face Hub (or local path)
pipeline = Trellis2ImageTo3DPipeline.from_pretrained("microsoft/TRELLIS.2-4B")
pipeline.cuda() # Move all models to GPU
# Run inference on input image
image = Image.open("assets/example_image/T.png")
mesh = pipeline.run(image)[0] # Returns list; take first mesh
mesh.simplify(16_777_216) # Limit faces for nvdiffrast compatibility
# Render 360-degree preview video
video = render_utils.make_pbr_vis_frames(
render_utils.render_video(mesh, envmap=envmap)
)
imageio.mimsave("sample.mp4", video, fps=15)
# Export to GLB for Blender/Three.js import
glb = o_voxel.postprocess.to_glb(
vertices=mesh.vertices,
faces=mesh.faces,
attr_volume=mesh.attrs,
coords=mesh.coords,
attr_layout=mesh.layout,
voxel_size=mesh.voxel_size,
aabb=[[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]],
decimation_target=1_000_000,
texture_size=4096,
remesh=True,
remesh_band=1,
remesh_project=0,
verbose=True,
)
glb.export("sample.glb", extension_webp=True)
Configuring Pipeline Types and Resolution
You can control the output quality by specifying the pipeline_type parameter in the run method. The 1536_cascade option utilizes cascaded flow models for maximum detail, while 512 offers fastest inference:
# High-quality mode for detailed assets
mesh = pipeline.run(
image,
pipeline_type='1536_cascade', # Options: '512', '1024', '1024_cascade', '1536_cascade'
num_samples=1,
seed=123
)[0]
Summary
Trellis2ImageTo3DPipeline.from_pretrained()loads complete inference configurations frompipeline.jsonand initializes only the models specified inmodel_names_to_load.- The pipeline starts on CPU by default; call
.cuda()to enable GPU acceleration for all internal components. - Inference follows a strict sequence: preprocessing → conditioning → sparse structure → shape latent → texture latent → decoding.
- Generated meshes return as
MeshWithVoxelobjects, which support direct export to GLB viao_voxel.postprocess.to_glb()or rendering through utilities intrellis2/utils/render_utils.py.
Frequently Asked Questions
What checkpoint formats does from_pretrained() support?
The method accepts either a local directory path containing pipeline.json and model weights, or a Hugging Face Hub repository identifier such as "microsoft/TRELLIS.2-4B". When using a Hub ID, the pipeline automatically downloads and caches the required files to your local Hugging Face cache directory.
How do I enable low VRAM mode when using the pipeline?
The pipeline respects the low_vram flag specified in the pipeline.json configuration. Additionally, you can manage memory manually by keeping the pipeline on CPU until needed, or by processing images one at a time rather than in batches. The expandable_segments:True environment variable in PyTorch also helps prevent OOM errors during large cascade sampling operations.
What are the differences between the pipeline_type options?
The 512 mode samples at 512³ voxel resolution for fastest generation. The 1024 and 1024_cascade modes increase resolution to 1024³, with the cascade variant using separate low- and high-resolution flow models for better quality. The 1536_cascade option provides the highest fidelity at 1536³ resolution but requires significantly more GPU memory and inference time.
How do I export generated meshes to standard 3D formats?
The resulting MeshWithVoxel objects contain all necessary geometry and texture data. Use o_voxel.postprocess.to_glb() to export directly to GLB format (compatible with Blender, Three.js, and game engines), or access the raw vertices, faces, and attrs tensors for custom export pipelines to OBJ, FBX, or USD formats.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →