Sparse‑Structure SLat vs. Shape SLat in TRELLIS.2: Key Differences Explained
Sparse‑structure SLat stores geometry as a compact sparse voxel grid, while shape SLat encodes surface‑level geometry as a tokenized latent decoded into high‑fidelity meshes.
TRELLIS.2, Microsoft's open-source 3D generative AI framework, provides two distinct latent representations for 3D scenes. Understanding the difference between sparse‑structure SLat and shape SLat is essential for selecting the right pipeline—whether you need efficient voxel-based processing or production-quality mesh generation.
What Is Sparse‑Structure SLat?
Sparse‑structure SLat represents 3D geometry as a field-free, sparse voxel grid. Only occupied voxels are stored, making memory usage scale with actual geometry rather than resolution.
Core Characteristics
- Storage format: Highly compact 3D tensor with raw occupancy and optional color data
- Dataset class:
SparseStructureLatentdefined in [trellis2/datasets/sparse_structure_latent.py](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/datasets/sparse_structure_latent.py) - Decoder:
ss_dec(sparse-structure decoder) returns dense voxel fields - Renderer:
VoxelRendererfrom [trellis2/renderers/voxel_renderer.py](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/renderers/voxel_renderer.py) - Flow model:
SparseStructureFlowModelin [trellis2/models/sparse_structure_flow.py](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_flow.py)
When to Use Sparse‑Structure SLat
- Quick voxel-based inspection or visualization
- Downstream tasks requiring raw density information
- Voxel-based collision detection or spatial queries
- Conversion to other voxel formats (e.g., for simulation pipelines)
What Is Shape SLat?
Shape SLat represents object geometry as a structured sequence of tokens decoded into a complete mesh with vertices, faces, and optional texture attributes.
Core Characteristics
- Storage format: Tokenized latent stream ("SLat" tokens) describing surface geometry
- Dataset class:
SLatShapedefined in [trellis2/datasets/structured_latent_shape.py](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/datasets/structured_latent_shape.py) - Decoder:
slat_dec(shape decoder) returns mesh objects - Renderer:
MeshRendereraccessed viaget_renderer(reps[0]) - Flow model:
SLatFlowModelin [trellis2/models/structured_latent_flow.py](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/structured_latent_flow.py)
When to Use Shape SLat
- High-fidelity mesh generation for graphics pipelines
- PBR (Physically Based Rendering) material workflows
- Standard 3D asset production and game engine integration
- Applications requiring surface normals and UV coordinates
Key Differences at a Glance
| Aspect | Sparse‑Structure SLat | Shape SLat |
|---|---|---|
| Core representation | Sparse voxel grid | Tokenized surface description |
| Memory efficiency | Scales with geometry, not resolution | Fixed-size token sequence |
| Decoder output | Dense voxel field (C×H×W×D) |
Mesh (vertices, faces, textures) |
| Rendering approach | Direct voxel color visualization | Full mesh shading with normals |
| Typical checkpoint | ss_latent / ss_dec_conv3d_16l8_fp16 |
shape_latent / shape_dec_next_dc_f16c32_fp16 |
| Primary use case | Voxel-centric processing, density analysis | Production mesh generation |
Code Examples: Loading and Visualizing Each SLat Type
Loading Shape SLat
import torch
from trellis2.datasets.structured_latent_shape import SLatShape
# Initialize dataset
shape_ds = SLatShape(
roots="path/to/shape_dataset",
resolution=512,
pretrained_slat_dec="microsoft/TRELLIS.2-4B/ckpts/shape_dec_next_dc_f16c32_fp16",
)
# Load and decode
sample = shape_ds[0]
meshes = shape_ds.decode_latent(sample["x_0"].cuda())
# Render four-view visualization
images = shape_ds.visualize_sample(sample) # (4, 3, 1024, 1024)
Key implementation details:
decode_latentinstructured_latent.pyloadsslat_decvisualize_sampleinstructured_latent_shape.pyconfigures camera rig and mesh renderer
Loading Sparse‑Structure SLat
import torch
from trellis2.datasets.sparse_structure_latent import SparseStructureLatent
# Initialize dataset
ss_ds = SparseStructureLatent(
roots="path/to/sparse_structure_dataset",
pretrained_ss_dec="JeffreyXiang/TRELLIS-image-large/ckpts/ss_dec_conv3d_16l8_fp16",
)
# Load and decode
sample = ss_ds[0]
voxel_field = ss_ds.decode_latent(sample["x_0"].cuda()) # C×H×W×D
# Render voxel visualization
images = ss_ds.visualize_sample(sample) # (4, 3, 1024, 1024)
Key implementation details:
decode_latentinvokesself.ss_decfor sparse-to-dense conversionVoxelRendererprocesses non-zero voxels into colored point renderings
How TRELLIS.2 Uses Both Representations
TRELLIS.2 supports both SLat types because they serve complementary generative pipelines. Sparse‑structure SLat enables memory-efficient, resolution-agnostic previews ideal for voxel-based learning and spatial analysis. Shape SLat provides the canonical mesh output expected by modern graphics pipelines, with full support for PBR materials and surface-based processing.
Both representations are trained with flow-based diffusion models, but operate on fundamentally different data modalities:
- Sparse‑structure flow: Operates on sparse tensors for efficient voxel processing
- Structured latent flow: Operates on dense token sequences for mesh generation
Summary
- Sparse‑structure SLat compresses 3D scenes into sparse voxel grids decoded by
ss_decand rendered withVoxelRenderer - Shape SLat encodes surface geometry as token streams decoded by
slat_decinto meshes for standard rendering - Both use dedicated flow models—
SparseStructureFlowModelandSLatFlowModel—but target different downstream workflows - Choose sparse‑structure for voxel-based analysis; choose shape SLat for production mesh generation
Frequently Asked Questions
Can I convert sparse‑structure SLat to shape SLat?
Direct conversion is not implemented in the base TRELLIS.2 repository. The two representations are learned separately with different flow models and decoders. For mesh output from voxel data, you would need to run surface reconstruction (e.g., marching cubes) on the decoded voxel field, or regenerate using the shape SLat pipeline from the same image or text conditioning.
Which SLat type is used in the official image-to-3D demos?
The majority of TRELLIS.2 image-to-3D generation demos use shape SLat as the primary output representation. The mesh-based output integrates directly with standard graphics pipelines and provides higher visual fidelity with PBR shading, making it the default choice for showcase applications.
How do memory requirements compare between the two SLat types?
Sparse‑structure SLat scales memory usage with the amount of occupied geometry rather than the full spatial resolution, making it efficient for sparse scenes. Shape SLat uses fixed-size token sequences independent of geometric complexity. For dense, highly detailed objects, shape SLat may be more compact; for sparse, large-scale scenes, sparse‑structure SLat typically wins.
Are the decoders interchangeable between SLat types?
No. The decoders are specialized to their respective representations. ss_dec (sparse-structure decoder) outputs voxel fields and cannot produce meshes. slat_dec (shape decoder) outputs mesh structures and does not handle voxel data. Attempting to use either decoder on the wrong latent type will raise dimension or modality mismatch errors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →