How the Structured Latent (SLat) Representation Encodes Shape and Texture in TRELLIS 2

TRELLIS 2 uses a dual-path structured latent (SLat) representation that stores geometry as sparse voxel occupancy tokens (Shape SLat) and appearance as flow-generated dense features (Texture SLat), enabling high-resolution 3D generation while maintaining memory efficiency.

The microsoft/TRELLIS.2 repository implements a novel 3D generation framework that decouples shape and texture into two complementary latent spaces. This article examines how the structured latent (SLat) representation actually encodes geometric and photometric information using sparse tensor operations and conditional generative models.

Understanding the SLat Architecture

At its core, the SLat system relies on sparse tensor data structures to represent 3D content efficiently. Unlike dense voxel grids that consume memory cubically, the SLat representation scales linearly with surface area.

Sparse Tensor Backbone

Both shape and texture modalities utilize SparseTensor objects defined in trellis2/modules/sparse/basic.py. Each sparse tensor contains:

  • coords — Integer voxel coordinates stored as a (N, 3) tensor
  • feats — Feature vectors (occupancy flags for shape, latent codes for texture) stored as (N, C) tensors

This sparse representation allows TRELLIS 2 to process resolutions up to 1024³ on standard GPU hardware, as memory allocation is proportional only to active (occupied) voxels rather than the entire bounding volume.

Dual-Path Design Philosophy

The SLat system splits 3D content into two distinct families:

  1. Shape SLat — Encodes geometry through occupancy and coarse surface normals
  2. Texture SLat — Encodes appearance through dense RGB and PBR (Physically Based Rendering) values

This separation allows the pipeline to generate geometry first, then condition texture generation on the established shape, ensuring alignment between form and appearance.

How Shape Information is Encoded

Shape encoding in TRELLIS 2 revolves around SLatShape, a specialized dataset class that prepares geometric tokens for variational autoencoder (VAE) training.

Shape SLat Creation and Storage

The SLatShape class in trellis2/datasets/structured_latent_shape.py inherits from both SLat and SLatVisMixin. During initialization (lines 68–84), it loads per-object voxel files from shape_latents/<name>/<sha256>.npz archives containing pre-computed coords and feats arrays.

The collate_fn method in trellis2/datasets/structured_latent.py (lines 71–87) packs individual instance coordinates into a batched SparseTensor identified as pack['x_0']. This batched sparse tensor serves as the input to the geometric encoder.

VAE Training for Geometry

The ShapeVaeTrainer class in trellis2/trainers/vae/shape_vae.py manages the encoding pipeline:

  • Encoder: Compresses the input sparse tensor into a latent distribution z
  • Decoder: Reconstructs voxel geometry from the latent code
  • Loss Function: Combines KL-divergence regularization with binary cross-entropy occupancy loss and geometry rendering terms (mask, depth, and normal consistency)

During training (lines 80–114), the model learns to map high-resolution sparse voxel grids into compact latent representations that preserve topological and geometric details.

Shape Sampling and Decoding

At inference time, the SLatFlowModel (defined in trellis2/models/structured_latent_flow.py) generates novel shape latents through flow-based sampling. The pipeline calls sample_shape_slat() to produce a latent code that the shape decoder expands back into a full sparse voxel grid representing occupancy and surface normals.

How Texture Information is Encoded

Texture encoding follows a different paradigm, using the established shape SLat as a geometric condition for appearance generation.

Conditioning on Shape SLat

The texture generation pipeline in trellis2/pipelines/trellis2_texturing.py receives the shape SLat as a mandatory conditioning input (cond). The flow model defined in trellis2/models/structured_latent_flow.py (lines 15–79) utilizes a Modulated Sparse Transformer architecture that processes both sparse geometric tokens and dense texture tokens simultaneously.

This conditioning ensures that texture features align with geometric surfaces, preventing appearance from being generated in empty space.

Flow-Based Texture Generation

The sample_tex_slat() method (lines 224–263 in trellis2_texturing.py) executes the forward diffusion process:

  1. Accepts a CLIP image embedding or other conditioning signal
  2. Receives the shape SLat as geometric context
  3. Iterates through the flow model using configurable time steps (t)
  4. Outputs a texture latent (slat) containing compressed appearance information

Texture Decoding Pipeline

Decoding occurs in lines 281–284 of the texturing pipeline:

tex_vox = tex_slat_decoder(slat) * 0.5 + 0.5

The tex_slat_decoder maps the latent texture code into per-voxel RGB and PBR feature vectors. Unlike the sparse shape representation, texture features are densely populated across occupied voxels to ensure continuous color variation across surfaces.

End-to-End Implementation Example

The complete pipeline in trellis2/pipelines/trellis2_image_to_3d.py (lines 551–591) demonstrates the integration of both SLat modalities:


# 1. Sample shape SLat from image conditioning

shape_slat = pipeline.sample_shape_slat(
    cond_shape, 
    shape_flow_model
)

# 2. Sample texture SLat conditioned on the shape

tex_slat = pipeline.sample_tex_slat(
    cond_tex,  # CLIP image embedding

    pipeline.models['tex_slat_flow_model_1024'],
    shape_slat,
    tex_slat_sampler_params,
)

# 3. Decode both representations

shape_vox = pipeline.decode_shape_slat(shape_slat)  # Occupancy + normals

tex_vox = pipeline.decode_tex_slat(tex_slat)        # RGB + PBR

# 4. Generate final textured mesh

mesh = pipeline.decode_latent(shape_vox, tex_vox, resolution=1024)

Loading Pre-computed Shape Data

To load existing shape tokens from disk:

from trellis2.datasets import SLatShape

shape_dataset = SLatShape(
    roots='/path/to/dataset',
    resolution=64,
    pretrained_slat_dec='microsoft/TRELLIS.2-4B/ckpts/shape_dec_next_dc_f16c32_fp16',
)

# Retrieve sparse tensor representation

example = shape_dataset[0]
shape_slat = example['x_0']  # SparseTensor with coords and feats

Visualizing SLat Representations

The SLatVisMixin provides built-in visualization through the visualize_sample method:


# Returns rendered normals from four viewpoints

images = shape_dataset.visualize_sample(example)

# Shape: (4, 3, 1024, 1024)

Summary

  • The structured latent (SLat) representation in TRELLIS 2 separates 3D content into sparse geometric tokens (Shape SLat) and dense appearance features (Texture SLat).
  • Shape SLat utilizes SparseTensor objects stored in trellis2/datasets/structured_latent_shape.py, trained via ShapeVaeTrainer to encode voxel occupancy and surface normals.
  • Texture SLat depends on the shape SLat as a conditioning signal, generated through the SLatFlowModel and decoded by tex_slat_decoder in the texturing pipeline.
  • Memory efficiency is achieved through sparse tensor operations in trellis2/modules/sparse/basic.py, enabling 1024³ resolution processing.
  • End-to-end generation follows a two-stage pipeline: shape sampling → texture conditioning → joint decoding into textured meshes.

Frequently Asked Questions

What is the difference between Shape SLat and Texture SLat?

Shape SLat encodes geometric information—specifically voxel occupancy and surface normals—using sparse tensor representations managed by the SLatShape class. Texture SLat encodes appearance information (RGB and PBR values) using a dense feature representation generated by a flow model conditioned on the Shape SLat. While shape uses VAE training (ShapeVaeTrainer), texture uses flow-based diffusion (SLatFlowModel) to generate latent codes that are decoded into per-voxel colors.

Why does TRELLIS 2 use sparse tensors for the SLat representation?

Sparse tensors, implemented in trellis2/modules/sparse/basic.py, store only occupied voxel coordinates (coords) and their associated features (feats). This reduces memory complexity from cubic $O(n^3)$ to linear $O(n)$ relative to surface area, enabling the processing of high-resolution grids (up to 1024³) that would be prohibitively expensive with dense representations.

How does texture generation depend on the shape SLat?

The texture flow model defined in trellis2/models/structured_latent_flow.py requires the shape SLat as a geometric condition (cond parameter). This ensures that texture latent codes are generated only where geometric surfaces exist, aligning appearance features with the underlying shape structure during the sample_tex_slat process in the texturing pipeline.

What resolution can the SLat representation support?

According to the implementation in trellis2/pipelines/trellis2_texturing.py, the SLat system supports multiple resolutions including 512³ and 1024³, depending on the specific flow model checkpoint used (e.g., tex_slat_flow_model_1024). The sparse tensor architecture ensures that higher resolutions do not linearly increase memory usage, as only surface voxels require storage and computation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →