What Are `ShapesLatDecoder` and `TexsLatDecoder` in the TRELLIS.2 Pipeline?

ShapesLatDecoder and TexsLatDecoder are specialized flow-based transformers that reconstruct 3D geometry and texture/material appearance from compact latent token sequences, enabling efficient storage and high-fidelity regeneration of 3D assets in the TRELLIS.2 system.

The TRELLIS.2 pipeline represents 3D assets using two distinct structural latent spaces—one for geometry and one for appearance. These decoders form the critical bridge between compressed latent representations and renderable 3D data, operating as the inverse of their corresponding VAE encoders.

How TRELLIS.2 Structures Its Latent Spaces

Microsoft's TRELLIS.2 employs a disentangled latent architecture that separates shape and texture information into independent encoding pathways. This design mirrors how 3D asset pipelines traditionally handle geometry and materials separately, but embeds both into learned latent spaces for generative modeling.

Latent Type Encodes Storage Token Dimensions
Shape latent Coarse geometric layout (voxel occupancy, structural primitives) shape_latents/…/*.npz N × D (short token sequence)
Texture latent Fine-grained surface appearance (BRDF parameters, albedo, normals) texture_latents/…/*.npz M × D (token sequence)

Both latents are produced by separate VAE encoders during training and consumed by their dedicated decoders during inference and rendering.

The Role of ShapesLatDecoder in Geometry Reconstruction

ShapesLatDecoder implements the geometry decoding pathway in trellis2/models/structured_latent_flow.py. This transformer-based decoder learns a flow-matching transformation that maps shape latent tokens back to dense geometric representations.

The decoder handles sparsity-aware reconstruction—a key requirement for efficient 3D generation. Rather than predicting full dense volumes directly, it leverages the structured latent flow architecture to decode from compact tokens to either:

  • Dense voxel grids representing occupancy or signed distance fields
  • Sparse mesh structures that can be meshed via marching cubes

# Geometry reconstruction from shape latent tokens

from trellis2.models.structured_latent_flow import ShapesLatDecoder

shape_decoder = ShapesLatDecoder(
    latent_dim=128,          # Dimension of each shape latent token

    hidden_dim=256,          # Hidden dimension for decoder transformer layers

    out_resolution=64        # Target voxel grid resolution

)

# shape_latent: (B, N, latent_dim) tensor from the shape VAE encoder

voxel_occupancy = shape_decoder(shape_latent)

# Output: (B, 1, 64, 64, 64) dense voxel grid

The decoder architecture incorporates flow-matching objectives that enable deterministic sampling from the learned geometric distribution, as implemented in the structured_latent_flow.py module.

The Role of TexsLatDecoder in Texture Reconstruction

TexsLatDecoder—defined in trellis2/models/sparse_structure_flow.py—performs the appearance decoding function. It transforms texture latent tokens into per-voxel or per-vertex material attributes that align spatially with the decoded geometry.

This decoder must resolve high-frequency texture details from relatively compact latent sequences, reconstructing:

  • Albedo (base color)
  • Normal maps (surface orientation)
  • Roughness/metallic parameters for physically-based rendering

# Texture reconstruction from texture latent tokens

from trellis2.models.sparse_structure_flow import TexsLatDecoder

tex_decoder = TexsLatDecoder(
    latent_dim=128,          # Matches shape latent dimension

    hidden_dim=256,          # Transformer hidden dimension

    out_channels=9           # RGB (3) + normal (3) + roughness/metallic (3)

)

# tex_latent: (B, M, latent_dim) tensor from the texture VAE encoder

appearance_features = tex_decoder(tex_latent)

# Output: (B, 9, 64, 64, 64) multi-channel texture field

The TexsLatDecoder operates on the sparse structure established by the geometry, ensuring texture predictions are conditioned on and aligned with the underlying shape.

Pipeline Integration: Where Decoders Fit

The decoders occupy a specific position in the end-to-end TRELLIS.2 inference pipeline, as orchestrated in trellis2/pipelines/trellis2_image_to_3d.py:


Image Input
    ↓
Image Encoder → Joint Latent (DIT transformer processing)
    ↓
Split to:                    Shape Latent Tokens    Texture Latent Tokens
                              ↓                      ↓
                              ShapesLatDecoder       TexsLatDecoder
                              ↓                      ↓
Dense Outputs:                Voxel Grid/Occupancy   Appearance Features (9-ch)
                              ↓                      ↓
                              └──→ PBRMeshRenderer ←─┘
                                    ↓
                              Final Rendered 3D Asset

The training logic in trellis2/trainers/vae/shape_vae.py and trellis2/trainers/vae/pbr_vae.py demonstrates how these decoders are optimized:

  • Reconstruction loss measures fidelity of decoded outputs against ground-truth geometry/textures
  • KL divergence regularizes the latent space produced by the paired encoder
  • Flow-matching losses stabilize the transformer-based decoding process

Key Implementation Files

Component Source File Function
Shape decoder trellis2/models/structured_latent_flow.py ShapesLatDecoder class with flow-matching transformer
Texture decoder trellis2/models/sparse_structure_flow.py TexsLatDecoder class for appearance reconstruction
Shape VAE training trellis2/trainers/vae/shape_vae.py Instantiates and trains shape encoder/decoder pair
Texture VAE training trellis2/trainers/vae/pbr_vae.py Instantiates and trains texture encoder/decoder pair
End-to-end pipeline trellis2/pipelines/trellis2_image_to_3d.py Orchestrates decoder calls for inference

Design Rationale: Why Separate Decoders?

The decoupled decoder architecture in TRELLIS.2 enables several capabilities:

  • Independent compression rates: Geometry and texture can use different token counts (N vs M) based on their complexity
  • Modular editing: Shape and appearance latents can be manipulated or swapped independently
  • Efficient caching: Pre-decoded geometry can be paired with varying textures without re-running ShapesLatDecoder
  • Specialized architectures: ShapesLatDecoder can prioritize spatial coherence for occupancy, while TexsLatDecoder optimizes for high-frequency detail recovery

Summary

  • ShapesLatDecoder (structured_latent_flow.py) reconstructs 3D geometry from shape latent tokens, outputting voxel grids or sparse structures via flow-matching transformers
  • TexsLatDecoder (sparse_structure_flow.py) reconstructs surface appearance (materials, textures) from texture latent tokens, producing multi-channel feature fields aligned with geometry
  • Both decoders are VAE companions, trained jointly with encoders via reconstruction and flow-matching objectives in trellis2/trainers/vae/
  • The decoders enable efficient 3D asset representation—compact latent codes that expand to full renderable assets on demand

Frequently Asked Questions

What is the difference between ShapesLatDecoder and TexsLatDecoder?

ShapesLatDecoder focuses exclusively on geometric reconstruction—voxel occupancy, signed distance fields, or mesh structures—while TexsLatDecoder handles appearance attributes like albedo, normals, and material parameters. They operate on independent latent sequences but produce spatially-aligned outputs that combine during rendering.

Where are these decoders trained in the TRELLIS.2 codebase?

Training occurs in separate VAE trainer modules: ShapesLatDecoder is trained in trellis2/trainers/vae/shape_vae.py alongside its encoder, and TexsLatDecoder is trained in trellis2/trainers/vae/pbr_vae.py. Both use reconstruction losses combined with flow-matching objectives specific to their output modalities.

Can ShapesLatDecoder and TexsLatDecoder operate independently?

Yes—the decoders are designed for modular operation. A pre-decoded geometry can be cached and rendered with different textures by running only TexsLatDecoder with new texture latents. This enables applications like material swapping and efficient multi-texture rendering from shared geometry.

What resolution do these decoders typically output?

Both decoders support configurable output resolutions set via the out_resolution parameter, with 64³ voxels being common for the shape decoder and matching spatial dimensions for texture features. Higher resolutions trade computational cost for finer geometric and textural detail.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →