# What Are `ShapesLatDecoder` and `TexsLatDecoder` in the TRELLIS.2 Pipeline?

> Discover how ShapesLatDecoder and TexsLatDecoder in TRELLIS.2 reconstruct 3D geometry and appearance from latent tokens, enabling efficient asset storage and high-fidelity regeneration.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: internals
- Published: 2026-08-04

---

**`ShapesLatDecoder` and `TexsLatDecoder` are specialized flow-based transformers that reconstruct 3D geometry and texture/material appearance from compact latent token sequences, enabling efficient storage and high-fidelity regeneration of 3D assets in the TRELLIS.2 system.**

The TRELLIS.2 pipeline represents 3D assets using two distinct **structural latent spaces**—one for geometry and one for appearance. These decoders form the critical bridge between compressed latent representations and renderable 3D data, operating as the inverse of their corresponding VAE encoders.

## How TRELLIS.2 Structures Its Latent Spaces

Microsoft's TRELLIS.2 employs a **disentangled latent architecture** that separates shape and texture information into independent encoding pathways. This design mirrors how 3D asset pipelines traditionally handle geometry and materials separately, but embeds both into learned latent spaces for generative modeling.

| Latent Type | Encodes | Storage | Token Dimensions |
|-------------|---------|---------|------------------|
| **Shape latent** | Coarse geometric layout (voxel occupancy, structural primitives) | `shape_latents/…/*.npz` | *N* × *D* (short token sequence) |
| **Texture latent** | Fine-grained surface appearance (BRDF parameters, albedo, normals) | `texture_latents/…/*.npz` | *M* × *D* (token sequence) |

Both latents are produced by separate **VAE encoders** during training and consumed by their dedicated decoders during inference and rendering.

## The Role of `ShapesLatDecoder` in Geometry Reconstruction

`ShapesLatDecoder` implements the **geometry decoding pathway** in [`trellis2/models/structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/structured_latent_flow.py). This transformer-based decoder learns a flow-matching transformation that maps shape latent tokens back to dense geometric representations.

The decoder handles **sparsity-aware reconstruction**—a key requirement for efficient 3D generation. Rather than predicting full dense volumes directly, it leverages the structured latent flow architecture to decode from compact tokens to either:

- Dense voxel grids representing occupancy or signed distance fields
- Sparse mesh structures that can be meshed via marching cubes

```python

# Geometry reconstruction from shape latent tokens

from trellis2.models.structured_latent_flow import ShapesLatDecoder

shape_decoder = ShapesLatDecoder(
    latent_dim=128,          # Dimension of each shape latent token

    hidden_dim=256,          # Hidden dimension for decoder transformer layers

    out_resolution=64        # Target voxel grid resolution

)

# shape_latent: (B, N, latent_dim) tensor from the shape VAE encoder

voxel_occupancy = shape_decoder(shape_latent)

# Output: (B, 1, 64, 64, 64) dense voxel grid

```

The decoder architecture incorporates **flow-matching objectives** that enable deterministic sampling from the learned geometric distribution, as implemented in the [`structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/structured_latent_flow.py) module.

## The Role of `TexsLatDecoder` in Texture Reconstruction

`TexsLatDecoder`—defined in [`trellis2/models/sparse_structure_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_flow.py)—performs the **appearance decoding** function. It transforms texture latent tokens into per-voxel or per-vertex material attributes that align spatially with the decoded geometry.

This decoder must resolve **high-frequency texture details** from relatively compact latent sequences, reconstructing:

- Albedo (base color)
- Normal maps (surface orientation)
- Roughness/metallic parameters for physically-based rendering

```python

# Texture reconstruction from texture latent tokens

from trellis2.models.sparse_structure_flow import TexsLatDecoder

tex_decoder = TexsLatDecoder(
    latent_dim=128,          # Matches shape latent dimension

    hidden_dim=256,          # Transformer hidden dimension

    out_channels=9           # RGB (3) + normal (3) + roughness/metallic (3)

)

# tex_latent: (B, M, latent_dim) tensor from the texture VAE encoder

appearance_features = tex_decoder(tex_latent)

# Output: (B, 9, 64, 64, 64) multi-channel texture field

```

The `TexsLatDecoder` operates on the **sparse structure** established by the geometry, ensuring texture predictions are conditioned on and aligned with the underlying shape.

## Pipeline Integration: Where Decoders Fit

The decoders occupy a specific position in the end-to-end TRELLIS.2 inference pipeline, as orchestrated in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py):

```

Image Input
    ↓
Image Encoder → Joint Latent (DIT transformer processing)
    ↓
Split to:                    Shape Latent Tokens    Texture Latent Tokens
                              ↓                      ↓
                              ShapesLatDecoder       TexsLatDecoder
                              ↓                      ↓
Dense Outputs:                Voxel Grid/Occupancy   Appearance Features (9-ch)
                              ↓                      ↓
                              └──→ PBRMeshRenderer ←─┘
                                    ↓
                              Final Rendered 3D Asset

```

The training logic in [`trellis2/trainers/vae/shape_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/trainers/vae/shape_vae.py) and [`trellis2/trainers/vae/pbr_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/trainers/vae/pbr_vae.py) demonstrates how these decoders are optimized:

- **Reconstruction loss** measures fidelity of decoded outputs against ground-truth geometry/textures
- **KL divergence** regularizes the latent space produced by the paired encoder
- **Flow-matching losses** stabilize the transformer-based decoding process

## Key Implementation Files

| Component | Source File | Function |
|-----------|-------------|----------|
| Shape decoder | [`trellis2/models/structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/structured_latent_flow.py) | `ShapesLatDecoder` class with flow-matching transformer |
| Texture decoder | [`trellis2/models/sparse_structure_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_flow.py) | `TexsLatDecoder` class for appearance reconstruction |
| Shape VAE training | [`trellis2/trainers/vae/shape_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/trainers/vae/shape_vae.py) | Instantiates and trains shape encoder/decoder pair |
| Texture VAE training | [`trellis2/trainers/vae/pbr_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/trainers/vae/pbr_vae.py) | Instantiates and trains texture encoder/decoder pair |
| End-to-end pipeline | [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) | Orchestrates decoder calls for inference |

## Design Rationale: Why Separate Decoders?

The decoupled decoder architecture in TRELLIS.2 enables several capabilities:

- **Independent compression rates**: Geometry and texture can use different token counts (*N* vs *M*) based on their complexity
- **Modular editing**: Shape and appearance latents can be manipulated or swapped independently
- **Efficient caching**: Pre-decoded geometry can be paired with varying textures without re-running `ShapesLatDecoder`
- **Specialized architectures**: `ShapesLatDecoder` can prioritize spatial coherence for occupancy, while `TexsLatDecoder` optimizes for high-frequency detail recovery

## Summary

- **`ShapesLatDecoder`** ([`structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/structured_latent_flow.py)) reconstructs **3D geometry** from shape latent tokens, outputting voxel grids or sparse structures via flow-matching transformers
- **`TexsLatDecoder`** ([`sparse_structure_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/sparse_structure_flow.py)) reconstructs **surface appearance** (materials, textures) from texture latent tokens, producing multi-channel feature fields aligned with geometry
- Both decoders are **VAE companions**, trained jointly with encoders via reconstruction and flow-matching objectives in `trellis2/trainers/vae/`
- The decoders enable **efficient 3D asset representation**—compact latent codes that expand to full renderable assets on demand

## Frequently Asked Questions

### What is the difference between `ShapesLatDecoder` and `TexsLatDecoder`?

`ShapesLatDecoder` focuses exclusively on geometric reconstruction—voxel occupancy, signed distance fields, or mesh structures—while `TexsLatDecoder` handles appearance attributes like albedo, normals, and material parameters. They operate on independent latent sequences but produce spatially-aligned outputs that combine during rendering.

### Where are these decoders trained in the TRELLIS.2 codebase?

Training occurs in separate VAE trainer modules: `ShapesLatDecoder` is trained in [`trellis2/trainers/vae/shape_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/trainers/vae/shape_vae.py) alongside its encoder, and `TexsLatDecoder` is trained in [`trellis2/trainers/vae/pbr_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/trainers/vae/pbr_vae.py). Both use reconstruction losses combined with flow-matching objectives specific to their output modalities.

### Can `ShapesLatDecoder` and `TexsLatDecoder` operate independently?

Yes—the decoders are designed for **modular operation**. A pre-decoded geometry can be cached and rendered with different textures by running only `TexsLatDecoder` with new texture latents. This enables applications like material swapping and efficient multi-texture rendering from shared geometry.

### What resolution do these decoders typically output?

Both decoders support configurable output resolutions set via the `out_resolution` parameter, with 64³ voxels being common for the shape decoder and matching spatial dimensions for texture features. Higher resolutions trade computational cost for finer geometric and textural detail.