# Understanding shape_slat_decoder and tex_slat_decoder in the TRELLIS.2 Generation Pipeline

> Discover how shape_slat_decoder and tex_slat_decoder in TRELLIS.2 reconstruct voxel grids and synthesize texture maps from SLAT sequences using the SLatFlowModel.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: internals
- Published: 2026-08-04

---

**The `shape_slat_decoder` and `tex_slat_decoder` in microsoft/TRELLIS.2 are implemented by the `SLatFlowModel` class in [`trellis2/models/structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/structured_latent_flow.py), where the shape decoder reconstructs sparse voxel grids and the texture decoder synthesizes dense RGB texture maps from Structured Latent Token (SLAT) sequences.**

In the microsoft/TRELLIS.2 open-source repository, high-quality 3D asset generation relies on compressing geometry and appearance into SLAT (Structured Latent Token) representations. These decoders serve as the critical final translation layers that convert compressed latent token streams into concrete 3D structures and surface textures, bridging the gap between abstract latent space and renderable output.

## How shape_slat_decoder and tex_slat_decoder Transform Latent Tokens

Both decoders share a unified architecture but target different output modalities. Rather than separate classes, they represent configured instances of the same flow-based model operating on distinct SLAT sequences with modality-specific parameters.

### The Core Implementation: SLatFlowModel

According to the source code in [`trellis2/models/structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/structured_latent_flow.py), both decoders are implemented by the **`SLatFlowModel`** class. This model processes input latent sequences through a stack of **Modulated Sparse Transformer Cross-Blocks** before reconstructing the final representation.

The forward pass, located at lines 71-99 of [`structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/structured_latent_flow.py), accepts latent tokens `z`, timestep `t`, and optional conditioning `cond`. The method signature follows:

```python
def forward(self, z, t, cond=None):
    # Implementation spans lines 71-99

    ...

```

This architecture uses sparse tensor operations for memory-efficient processing of 3D volumetric data.

### shape_slat_decoder: Sparse Voxel Grid Reconstruction

The shape decoder configures `SLatFlowModel` with parameters optimized for geometric reconstruction. It typically operates at a **resolution of 64³** and outputs a `SparseTensor` containing feature channels representing the object's volumetric structure.

When processing shape SLAT tokens, the decoder expands the compressed latent sequence back into spatial coordinates, filling voxel grids that can be subsequently converted to meshes via neural rendering techniques. The output sparse tensor preserves memory efficiency by only storing occupied voxels and their associated features.

### tex_slat_decoder: Dense Texture Map Generation

For texture synthesis, the same `SLatFlowModel` class is instantiated with **higher resolution settings (typically 256)** and **`out_channels=3`** to produce RGB values. Unlike the shape decoder's sparse output, the texture decoder generates a dense tensor representing a UV texture map or per-voxel color features.

This decoder processes texture-specific SLAT tokens to restore high-frequency surface details, outputting a dense tensor of shape `[B, 3, H, W]` suitable for direct application to the geometry produced by the shape decoder.

## Integration in the TRELLIS.2 Generation Pipeline

The decoders operate as the final stage after latent diffusion or flow-based generation completes. According to the pipeline implementation referenced in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), the generation process follows this sequence:

1. **Latent Generation**: A diffusion model produces separate SLAT token sequences for shape and texture, configured via [`configs/gen/slat_flow_imgshape2tex_dit_1_3B_512_bf16.json`](https://github.com/microsoft/TRELLIS.2/blob/main/configs/gen/slat_flow_imgshape2tex_dit_1_3B_512_bf16.json).
2. **Shape Decoding**: The `shape_slat_decoder` expands shape tokens into a sparse voxel grid.
3. **Texture Decoding**: The `tex_slat_decoder` generates the corresponding dense texture map.
4. **Rendering**: The decoded geometry and texture combine in the rendering module to produce the final 3D asset.

This modular design allows the same underlying `SLatFlowModel` architecture to handle both geometric and appearance decoding by varying initialization parameters.

## Code Implementation Examples

The following examples demonstrate how to instantiate both decoders based on the actual implementation in [`trellis2/models/structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/structured_latent_flow.py).

### Instantiating the Shape Decoder

```python
import torch
from trellis2.models.structured_latent_flow import SLatFlowModel
from trellis2.modules import sparse as sp

# Configure shape_slat_decoder for 64^3 voxel output

shape_slat_decoder = SLatFlowModel(
    resolution=64,
    in_channels=16,          # Latent token dimension

    model_channels=128,
    cond_channels=0,
    out_channels=sp.SparseLinear.out_features,
    num_blocks=12,
    dtype='float32',
)

# z_shape: [B, T, C] tensor from diffusion model

# t: diffusion timestep

decoded_shape = shape_slat_decoder(z_shape, t, cond=None)

# Returns: SparseTensor ready for mesh extraction

```

### Configuring the Texture Decoder

```python
import torch
from trellis2.models.structured_latent_flow import SLatFlowModel

# Configure tex_slat_decoder for 256x256 RGB texture

tex_slat_decoder = SLatFlowModel(
    resolution=256,        # Higher resolution for texture detail

    in_channels=16,
    model_channels=256,
    cond_channels=0,
    out_channels=3,      # RGB channels

    num_blocks=12,
    dtype='float32',
)

# z_tex: texture SLAT tokens [B, T, C]

decoded_texture = tex_slat_decoder(z_tex, t, cond=None)

# Returns: Dense tensor [B, 3, 256, 256] representing UV map

```

Both instances use identical class definitions but target fundamentally different representations—sparse volumetric data versus dense surface appearance.

## Summary

- **Unified Architecture**: Both `shape_slat_decoder` and `tex_slat_decoder` are implemented by the `SLatFlowModel` class in [`trellis2/models/structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/structured_latent_flow.py).
- **Sparse vs. Dense**: The shape decoder generates sparse voxel grids at 64³ resolution, while the texture decoder produces dense RGB maps at 256² resolution.
- **Transformer Core**: Both utilize Modulated Sparse Transformer Cross-Blocks to process SLAT token sequences as implemented in lines 71-99 of the source file.
- **Pipeline Position**: Located at the end of the generation pipeline ([`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py)), these decoders transform latent tokens into renderable 3D assets.

## Frequently Asked Questions

### What is the primary difference between shape_slat_decoder and tex_slat_decoder?

The primary difference lies in output modalities and spatial configurations. While both leverage the same `SLatFlowModel` implementation, `shape_slat_decoder` outputs sparse tensors representing 3D voxel grids at 64³ resolution, whereas `tex_slat_decoder` outputs dense RGB tensors representing 2D texture maps at 256² resolution. The distinction is controlled via the `resolution` and `out_channels` initialization parameters.

### Which Python class implements these decoders in the TRELLIS.2 codebase?

Both decoders are implemented by the **`SLatFlowModel`** class located in [`trellis2/models/structured_latent_flow.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/structured_latent_flow.py). This class handles the flow-based transformation of SLAT tokens through modulated sparse transformer blocks, with specific decoding behavior determined by configuration values passed during instantiation.

### How do the decoders fit into the end-to-end generation pipeline?

According to [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), the decoders serve as the final reconstruction stage. After a diffusion model generates SLAT token sequences (configured via files like [`configs/gen/slat_flow_imgshape2tex_dit_1_3B_512_bf16.json`](https://github.com/microsoft/TRELLIS.2/blob/main/configs/gen/slat_flow_imgshape2tex_dit_1_3B_512_bf16.json)), the `shape_slat_decoder` converts shape tokens into voxel geometry while the `tex_slat_decoder` generates corresponding surface textures. These outputs are then fused for final rendering.

### Can the same decoder instance handle both shape and texture generation?

No, separate instances are required because architectural hyperparameters differ significantly between modalities. The shape decoder requires `out_channels` matching sparse feature dimensions and cubic resolutions (64), while the texture decoder requires `out_channels=3` for RGB and higher planar resolutions (256). Each modality requires independent instantiation of `SLatFlowModel` with specific configuration dictionaries.