# How the Shape-Conditioned Texture Pipeline Encodes Mesh Geometry in TRELLIS.2

> Discover how TRELLIS.2's shape-conditioned texture pipeline encodes mesh geometry using a dual-grid voxel representation and a shape encoder to create a geometry-aware latent shape-slat.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: deep-dive
- Published: 2026-08-03

---

**The shape-conditioned texture pipeline encodes mesh geometry by normalizing the input mesh, rasterizing it into a flexible dual-grid voxel representation, and processing the resulting sparse tensor through a dedicated shape encoder to produce a geometry-aware latent called a shape-slat.**

The **shape-conditioned texture pipeline** in Microsoft's TRELLIS.2 repository enables high-fidelity texture generation by first converting raw 3D meshes into structured geometric latents. This encoding process, implemented in the `Trellis2TexturingPipeline` class, transforms arbitrary mesh topologies into a sparse voxel representation that conditions the subsequent texture synthesis models. Understanding this encoding mechanism is essential for developers extending the pipeline or optimizing memory usage for high-resolution meshes.

## The Three-Stage Geometry Encoding Process

The encoding pipeline operates through three distinct phases implemented in [`trellis2/pipelines/trellis2_texturing.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_texturing.py). Each stage progressively transforms the mesh from raw vertex data into a compact latent representation suitable for neural processing.

### Step 1: Mesh Normalization and Coordinate Alignment

The pipeline begins with the `preprocess_mesh` function (lines 6-20), which standardizes input meshes to a consistent spatial domain. This normalization ensures geometric consistency regardless of the mesh's original scale or orientation.

The function performs three critical operations:

- **Centering**: Computes the bounding box center and translates vertices to the origin
- **Uniform scaling**: Scales the mesh to fit within a -0.5 to 0.5 cube using a scale factor of `0.99999 / (vertices_max - vertices_min).max()`
- **Coordinate system alignment**: Swaps the Y and Z axes to match the renderer's coordinate conventions

```python
vertices = mesh.vertices
vertices_min = vertices.min(axis=0)
vertices_max = vertices.max(axis=0)
center = (vertices_min + vertices_max) / 2
scale = 0.99999 / (vertices_max - vertices_min).max()
vertices = (vertices - center) * scale

# Swap Y/Z to align with rendering conventions

tmp = vertices[:, 1].copy()
vertices[:, 1] = -vertices[:, 2]
vertices[:, 2] = tmp

```

### Step 2: Dual-Grid Voxelization

Following normalization, the `encode_shape_slat` method (lines 84-115) rasterizes the mesh into a **flexible dual grid** using `o_voxel.convert.mesh_to_flexible_dual_grid`. This voxelization process converts the continuous surface representation into a discrete sparse grid while preserving geometric details through dual vertices.

The voxelization produces three essential tensors:

- `voxel_indices`: Integer grid coordinates for each occupied voxel
- `dual_vertices`: Per-voxel continuous coordinates representing the dual grid vertices  
- `intersected`: A binary mask indicating which voxels intersect the mesh surface

```python
voxel_indices, dual_vertices, intersected = o_voxel.convert.mesh_to_flexible_dual_grid(
    vertices.cpu(), faces.cpu(),
    grid_size=resolution,
    aabb=[[-0.5,-0.5,-0.5],[0.5,0.5,0.5]],
    face_weight=1.0,
    boundary_weight=0.2,
    regularization_weight=1e-2,
    timing=True,
)

```

### Step 3: Sparse Tensor Construction and Shape Encoding

The final stage constructs a `SparseTensor` (defined in [`trellis2/modules/sparse/__init__.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/modules/sparse/__init__.py)) and processes it through the `shape_slat_encoder` network. This transformation occurs in lines 116-124 of the pipeline file.

The construction involves two key operations:

1. **Feature calculation**: Converting absolute dual vertices into relative voxel offsets using `dual_vertices * resolution - voxel_indices`
2. **Coordinate formatting**: Prepending a batch dimension of zeros to the voxel indices

```python
vertices = SparseTensor(
    feats=dual_vertices * resolution - voxel_indices,
    coords=torch.cat([torch.zeros_like(voxel_indices[:, 0:1]), voxel_indices], dim=-1)
).to(self.device)
intersected = vertices.replace(intersected).to(self.device)

```

The **shape-slat encoder**—a sparse-aware UNet-style network implemented in [`trellis2/models/sparse_structure_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_vae.py)—processes these tensors to generate the final geometry latent:

```python
shape_slat = self.models['shape_slat_encoder'](vertices, intersected)

```

## Practical Implementation Examples

To encode mesh geometry using the shape-conditioned texture pipeline, initialize the `Trellis2TexturingPipeline` and call the encoding methods directly.

### Encoding a Mesh to Shape-Slat

```python
import trimesh
from trellis2.pipelines.trellis2_texturing import Trellis2TexturingPipeline

# Load mesh and initialize pipeline

mesh = trimesh.load('example_mesh.obj')
pipeline = Trellis2TexturingPipeline.from_pretrained('path/to/checkpoint')
pipeline.to('cuda')

# Generate geometry latent

shape_slat = pipeline.encode_shape_slat(mesh, resolution=1024)
print(shape_slat.feats.shape)  # (num_voxels, latent_dim)

```

### End-to-End Texture Generation

For complete texture synthesis, the pipeline automatically handles encoding internally:

```python
import trimesh
from PIL import Image
from trellis2.pipelines.trellis2_texturing import Trellis2TexturingPipeline

mesh = trimesh.load('chair.obj')
image = Image.open('chair_photo.jpg')

pipeline = Trellis2TexturingPipeline.from_pretrained('path/to/checkpoint')
pipeline.to('cuda')
textured_mesh = pipeline.run(mesh, image, resolution=1024, texture_size=2048)
textured_mesh.export('textured_chair.glb')

```

## Key Source Files and Components

The encoding pipeline relies on several critical components within the TRELLIS.2 codebase:

- **[`trellis2/pipelines/trellis2_texturing.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_texturing.py)**: Contains the main `Trellis2TexturingPipeline` class with `preprocess_mesh` and `encode_shape_slat` methods
- **[`trellis2/modules/sparse/__init__.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/modules/sparse/__init__.py)**: Defines the `SparseTensor` data structure for efficient sparse grid representation
- **[`trellis2/models/sparse_structure_vae.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/models/sparse_structure_vae.py)**: Implements the shape-slat encoder network architecture
- **[`o-voxel/o_voxel/convert.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o-voxel/o_voxel/convert.py)**: Provides the `mesh_to_flexible_dual_grid` voxelization routine

## Summary

- The **shape-conditioned texture pipeline** encodes geometry through normalization, voxelization, and sparse tensor processing
- **Mesh normalization** centers, scales, and reorients vertices to a standard -0.5 to 0.5 cube with Y/Z axis swapping
- **Dual-grid voxelization** rasterizes surfaces into `voxel_indices`, `dual_vertices`, and `intersected` masks via `mesh_to_flexible_dual_grid`
- **SparseTensor construction** converts absolute coordinates to relative voxel offsets before encoding
- The **shape-slat encoder** processes sparse geometric data to produce conditioning latents for texture generation

## Frequently Asked Questions

### What is the resolution parameter in encode_shape_slat?

The `resolution` parameter controls the grid size for voxelization, determining the spatial granularity of the shape-slat. Higher resolutions (e.g., 1024) capture finer geometric details but require more memory and computation. The parameter is passed directly to `mesh_to_flexible_dual_grid` as `grid_size`.

### Why does the pipeline swap the Y and Z axes during preprocessing?

The coordinate swap aligns the input mesh with the rendering conventions used by the TRELLIS.2 renderer. This ensures that the encoded geometry matches the expected orientation when the texture generation model projects features onto the surface, preventing texture misalignment during synthesis.

### What is the difference between dual_vertices and voxel_indices?

`voxel_indices` represent discrete integer coordinates on the sparse grid, while `dual_vertices` contain continuous coordinates representing the precise intersection points within each voxel. The pipeline converts these to relative offsets (`dual_vertices * resolution - voxel_indices`) to provide sub-voxel positional information to the neural encoder.

### Can I use the shape-slat encoding without generating textures?

Yes. The `encode_shape_slat` method is publicly exposed in the `Trellis2TexturingPipeline` class, allowing you to extract geometry latents independently. This enables using the encoded representation for other downstream tasks such as shape analysis, retrieval, or as input to custom texture models.