How the Shape-Conditioned Texture Pipeline Encodes Mesh Geometry in TRELLIS.2

The shape-conditioned texture pipeline encodes mesh geometry by normalizing the input mesh, rasterizing it into a flexible dual-grid voxel representation, and processing the resulting sparse tensor through a dedicated shape encoder to produce a geometry-aware latent called a shape-slat.

The shape-conditioned texture pipeline in Microsoft's TRELLIS.2 repository enables high-fidelity texture generation by first converting raw 3D meshes into structured geometric latents. This encoding process, implemented in the Trellis2TexturingPipeline class, transforms arbitrary mesh topologies into a sparse voxel representation that conditions the subsequent texture synthesis models. Understanding this encoding mechanism is essential for developers extending the pipeline or optimizing memory usage for high-resolution meshes.

The Three-Stage Geometry Encoding Process

The encoding pipeline operates through three distinct phases implemented in trellis2/pipelines/trellis2_texturing.py. Each stage progressively transforms the mesh from raw vertex data into a compact latent representation suitable for neural processing.

Step 1: Mesh Normalization and Coordinate Alignment

The pipeline begins with the preprocess_mesh function (lines 6-20), which standardizes input meshes to a consistent spatial domain. This normalization ensures geometric consistency regardless of the mesh's original scale or orientation.

The function performs three critical operations:

  • Centering: Computes the bounding box center and translates vertices to the origin
  • Uniform scaling: Scales the mesh to fit within a -0.5 to 0.5 cube using a scale factor of 0.99999 / (vertices_max - vertices_min).max()
  • Coordinate system alignment: Swaps the Y and Z axes to match the renderer's coordinate conventions
vertices = mesh.vertices
vertices_min = vertices.min(axis=0)
vertices_max = vertices.max(axis=0)
center = (vertices_min + vertices_max) / 2
scale = 0.99999 / (vertices_max - vertices_min).max()
vertices = (vertices - center) * scale

# Swap Y/Z to align with rendering conventions

tmp = vertices[:, 1].copy()
vertices[:, 1] = -vertices[:, 2]
vertices[:, 2] = tmp

Step 2: Dual-Grid Voxelization

Following normalization, the encode_shape_slat method (lines 84-115) rasterizes the mesh into a flexible dual grid using o_voxel.convert.mesh_to_flexible_dual_grid. This voxelization process converts the continuous surface representation into a discrete sparse grid while preserving geometric details through dual vertices.

The voxelization produces three essential tensors:

  • voxel_indices: Integer grid coordinates for each occupied voxel
  • dual_vertices: Per-voxel continuous coordinates representing the dual grid vertices
  • intersected: A binary mask indicating which voxels intersect the mesh surface
voxel_indices, dual_vertices, intersected = o_voxel.convert.mesh_to_flexible_dual_grid(
    vertices.cpu(), faces.cpu(),
    grid_size=resolution,
    aabb=[[-0.5,-0.5,-0.5],[0.5,0.5,0.5]],
    face_weight=1.0,
    boundary_weight=0.2,
    regularization_weight=1e-2,
    timing=True,
)

Step 3: Sparse Tensor Construction and Shape Encoding

The final stage constructs a SparseTensor (defined in trellis2/modules/sparse/__init__.py) and processes it through the shape_slat_encoder network. This transformation occurs in lines 116-124 of the pipeline file.

The construction involves two key operations:

  1. Feature calculation: Converting absolute dual vertices into relative voxel offsets using dual_vertices * resolution - voxel_indices
  2. Coordinate formatting: Prepending a batch dimension of zeros to the voxel indices
vertices = SparseTensor(
    feats=dual_vertices * resolution - voxel_indices,
    coords=torch.cat([torch.zeros_like(voxel_indices[:, 0:1]), voxel_indices], dim=-1)
).to(self.device)
intersected = vertices.replace(intersected).to(self.device)

The shape-slat encoder—a sparse-aware UNet-style network implemented in trellis2/models/sparse_structure_vae.py—processes these tensors to generate the final geometry latent:

shape_slat = self.models['shape_slat_encoder'](vertices, intersected)

Practical Implementation Examples

To encode mesh geometry using the shape-conditioned texture pipeline, initialize the Trellis2TexturingPipeline and call the encoding methods directly.

Encoding a Mesh to Shape-Slat

import trimesh
from trellis2.pipelines.trellis2_texturing import Trellis2TexturingPipeline

# Load mesh and initialize pipeline

mesh = trimesh.load('example_mesh.obj')
pipeline = Trellis2TexturingPipeline.from_pretrained('path/to/checkpoint')
pipeline.to('cuda')

# Generate geometry latent

shape_slat = pipeline.encode_shape_slat(mesh, resolution=1024)
print(shape_slat.feats.shape)  # (num_voxels, latent_dim)

End-to-End Texture Generation

For complete texture synthesis, the pipeline automatically handles encoding internally:

import trimesh
from PIL import Image
from trellis2.pipelines.trellis2_texturing import Trellis2TexturingPipeline

mesh = trimesh.load('chair.obj')
image = Image.open('chair_photo.jpg')

pipeline = Trellis2TexturingPipeline.from_pretrained('path/to/checkpoint')
pipeline.to('cuda')
textured_mesh = pipeline.run(mesh, image, resolution=1024, texture_size=2048)
textured_mesh.export('textured_chair.glb')

Key Source Files and Components

The encoding pipeline relies on several critical components within the TRELLIS.2 codebase:

Summary

  • The shape-conditioned texture pipeline encodes geometry through normalization, voxelization, and sparse tensor processing
  • Mesh normalization centers, scales, and reorients vertices to a standard -0.5 to 0.5 cube with Y/Z axis swapping
  • Dual-grid voxelization rasterizes surfaces into voxel_indices, dual_vertices, and intersected masks via mesh_to_flexible_dual_grid
  • SparseTensor construction converts absolute coordinates to relative voxel offsets before encoding
  • The shape-slat encoder processes sparse geometric data to produce conditioning latents for texture generation

Frequently Asked Questions

What is the resolution parameter in encode_shape_slat?

The resolution parameter controls the grid size for voxelization, determining the spatial granularity of the shape-slat. Higher resolutions (e.g., 1024) capture finer geometric details but require more memory and computation. The parameter is passed directly to mesh_to_flexible_dual_grid as grid_size.

Why does the pipeline swap the Y and Z axes during preprocessing?

The coordinate swap aligns the input mesh with the rendering conventions used by the TRELLIS.2 renderer. This ensures that the encoded geometry matches the expected orientation when the texture generation model projects features onto the surface, preventing texture misalignment during synthesis.

What is the difference between dual_vertices and voxel_indices?

voxel_indices represent discrete integer coordinates on the sparse grid, while dual_vertices contain continuous coordinates representing the precise intersection points within each voxel. The pipeline converts these to relative offsets (dual_vertices * resolution - voxel_indices) to provide sub-voxel positional information to the neural encoder.

Can I use the shape-slat encoding without generating textures?

Yes. The encode_shape_slat method is publicly exposed in the Trellis2TexturingPipeline class, allowing you to extract geometry latents independently. This enables using the encoded representation for other downstream tasks such as shape analysis, retrieval, or as input to custom texture models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →