SparseStructureFlowModel vs ShapesLatFlowModel vs TexSLatFlowModel: TRELLIS.2 Architecture Guide
SparseStructureFlowModel is the generic conditional diffusion transformer backbone for sparse 3-D tensors, while ShapesLatFlowModel (via SLatFlowModel) and TexSLatFlowModel (via ElasticSLatFlowModel) are thin specialization wrappers that configure it for geometry (single-channel occupancy) and texture (three-channel RGB) latents respectively.
The TRELLIS.2 repository implements a unified latent flow architecture for 3D asset generation. Understanding the differences between sparse_structure_flow_model, shape_slat_flow_model, and tex_slat_flow_model is critical for customizing the diffusion pipeline or extending the framework. While these components share identical transformer internals, they diverge in latent tensor shapes, conditioning mechanisms, and default hyperparameters baked into their configuration files.
Architectural Hierarchy: Generic Backbone vs. Specialization Wrappers
The codebase employs a hierarchical design where SparseStructureFlowModel serves as the foundational implementation. Located in trellis2/models/sparse_structure_flow.py (lines 56-248), this class implements the full conditional diffusion transformer logic including timestep embedding, transformer blocks, and sparse tensor handling.
Both ShapesLatFlowModel and TexSLatFlowModel are not re-implementations. Instead, they are thin wrapper classes defined in trellis2/models/structured_latent_flow.py that inherit from SparseStructureFlowModel and hardcode specific arguments for their respective domains. This design ensures that weight initialization routines, checkpointing behavior, and forward pass logic remain identical across all three variants.
Key Technical Differences
Input Channels and Latent Resolution
The primary distinction lies in the shape of the latent tensors each model consumes:
- ShapesLatFlowModel: Processes 1-channel occupancy grids representing geometry, typically at
resolution=64. The wrapper setsin_channels=1andout_channels=1viaSLatFlowModel. - TexSLatFlowModel: Handles 3-channel RGB texture grids, usually at
resolution=128. TheElasticSLatFlowModelwrapper configuresin_channels=3andout_channels=3. - SparseStructureFlowModel: Accepts any channel count defined by the
in_channelsconstructor argument, making it suitable for custom latent spaces.
Conditioning Channels and Context
Each flow model expects different conditioning vectors:
- Shape flow: Uses
cond_channels=256by default, conditioning on low-dimensional shape latents or geometric context. - Texture flow: Employs
cond_channels=512, typically conditioning on style vectors or appearance codes. - Generic model: The
cond_channelsparameter is configurable at instantiation.
Positional Embedding Strategies (pe_mode)
The pe_mode parameter controls how positional information is injected:
- ShapesLatFlowModel: Defaults to
ape(absolute positional embedding) as specified inshape_vae_next_dc_f16c32_fp16.json. - TexSLatFlowModel: Defaults to
rope(rotary positional embedding) for finer spatial control in texture synthesis, configured intex_vae_next_dc_f16c32_fp16.json. - SparseStructureFlowModel: Supports both modes via the
pe_modeargument but requires explicit selection.
Hyperparameter Defaults and Initialization
The configuration files reveal significant architectural scaling differences:
- Shape configs: Set
model_channels=256,num_blocks=12, andinitialization=vanilla. - Texture configs: Use
model_channels=512,num_blocks=16, and ofteninitialization=scaledto handle larger channel widths. - Training entry points: The
train.pyscript loads these defaults when invoked with--model shapeor--model tex, instantiating the appropriate wrapper class.
Implementation in Source Files
Generic Backbone Implementation
In trellis2/models/sparse_structure_flow.py, the SparseStructureFlowModel class implements the complete transformer pipeline:
class SparseStructureFlowModel(nn.Module):
def __init__(self, resolution, in_channels, cond_channels,
model_channels, num_blocks, pe_mode="ape", ...):
# Full implementation of diffusion transformer blocks
# Handles sparse tensor operations and modulation
Shape Flow Wrapper
The SLatFlowModel class in trellis2/models/structured_latent_flow.py hardcodes geometry-specific defaults:
class SLatFlowModel(SparseStructureFlowModel):
def __init__(self, **kwargs):
super().__init__(
resolution=64,
in_channels=1, # Single-channel occupancy grid
cond_channels=256, # Shape conditioning dimension
out_channels=1,
**kwargs,
)
Texture Flow Wrapper with Elastic Mixin
The ElasticSLatFlowModel extends the base class for appearance modeling:
class ElasticSLatFlowModel(SparseStructureFlowModel):
def __init__(self, **kwargs):
super().__init__(
resolution=128,
in_channels=3, # RGB texture channels
cond_channels=512,
out_channels=3,
**kwargs,
)
# Additional elasticity logic from sparse_elastic_mixin.py
Configuration Files and Model Selection
The repository separates architecture definition from hyperparameter specification:
- Shape flow configuration:
configs/gen/slat_flow_imgshape2tex_dit_1_3B_512_bf16.jsonandshape_vae_next_dc_f16c32_fp16.jsondefine the 64-resolution, 256-channel setup. - Texture flow configuration:
configs/gen/slat_flow_imgshape2tex_dit_1_3B_512_bf16_ft1024.jsonandtex_vae_next_dc_f16c32_fp16.jsonspecify the 128-resolution, 512-channel architecture with rotary embeddings.
When executing train.py, the --model flag determines which wrapper class and config file are loaded, ensuring the correct channel dimensions and conditioning setup for the target task.
Practical Usage Examples
Instantiating a ShapesLatFlowModel
from trellis2.models.structured_latent_flow import SLatFlowModel
# Configuration loaded from shape_vae_next_dc_f16c32_fp16.json
shape_flow = SLatFlowModel(
resolution=64,
model_channels=256,
num_blocks=12,
num_head_channels=64,
pe_mode="ape", # Absolute positional embedding
dtype="float32",
)
Instantiating a TexSLatFlowModel
from trellis2.models.structured_latent_flow import ElasticSLatFlowModel
# Configuration loaded from tex_vae_next_dc_f16c32_fp16.json
tex_flow = ElasticSLatFlowModel(
resolution=128,
model_channels=512,
num_blocks=16,
num_head_channels=64,
pe_mode="rope", # Rotary positional embedding
dtype="float16",
share_mod=True, # Enable shared modulation for texture
)
Running a Forward Pass
Both models share the identical API defined in the generic backbone:
import torch
# Dummy sparse tensor data: (B, C, D, H, W)
x = torch.randn(4, shape_flow.in_channels, shape_flow.resolution,
shape_flow.resolution, shape_flow.resolution)
t = torch.randint(0, 1000, (4,)) # Diffusion timestep
cond = torch.randn(4, shape_flow.cond_channels)
# Forward pass applies the same transformer logic regardless of specialization
out = shape_flow(x, t, cond) # Output shape: (B, C, D, H, W)
Summary
- SparseStructureFlowModel is the complete conditional diffusion transformer implementation in
trellis2/models/sparse_structure_flow.py, handling all sparse tensor operations and modulation logic. - ShapesLatFlowModel (
SLatFlowModel) is a configuration wrapper that instantiates the generic model within_channels=1,resolution=64, and absolute positional embeddings for geometry generation. - TexSLatFlowModel (
ElasticSLatFlowModel) configures the same backbone within_channels=3,resolution=128, rotary embeddings, and elastic mixin support for texture synthesis. - All three share identical forward pass implementations, checkpointing behavior, and transformer block architectures, differing only in constructor arguments and JSON configuration files.
Frequently Asked Questions
Can I instantiate SparseStructureFlowModel directly instead of using the wrappers?
Yes. You can import and instantiate SparseStructureFlowModel directly from trellis2.models.sparse_structure_flow if you need custom channel counts, resolutions, or conditioning dimensions that differ from the standard shape or texture configurations. The wrappers simply provide convenient default presets.
Why does the texture model use rotary positional embeddings (rope) while the shape model uses absolute (ape)?
According to the TRELLIS.2 configuration files, texture generation benefits from rope (rotary positional embedding) because it provides finer spatial control and better extrapolation for high-resolution RGB grids (128³). The shape model uses ape (absolute positional embedding) as the default because occupancy grids at lower resolution (64³) do not require the same spatial granularity, though both modes are technically supported in the generic backbone.
What is the elastic mixin referenced in TexSLatFlowModel?
The ElasticSLatFlowModel incorporates logic from trellis2/modules/sparse/elastic_mixin.py (or sparse_elastic_mixin.py) that adds elasticity behavior to the texture flow. This mixin enables adaptive handling of texture-specific modulation and sparse tensor operations that differ from rigid occupancy grids, though the core transformer blocks remain unchanged from the generic implementation.
How do I switch between shape and texture flow models during training?
The train.py script controls model selection via the --model argument. Pass --model shape to instantiate SLatFlowModel with the shape-specific JSON configs, or --model tex to instantiate ElasticSLatFlowModel with texture-specific hyperparameters. The script automatically loads the appropriate in_channels, resolution, and pe_mode settings from the corresponding configuration files in configs/gen/.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →