How to Configure pipeline_type Options for Different Resolutions in TRELLIS.2

The Trellis2ImageTo3DPipeline class accepts a pipeline_type parameter that selects from four resolution presets—SI2 (512³), IO24 (1024³), IO24-cascade (1024³ two-stage), and IS36-cascade (1536³)—each mapping to specific voxel densities and latent grid sizes defined in the source code.

The microsoft/TRELLIS.2 3D reconstruction framework exposes resolution control through the pipeline_type configuration option. This parameter routes to internal identifiers that determine the spatial resolution of the generated mesh and the underlying structured latent field, allowing you to trade inference speed for geometric fidelity.

Resolution Presets and Technical Mapping

TRELLIS.2 implements four built-in pipeline_type presets that control the internal resolution of latent representations. The mapping between human-readable aliases and internal identifiers is handled in trellis2/pipelines/trellis2_image_to_3d.py.

SI2 (512³) – Low-Resolution Fast Inference

The SI2 preset corresponds to a 512 × 512 × 512 voxel grid. It uses a latent grid size (ss_res) of 32, making it the fastest option for prototyping or low-detail reconstructions. Internally, this maps to the identifier '512'.

IO24 (1024³) – Standard High-Quality Mode

The IO24 preset generates a 1024 × 1024 × 1024 voxel mesh using a latent grid size (ss_res) of 64. This is the standard high-quality mode for balanced inference speed and detail. It maps to the internal identifier '1024'.

IO24-cascade (1024³) – Two-Stage Refinement

The IO24-cascade preset also targets 1024³ voxels but employs a two-stage sampling strategy. It first generates a coarse prediction, then refines it at full resolution. Despite the higher output resolution, it uses a latent grid size (ss_res) of 32 for efficiency. This maps to '1024_cascade'.

IS36-cascade (1536³) – Maximum Resolution Cascade

The IS36-cascade preset delivers the highest quality at 1536 × 1536 × 1536 voxels. Like the 1024_cascade mode, it uses cascaded sampling but operates at the maximum supported resolution. It uses a latent grid size (ss_res) of 32 and maps to '1536_cascade'.

Source Code Implementation

The resolution logic is implemented in the pipeline's run method and configuration loading system.

Resolution Mapping in trellis2_image_to_3d.py

The Trellis2ImageTo3DPipeline class in trellis2/pipelines/trellis2_image_to_3d.py defines the relationship between pipeline_type strings and their underlying ss_res (structured latent resolution) values:

  • '512' → ss_res=32
  • '1024' → ss_res=64
  • '1024_cascade' → ss_res=32
  • '1536_cascade' → ss_res=32

The ss_res parameter determines the resolution of the structured-latent field that the network predicts before converting to the final mesh.

Base Pipeline Loading

The trellis2/pipelines/base.py file provides the generic Pipeline base class that handles loading of pretrained configurations. When you call from_pretrained(), the base class loads the model weights and prepares the pipeline to accept the pipeline_type argument at inference time.

Configuring pipeline_type in Python

You configure the resolution by passing the pipeline_type argument to the run() method after loading your pretrained model.

Basic Configuration Syntax

Load the model and select a specific resolution preset:

from trellis2.pipelines import Trellis2ImageTo3DPipeline

# Load the pretrained model (e.g., microsoft/TRELLIS.2-4B)

pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    "microsoft/TRELLIS.2-4B"
)

pipeline.cuda()

# Configure for standard 1024³ resolution

mesh = pipeline.run(image, pipeline_type="IO24")[0]

Enabling Cascade Sampling

For higher quality with the cascade modes, specify the cascade preset explicitly:


# Generate using two-stage 1024³ cascade (coarse-to-fine)

mesh = pipeline.run(image, pipeline_type="IO24-cascade")[0]

# Or use the maximum 1536³ resolution

mesh_high_res = pipeline.run(image, pipeline_type="IS36-cascade")[0]

Benchmarking Multiple Resolutions

You can switch resolutions at runtime without reloading the model:

resolutions = ["SI2", "IO24", "IO24-cascade", "IS36-cascade"]

for preset in resolutions:
    mesh = pipeline.run(image, pipeline_type=preset)[0]
    vertex_count = mesh.vertices.shape[0]
    print(f"{preset}: {vertex_count} vertices")

Key Implementation Files

Understanding these files helps with advanced configuration:

  • trellis2/pipelines/trellis2_image_to_3d.py: Contains the run method implementation, the pipeline_type resolution logic, and the ss_res mapping dictionary.
  • trellis2/pipelines/base.py: Provides the base Pipeline class responsible for loading pretrained configs and checkpoint management.
  • example.py: Demonstrates minimal end-to-end usage with configurable pipeline types.
  • app.py: Full-featured demo application that forwards user-selected pipeline_type values to the underlying pipeline.

Summary

  • Four presets are available: SI2 (512³), IO24 (1024³), IO24-cascade (1024³), and IS36-cascade (1536³).
  • Latent grid size varies by preset: standard IO24 uses ss_res=64, while all others use ss_res=32.
  • Cascade modes (IO24-cascade, IS36-cascade) run coarse-to-fine sampling for improved quality at the cost of inference time.
  • Configuration occurs at inference time via the pipeline_type parameter in trellis2/pipelines/trellis2_image_to_3d.py.
  • No model reloading is required to switch between resolutions; change the argument passed to run().

Frequently Asked Questions

What is the difference between IO24 and IO24-cascade?

IO24 generates a 1024³ mesh in a single pass using a latent grid size of 64, while IO24-cascade uses a two-stage process with a latent grid size of 32, first predicting a coarse structure then refining it to 1024³. The cascade mode typically produces finer geometric details but requires longer inference time.

How does the latent grid size (ss_res) affect reconstruction?

The ss_res parameter in trellis2/pipelines/trellis2_image_to_3d.py determines the resolution of the structured-latent field before voxel decoding. A value of 64 (used in IO24) provides higher-capacity latent representations but demands more memory, while 32 (used in SI2 and cascade modes) is more memory-efficient and faster but may capture less fine detail in complex geometries.

Can I use custom resolutions not in the preset list?

No. The pipeline_type argument strictly accepts the four predefined aliases (SI2, IO24, IO24-cascade, IS36-cascade) which map to internal identifiers (512, 1024, 1024_cascade, 1536_cascade). To use arbitrary resolutions, you would need to modify the ss_res dictionary and potentially retrain the decoder in trellis2_image_to_3d.py.

Where is the pipeline_type validation handled?

Input validation and resolution mapping occur within the run method of Trellis2ImageTo3DPipeline in trellis2/pipelines/trellis2_image_to_3d.py. The method resolves the string alias to its internal identifier and selects the corresponding ss_res value before executing the forward pass.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →