How to Fine-Tune TRELLIS.2 on Custom Datasets Using O-Voxel Preprocessing

Fine-tuning TRELLIS.2 requires converting your 3D assets to the O-Voxel (VXZ) format using the provided rasterization utilities, implementing a PyTorch Dataset class that loads these sparse voxel files, and executing the training pipeline via train.py with a custom configuration pointing to your preprocessed data.

The microsoft/TRELLIS.2 repository provides a modular 3D generation framework that relies on the efficient O-Voxel format for data storage. To fine-tune TRELLIS.2 on your own meshes, point clouds, or multi-view images, you must first preprocess these assets into compressed sparse voxel grids. This guide walks through the complete pipeline from raw asset conversion to model training, referencing the actual source implementation in o_voxel/io/vxz.py and trellis2/datasets/sparse_voxel_pbr.py.

Convert Raw Assets to O-Voxel Format

TRELLIS.2 expects all training data in the O-Voxel (VXZ) format, a compact binary representation of sparse voxel grids. The preprocessing step rasterizes raw geometry into voxel coordinates and serializes them using compression and chunking to optimize I/O performance.

Rasterize Meshes to VXZ Files

The voxelize_pbr.py script in the repository demonstrates how to read geometry and rasterize it into a sparse voxel grid. At the core of this process are two functions from o_voxel/io/vxz.py: rasterize_mesh converts vertices and faces into voxel indices, while write_vxz serializes the data to disk with optional compression.


# Minimal conversion example adapted from o-voxel/examples/mesh2ovox.py

from o_voxel.io import write_vxz
from o_voxel.rasterize import rasterize_mesh
import torch

# Load your mesh (implementation-specific loader)

verts, faces = load_my_mesh("my_object.obj")

# Rasterize into sparse voxel grid with default chunk size of 256

voxel_indices, attrs = rasterize_mesh(verts, faces, chunk_size=256)

# Write compressed VXZ file

write_vxz(
    "my_dataset/obj_001.vxz",
    coord=torch.from_numpy(voxel_indices).int(),
    attr={"rgb": torch.from_numpy(attrs["rgb"]).uint8()},
    chunk_size=256,
    compression="zstd",  # Options: none, deflate, lzma, zstd

)

The VXZ format stores coordinates, per-voxel attributes (such as RGB color), chunking metadata, and compression headers (supporting deflate, lzma, or zstd algorithms). This structure ensures fast random access and minimal memory footprint when loading large 3D datasets during training.

Create a Custom PyTorch Dataset

Once your assets are converted to VXZ files, you must implement a dataset class that the TRELLIS.2 training loop can consume. The repository provides a reference implementation in trellis2/datasets/sparse_voxel_pbr.py called SparseVoxelPBR, which demonstrates the required interface.

You can subclass SparseVoxelPBR or create a custom Dataset that utilizes read_vxz from o_voxel/io/vxz.py to deserialize the sparse tensors:

import torch
from torch.utils.data import Dataset
from o_voxel.io import read_vxz
import os

class MyVoxelDataset(Dataset):
    def __init__(self, root_dir):
        self.files = sorted([
            os.path.join(root_dir, f) 
            for f in os.listdir(root_dir) 
            if f.endswith(".vxz")
        ])

    def __len__(self):
        return len(self.files)

    def __getitem__(self, idx):
        coord, attr = read_vxz(self.files[idx])
        # Normalize RGB attributes to [0, 1]

        rgb = attr["rgb"].float() / 255.0
        return {"coord": coord, "rgb": rgb}

The read_vxz function automatically handles decompression and chunk reconstruction, returning coordinate tensors and attribute dictionaries that map directly to the structure expected by TRELLIS.2 models like SparseStructureVAE.

Configure the Training Pipeline

The training entry point train.py accepts configuration files (JSON or YAML) and command-line arguments to specify dataset paths, model architectures, and hyperparameters. To fine-tune on your custom data, you must update the dataset configuration entry to point to your VXZ directory and specify your custom dataset class.

A minimal configuration file for fine-tuning looks like this:

model:
  name: SparseStructureVAE
  latent_dim: 256

train:
  batch_size: 4
  learning_rate: 3e-4
  epochs: 100

dataset:
  class: MyVoxelDataset  # Import path or class name

  root: /path/to/my_dataset

The training script instantiates the model, wraps your dataset with a PyTorch DataLoader, and executes the standard TRELLIS.2 training loop. By default, the pipeline automatically applies any required filtering and compression handling through the I/O utilities in o_voxel/io/vxz.py.

Launch the Fine-Tuning Job

Execute the fine-tuning process by invoking train.py with your updated configuration:

python train.py \
    --config configs/fine_tune.yaml \
    --dataset_root /path/to/my_dataset \
    --output_dir ./fine_tuned_checkpoints

During execution, the training loop calls your dataset's __getitem__ method, which invokes read_vxz to stream voxel coordinates and attributes into the model. The SparseStructureVAE architecture (or whichever model you specify) processes these sparse tensors to learn the distribution of your custom 3D data.

Summary

  • Convert assets to the O-Voxel format using rasterize_mesh and write_vxz from o_voxel/io/vxz.py, storing compressed sparse voxel grids with optional zstd compression.
  • Implement a Dataset class that utilizes read_vxz to load coordinate and attribute tensors from your VXZ files, following the interface demonstrated in trellis2/datasets/sparse_voxel_pbr.py.
  • Configure the training run via YAML or JSON configs specifying your dataset class, model architecture (e.g., SparseStructureVAE), and hyperparameters for train.py.
  • Execute fine-tuning by launching train.py with the custom configuration, allowing the pipeline to automatically handle decompression and tensor loading.

Frequently Asked Questions

What is the O-Voxel (VXZ) format and why does TRELLIS.2 require it?

The O-Voxel format is a compact binary specification for sparse 3D voxel grids that encodes coordinates, per-voxel attributes, chunking metadata, and compression headers. TRELLIS.2 uses VXZ files because they enable fast I/O and memory-efficient storage of large 3D datasets, with support for multiple compression algorithms (deflate, lzma, zstd) and chunk-based access patterns implemented in o_voxel/io/vxz.py.

Can I fine-tune on point clouds or multi-view images instead of meshes?

Yes, the preprocessing pipeline is input-agnostic as long as you rasterize the data into sparse voxel coordinates before serialization. You can adapt the voxelize_pbr.py logic or the mesh2ovox.py example to sample point clouds into voxel grids or perform multi-view reconstruction into volumetric representations before writing them to VXZ files using write_vxz.

How does the training pipeline handle large datasets that exceed system memory?

The VXZ format supports chunked storage with a default chunk size of 256, and the read_vxz function streams only necessary chunks during training. Combined with compression options like zstd, this allows TRELLIS.2 to train on massive datasets while keeping only active batches in GPU memory, as handled by the PyTorch DataLoader integration in train.py.

Where is the fine-tuning logic implemented in the source code?

The primary training logic resides in train.py, which parses configurations, instantiates the model (such as SparseStructureVAE), and manages the training loop. The data loading interface is defined in trellis2/datasets/sparse_voxel_pbr.py, while low-level voxel I/O operations are contained in o_voxel/io/vxz.py, specifically the read_vxz and write_vxz functions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →