# How to Fine-Tune TRELLIS.2 on Custom Datasets Using O-Voxel Preprocessing

> Learn to fine-tune TRELLIS.2 on custom datasets. Convert 3D assets to O-Voxel format, create a PyTorch Dataset, and train using a custom configuration for optimal results.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: how-to-guide
- Published: 2026-08-04

---

**Fine-tuning TRELLIS.2 requires converting your 3D assets to the O-Voxel (VXZ) format using the provided rasterization utilities, implementing a PyTorch Dataset class that loads these sparse voxel files, and executing the training pipeline via [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py) with a custom configuration pointing to your preprocessed data.**

The microsoft/TRELLIS.2 repository provides a modular 3D generation framework that relies on the efficient O-Voxel format for data storage. To fine-tune TRELLIS.2 on your own meshes, point clouds, or multi-view images, you must first preprocess these assets into compressed sparse voxel grids. This guide walks through the complete pipeline from raw asset conversion to model training, referencing the actual source implementation in [`o_voxel/io/vxz.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/io/vxz.py) and [`trellis2/datasets/sparse_voxel_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/datasets/sparse_voxel_pbr.py).

## Convert Raw Assets to O-Voxel Format

TRELLIS.2 expects all training data in the **O-Voxel (VXZ)** format, a compact binary representation of sparse voxel grids. The preprocessing step rasterizes raw geometry into voxel coordinates and serializes them using compression and chunking to optimize I/O performance.

### Rasterize Meshes to VXZ Files

The [`voxelize_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/voxelize_pbr.py) script in the repository demonstrates how to read geometry and rasterize it into a sparse voxel grid. At the core of this process are two functions from [`o_voxel/io/vxz.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/io/vxz.py): `rasterize_mesh` converts vertices and faces into voxel indices, while `write_vxz` serializes the data to disk with optional compression.

```python

# Minimal conversion example adapted from o-voxel/examples/mesh2ovox.py

from o_voxel.io import write_vxz
from o_voxel.rasterize import rasterize_mesh
import torch

# Load your mesh (implementation-specific loader)

verts, faces = load_my_mesh("my_object.obj")

# Rasterize into sparse voxel grid with default chunk size of 256

voxel_indices, attrs = rasterize_mesh(verts, faces, chunk_size=256)

# Write compressed VXZ file

write_vxz(
    "my_dataset/obj_001.vxz",
    coord=torch.from_numpy(voxel_indices).int(),
    attr={"rgb": torch.from_numpy(attrs["rgb"]).uint8()},
    chunk_size=256,
    compression="zstd",  # Options: none, deflate, lzma, zstd

)

```

The VXZ format stores **coordinates**, **per-voxel attributes** (such as RGB color), **chunking metadata**, and **compression headers** (supporting deflate, lzma, or zstd algorithms). This structure ensures fast random access and minimal memory footprint when loading large 3D datasets during training.

## Create a Custom PyTorch Dataset

Once your assets are converted to VXZ files, you must implement a dataset class that the TRELLIS.2 training loop can consume. The repository provides a reference implementation in [`trellis2/datasets/sparse_voxel_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/datasets/sparse_voxel_pbr.py) called `SparseVoxelPBR`, which demonstrates the required interface.

You can subclass `SparseVoxelPBR` or create a custom `Dataset` that utilizes `read_vxz` from [`o_voxel/io/vxz.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/io/vxz.py) to deserialize the sparse tensors:

```python
import torch
from torch.utils.data import Dataset
from o_voxel.io import read_vxz
import os

class MyVoxelDataset(Dataset):
    def __init__(self, root_dir):
        self.files = sorted([
            os.path.join(root_dir, f) 
            for f in os.listdir(root_dir) 
            if f.endswith(".vxz")
        ])

    def __len__(self):
        return len(self.files)

    def __getitem__(self, idx):
        coord, attr = read_vxz(self.files[idx])
        # Normalize RGB attributes to [0, 1]

        rgb = attr["rgb"].float() / 255.0
        return {"coord": coord, "rgb": rgb}

```

The `read_vxz` function automatically handles decompression and chunk reconstruction, returning coordinate tensors and attribute dictionaries that map directly to the structure expected by TRELLIS.2 models like **SparseStructureVAE**.

## Configure the Training Pipeline

The training entry point [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py) accepts configuration files (JSON or YAML) and command-line arguments to specify dataset paths, model architectures, and hyperparameters. To fine-tune on your custom data, you must update the `dataset` configuration entry to point to your VXZ directory and specify your custom dataset class.

A minimal configuration file for fine-tuning looks like this:

```yaml
model:
  name: SparseStructureVAE
  latent_dim: 256

train:
  batch_size: 4
  learning_rate: 3e-4
  epochs: 100

dataset:
  class: MyVoxelDataset  # Import path or class name

  root: /path/to/my_dataset

```

The training script instantiates the model, wraps your dataset with a PyTorch DataLoader, and executes the standard TRELLIS.2 training loop. By default, the pipeline automatically applies any required filtering and compression handling through the I/O utilities in [`o_voxel/io/vxz.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/io/vxz.py).

## Launch the Fine-Tuning Job

Execute the fine-tuning process by invoking [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py) with your updated configuration:

```bash
python train.py \
    --config configs/fine_tune.yaml \
    --dataset_root /path/to/my_dataset \
    --output_dir ./fine_tuned_checkpoints

```

During execution, the training loop calls your dataset's `__getitem__` method, which invokes `read_vxz` to stream voxel coordinates and attributes into the model. The **SparseStructureVAE** architecture (or whichever model you specify) processes these sparse tensors to learn the distribution of your custom 3D data.

## Summary

- **Convert assets** to the O-Voxel format using `rasterize_mesh` and `write_vxz` from [`o_voxel/io/vxz.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/io/vxz.py), storing compressed sparse voxel grids with optional zstd compression.
- **Implement a Dataset class** that utilizes `read_vxz` to load coordinate and attribute tensors from your VXZ files, following the interface demonstrated in [`trellis2/datasets/sparse_voxel_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/datasets/sparse_voxel_pbr.py).
- **Configure the training run** via YAML or JSON configs specifying your dataset class, model architecture (e.g., `SparseStructureVAE`), and hyperparameters for [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py).
- **Execute fine-tuning** by launching [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py) with the custom configuration, allowing the pipeline to automatically handle decompression and tensor loading.

## Frequently Asked Questions

### What is the O-Voxel (VXZ) format and why does TRELLIS.2 require it?

The O-Voxel format is a compact binary specification for sparse 3D voxel grids that encodes coordinates, per-voxel attributes, chunking metadata, and compression headers. TRELLIS.2 uses VXZ files because they enable fast I/O and memory-efficient storage of large 3D datasets, with support for multiple compression algorithms (deflate, lzma, zstd) and chunk-based access patterns implemented in [`o_voxel/io/vxz.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/io/vxz.py).

### Can I fine-tune on point clouds or multi-view images instead of meshes?

Yes, the preprocessing pipeline is input-agnostic as long as you rasterize the data into sparse voxel coordinates before serialization. You can adapt the [`voxelize_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/voxelize_pbr.py) logic or the [`mesh2ovox.py`](https://github.com/microsoft/TRELLIS.2/blob/main/mesh2ovox.py) example to sample point clouds into voxel grids or perform multi-view reconstruction into volumetric representations before writing them to VXZ files using `write_vxz`.

### How does the training pipeline handle large datasets that exceed system memory?

The VXZ format supports chunked storage with a default chunk size of 256, and the `read_vxz` function streams only necessary chunks during training. Combined with compression options like zstd, this allows TRELLIS.2 to train on massive datasets while keeping only active batches in GPU memory, as handled by the PyTorch DataLoader integration in [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py).

### Where is the fine-tuning logic implemented in the source code?

The primary training logic resides in [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py), which parses configurations, instantiates the model (such as `SparseStructureVAE`), and manages the training loop. The data loading interface is defined in [`trellis2/datasets/sparse_voxel_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/datasets/sparse_voxel_pbr.py), while low-level voxel I/O operations are contained in [`o_voxel/io/vxz.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/io/vxz.py), specifically the `read_vxz` and `write_vxz` functions.