# How to Preprocess Training Data Using the DataToolkit for O-Voxel Conversion

> Learn to preprocess training data for O-Voxel conversion using TRELLIS.2's DataToolkit. Discover, normalize, and aggregate your PBR mesh dumps efficiently.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: how-to-guide
- Published: 2026-08-04

---

**The DataToolkit in microsoft/TRELLIS.2 provides the [`voxelize_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/voxelize_pbr.py) script to convert PBR mesh dumps into O-Voxel volumetric representations through a three-phase workflow involving dataset discovery, geometry normalization, and result aggregation.**

Preparing training data for the TRELLIS.2 framework requires converting physically-based rendered (PBR) mesh dumps into structured O-Voxel formats. The repository's DataToolkit automates this transformation through the [`data_toolkit/voxelize_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/voxelize_pbr.py) module, which orchestrates parallel processing from metadata filtering to volumetric rasterization. This guide explains how to preprocess training data using the data_toolkit for O-Voxel conversion with implementation details derived from the source code.

## Prerequisites and Dataset Structure

Before invoking the voxelization pipeline, organize your dataset with a **metadata.csv** file and a **pbr_dumps/** folder containing pickle files named by SHA256 hash. Each pickle file contains Blender-exported scene data with `materials` and `objects` fields describing the mesh geometry and surface properties.

## The Three-Phase Voxelization Workflow

The conversion logic in [`data_toolkit/voxelize_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/voxelize_pbr.py) executes three sequential phases to transform raw meshes into training-ready voxel grids.

### Phase 1: Dataset Discovery and Filtering

The script initializes by loading dataset metadata to enable targeted filtering. According to lines 15-22 of [`data_toolkit/voxelize_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/voxelize_pbr.py), the implementation uses pandas to index entries by SHA256 hash:

```python
metadata = pd.read_csv(...).set_index('sha256')

```

You can apply constraints using `--filter_low_aesthetic_score` to skip low-quality models below a threshold, or `--instances` to whitelist specific SHA256 hashes from a CSV file for selective processing.

### Phase 2: Geometry Normalization and Voxelization

For each selected dump, the script invokes **`_pbr_voxelize`** (lines 55-73), which performs critical preprocessing before voxelization. The function normalizes geometry into a **[-0.5, 0.5]** bounding cube and applies fallback materials when needed. The core conversion delegates to **`o_voxel.convert.blender_dump_to_volumetric_attr`**, which rasterizes the normalized mesh into dense voxel coordinates and attribute tensors at the requested resolution.

### Phase 3: Result Aggregation and Output

Successful conversions write binary **.vxz** files using **`o_voxel.io.write_vxz`** while maintaining CSV ledgers at `pbr_voxels_<res>/new_records/part_<rank>.csv`. These records track the `sha256` hash, `pbr_voxelized` status, and `num_pbr_voxels` count. Any conversion failures are captured in **errors.txt** for later inspection without halting the batch process.

## Running the Voxelization Script

Execute the complete pipeline using the module command with dataset-specific parameters:

```bash
python -m data_toolkit.voxelize_pbr \
    structured_latent_pbr \
    --root /path/to/your/dataset_root \
    --pbr_dump_root /path/to/pbr_dumps \
    --pbr_voxel_root /path/to/output_voxels \
    --resolution 256,512,1024 \
    --max_workers 8

```

The first positional argument selects a dataset helper module (e.g., `datasets.structured_latent_pbr`) that implements **`add_args`** and **`foreach_instance`** for parallel iteration. Adjust **`max_workers`** to control multiprocessing—set to `0` for serial execution.

## Implementing Custom Voxelization Pipelines

For integration into custom workflows, replicate the core logic from `_pbr_voxelize` using the O-Voxel library directly:

```python
import torch, numpy as np, o_voxel, pickle
from pathlib import Path

def voxelize_one(pickle_path: Path, out_path: Path, grid_size: int = 256):
    # Load dump

    dump = pickle.loads(pickle_path.read_bytes())
    
    # Normalise geometry to [-0.5, 0.5]^3

    verts = torch.from_numpy(np.concatenate([obj['vertices'] for obj in dump['objects']]))
    center = (verts.min(0)[0] + verts.max(0)[0]) / 2
    scale = 0.99999 / (verts.max(0)[0] - verts.min(0)[0]).max()
    
    for obj in dump['objects']:
        obj['vertices'] = (torch.from_numpy(obj['vertices']).float() - center) * scale
        obj['vertices'] = obj['vertices'].numpy()
    
    # Convert to voxel grid

    coord, attr = o_voxel.convert.blender_dump_to_volumetric_attr(
        dump, 
        grid_size=grid_size,
        aabb=[[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]],
        mip_level_offset=0,
        verbose=False,
        timing=False
    )
    
    # Remove unused attributes

    del attr['normal']
    del attr['emissive']
    
    # Write binary voxel file

    o_voxel.io.write_vxz(str(out_path), coord, attr)

```

## Key Source Files and Functions

Understanding these components helps when debugging or extending the pipeline:

- **[`data_toolkit/voxelize_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/voxelize_pbr.py)**: Main entry point handling CLI arguments and parallel execution via `_pbr_voxelize`
- **`o_voxel/convert/blender_dump_to_volumetric_attr`**: Core rasterization function converting Blender dumps to dense voxel grids with attribute channels
- **[`o_voxel/io/vxz.py`](https://github.com/microsoft/TRELLIS.2/blob/main/o_voxel/io/vxz.py)**: Binary I/O operations for the `.vxz` format used throughout TRELLIS.2 training pipelines
- **[`datasets/structured_latent_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/datasets/structured_latent_pbr.py)**: Example dataset module implementing `foreach_instance` for distributed processing across workers

## Summary

- Prepare input data with `metadata.csv` and `pbr_dumps/` folder containing SHA256-named pickle files exported from Blender
- Execute `python -m data_toolkit.voxelize_pbr` with dataset name, root path, and target resolutions to run the three-phase workflow
- The `_pbr_voxelize` function normalizes geometry to `[-0.5, 0.5]` before calling `o_voxel.convert.blender_dump_to_volumetric_attr` for rasterization
- Outputs include `.vxz` binary files, CSV ledgers in `pbr_voxels_<res>/new_records/`, and [`errors.txt`](https://github.com/microsoft/TRELLIS.2/blob/main/errors.txt) for troubleshooting failed conversions
- Control parallelization with `max_workers` and filter datasets using aesthetic score thresholds or instance whitelists

## Frequently Asked Questions

### What file format does the O-Voxel conversion produce?

The pipeline generates **.vxz** files, a binary format containing voxel coordinates and attribute tensors. These files are written using `o_voxel.io.write_vxz` and store the volumetric representations required for TRELLIS.2 training, with processing metadata tracked in accompanying CSV files under the `new_records` directory.

### How does the script handle geometry normalization?

The `_pbr_voxelize` function in [`data_toolkit/voxelize_pbr.py`](https://github.com/microsoft/TRELLIS.2/blob/main/data_toolkit/voxelize_pbr.py) computes the bounding box center and scale factor to fit meshes into a `[-0.5, 0.5]` cube. It applies a 0.99999 scaling factor to prevent boundary edge cases while maintaining aspect ratios, ensuring consistent voxel grid alignment across different mesh sizes before calling the O-Voxel conversion routines.

### Can I process multiple resolutions simultaneously?

Yes, pass comma-separated values to the `--resolution` argument (e.g., `--resolution 256,512,1024`). The script voxelizes each dump at all specified resolutions, creating separate output directories (`pbr_voxels_256`, `pbr_voxels_512`, etc.) with corresponding CSV ledgers for each grid size.

### What causes entries in the errors.txt file?

Failed conversions typically result from malformed pickle dumps, missing geometric data, or memory exhaustion during high-resolution voxelization. The script catches exceptions during `_pbr_voxelize` execution and writes the SHA256 hash and error traceback to [`errors.txt`](https://github.com/microsoft/TRELLIS.2/blob/main/errors.txt) without interrupting the batch processing of remaining instances.