How to Preprocess Training Data Using the DataToolkit for O-Voxel Conversion

The DataToolkit in microsoft/TRELLIS.2 provides the voxelize_pbr.py script to convert PBR mesh dumps into O-Voxel volumetric representations through a three-phase workflow involving dataset discovery, geometry normalization, and result aggregation.

Preparing training data for the TRELLIS.2 framework requires converting physically-based rendered (PBR) mesh dumps into structured O-Voxel formats. The repository's DataToolkit automates this transformation through the data_toolkit/voxelize_pbr.py module, which orchestrates parallel processing from metadata filtering to volumetric rasterization. This guide explains how to preprocess training data using the data_toolkit for O-Voxel conversion with implementation details derived from the source code.

Prerequisites and Dataset Structure

Before invoking the voxelization pipeline, organize your dataset with a metadata.csv file and a pbr_dumps/ folder containing pickle files named by SHA256 hash. Each pickle file contains Blender-exported scene data with materials and objects fields describing the mesh geometry and surface properties.

The Three-Phase Voxelization Workflow

The conversion logic in data_toolkit/voxelize_pbr.py executes three sequential phases to transform raw meshes into training-ready voxel grids.

Phase 1: Dataset Discovery and Filtering

The script initializes by loading dataset metadata to enable targeted filtering. According to lines 15-22 of data_toolkit/voxelize_pbr.py, the implementation uses pandas to index entries by SHA256 hash:

metadata = pd.read_csv(...).set_index('sha256')

You can apply constraints using --filter_low_aesthetic_score to skip low-quality models below a threshold, or --instances to whitelist specific SHA256 hashes from a CSV file for selective processing.

Phase 2: Geometry Normalization and Voxelization

For each selected dump, the script invokes _pbr_voxelize (lines 55-73), which performs critical preprocessing before voxelization. The function normalizes geometry into a [-0.5, 0.5] bounding cube and applies fallback materials when needed. The core conversion delegates to o_voxel.convert.blender_dump_to_volumetric_attr, which rasterizes the normalized mesh into dense voxel coordinates and attribute tensors at the requested resolution.

Phase 3: Result Aggregation and Output

Successful conversions write binary .vxz files using o_voxel.io.write_vxz while maintaining CSV ledgers at pbr_voxels_<res>/new_records/part_<rank>.csv. These records track the sha256 hash, pbr_voxelized status, and num_pbr_voxels count. Any conversion failures are captured in errors.txt for later inspection without halting the batch process.

Running the Voxelization Script

Execute the complete pipeline using the module command with dataset-specific parameters:

python -m data_toolkit.voxelize_pbr \
    structured_latent_pbr \
    --root /path/to/your/dataset_root \
    --pbr_dump_root /path/to/pbr_dumps \
    --pbr_voxel_root /path/to/output_voxels \
    --resolution 256,512,1024 \
    --max_workers 8

The first positional argument selects a dataset helper module (e.g., datasets.structured_latent_pbr) that implements add_args and foreach_instance for parallel iteration. Adjust max_workers to control multiprocessing—set to 0 for serial execution.

Implementing Custom Voxelization Pipelines

For integration into custom workflows, replicate the core logic from _pbr_voxelize using the O-Voxel library directly:

import torch, numpy as np, o_voxel, pickle
from pathlib import Path

def voxelize_one(pickle_path: Path, out_path: Path, grid_size: int = 256):
    # Load dump

    dump = pickle.loads(pickle_path.read_bytes())
    
    # Normalise geometry to [-0.5, 0.5]^3

    verts = torch.from_numpy(np.concatenate([obj['vertices'] for obj in dump['objects']]))
    center = (verts.min(0)[0] + verts.max(0)[0]) / 2
    scale = 0.99999 / (verts.max(0)[0] - verts.min(0)[0]).max()
    
    for obj in dump['objects']:
        obj['vertices'] = (torch.from_numpy(obj['vertices']).float() - center) * scale
        obj['vertices'] = obj['vertices'].numpy()
    
    # Convert to voxel grid

    coord, attr = o_voxel.convert.blender_dump_to_volumetric_attr(
        dump, 
        grid_size=grid_size,
        aabb=[[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]],
        mip_level_offset=0,
        verbose=False,
        timing=False
    )
    
    # Remove unused attributes

    del attr['normal']
    del attr['emissive']
    
    # Write binary voxel file

    o_voxel.io.write_vxz(str(out_path), coord, attr)

Key Source Files and Functions

Understanding these components helps when debugging or extending the pipeline:

  • data_toolkit/voxelize_pbr.py: Main entry point handling CLI arguments and parallel execution via _pbr_voxelize
  • o_voxel/convert/blender_dump_to_volumetric_attr: Core rasterization function converting Blender dumps to dense voxel grids with attribute channels
  • o_voxel/io/vxz.py: Binary I/O operations for the .vxz format used throughout TRELLIS.2 training pipelines
  • datasets/structured_latent_pbr.py: Example dataset module implementing foreach_instance for distributed processing across workers

Summary

  • Prepare input data with metadata.csv and pbr_dumps/ folder containing SHA256-named pickle files exported from Blender
  • Execute python -m data_toolkit.voxelize_pbr with dataset name, root path, and target resolutions to run the three-phase workflow
  • The _pbr_voxelize function normalizes geometry to [-0.5, 0.5] before calling o_voxel.convert.blender_dump_to_volumetric_attr for rasterization
  • Outputs include .vxz binary files, CSV ledgers in pbr_voxels_<res>/new_records/, and errors.txt for troubleshooting failed conversions
  • Control parallelization with max_workers and filter datasets using aesthetic score thresholds or instance whitelists

Frequently Asked Questions

What file format does the O-Voxel conversion produce?

The pipeline generates .vxz files, a binary format containing voxel coordinates and attribute tensors. These files are written using o_voxel.io.write_vxz and store the volumetric representations required for TRELLIS.2 training, with processing metadata tracked in accompanying CSV files under the new_records directory.

How does the script handle geometry normalization?

The _pbr_voxelize function in data_toolkit/voxelize_pbr.py computes the bounding box center and scale factor to fit meshes into a [-0.5, 0.5] cube. It applies a 0.99999 scaling factor to prevent boundary edge cases while maintaining aspect ratios, ensuring consistent voxel grid alignment across different mesh sizes before calling the O-Voxel conversion routines.

Can I process multiple resolutions simultaneously?

Yes, pass comma-separated values to the --resolution argument (e.g., --resolution 256,512,1024). The script voxelizes each dump at all specified resolutions, creating separate output directories (pbr_voxels_256, pbr_voxels_512, etc.) with corresponding CSV ledgers for each grid size.

What causes entries in the errors.txt file?

Failed conversions typically result from malformed pickle dumps, missing geometric data, or memory exhaustion during high-resolution voxelization. The script catches exceptions during _pbr_voxelize execution and writes the SHA256 hash and error traceback to errors.txt without interrupting the batch processing of remaining instances.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →