How to Export Microduck RL Policies to ONNX

Microduck RL policies trained with rsl_rl can be exported to ONNX by running the scripts/export.py CLI wrapper, which downloads the checkpoint from Weights & Biases, rebuilds the policy network with baked-in observation normalization, and outputs a self-contained <TASK_ID>.onnx file.

The pollen-robotics/microduck_rl repository uses the rsl_rl PPO trainer to learn locomotion policies for the Microduck robot. To deploy these policies on physical hardware, you must convert the PyTorch checkpoint files (e.g., model_3000.pt) into ONNX format with the observation normalizer baked directly into the model graph. This process ensures the exported model is self-contained and requires no additional preprocessing at runtime.

Export Workflow Overview

The export pipeline follows a four-stage workflow:

  1. Identify the task configuration and source WandB run containing the desired checkpoint.
  2. Execute the CLI wrapper at scripts/export.py with the appropriate arguments.
  3. Process the checkpoint through mjlab_microduck.export.main, which handles reconstruction and normalization.
  4. Validate the resulting ONNX file to ensure inference parity with the original PyTorch policy.

Step-by-Step Export Process

1. Identify the Task and WandB Run

Each training run is logged to Weights & Biases (WandB). You need two pieces of information to locate your policy:

  • Task ID: Corresponds to an environment configuration module under src/mjlab_microduck/tasks/ (e.g., microduck_velocity).
  • WandB Run Path: Follows the format <entity>/<project>/<run_id> (e.g., pollen-robotics/microduck/rl-run-2024-09-01).

Checkpoints are stored as .pt files (e.g., model_3000.pt) within the WandB artifact storage associated with the run.

2. Run the Export Script

The repository provides a thin CLI entry point at scripts/export.py that forwards arguments to the core library implementation. The script imports and executes mjlab_microduck.export.main:


# scripts/export.py

from mjlab_microduck.export import main

if __name__ == "__main__":
    main()

Execute the export via the command line:

uv run scripts/export.py <TASK_ID> \
    --wandb-run-path <entity>/<project>/<run_id> \
    [--checkpoint <STEP>]
  • --wandb-run-path: Specifies the source WandB run. The exporter fetches the latest checkpoint automatically if --checkpoint is omitted.
  • --checkpoint: Optional integer specifying the training step (e.g., 3000 to retrieve model_3000.pt).

3. Core Logic Inside mjlab_microduck.export

The main function in src/mjlab_microduck/export.py performs the conversion through the following sequence:

  • Load Checkpoint: Downloads the checkpoint tarball from WandB or reads a local file if a local path is provided via --checkpoint.
  • Rebuild Policy: Uses the task configuration to reconstruct the policy network architecture and restore saved weights from the checkpoint.
  • Bake Observation Normalizer: The training pipeline uses a learned normalizer for the 61-dimensional observation vector. The exporter attaches this normalizer as the first computational layer within the exported graph, ensuring the ONNX model receives raw observations and produces normalized actions internally.
  • Export to ONNX: Invokes torch.onnx.export to serialize the policy (with baked normalizer) to <TASK_ID>.onnx in the current working directory.
  • Validation: Runs a single inference step through both the PyTorch policy and the exported ONNX model to verify output shape and numerical parity.

CLI and Programmatic Usage Examples

Export the Latest Checkpoint

uv run scripts/export.py microduck_velocity \
    --wandb-run-path pollen-robotics/microduck/rl-run-2024-09-01

Export a Specific Training Step

uv run scripts/export.py microduck_velocity \
    --wandb-run-path pollen-robotics/microduck/rl-run-2024-09-01 \
    --checkpoint 3000

Programmatic Export (Python)

You can invoke the export logic directly from Python:

from mjlab_microduck.export import main
import argparse

# Configure arguments programmatically

parser = argparse.ArgumentParser()
parser.add_argument('task_id')
parser.add_argument('--wandb-run-path', required=True)
parser.add_argument('--checkpoint', type=int, default=None)

args = parser.parse_args([
    'microduck_velocity',
    '--wandb-run-path', 'pollen-robotics/microduck/rl-run-2024-09-01',
    '--checkpoint', '3000'
])

main(args)  # Generates microduck_velocity.onnx

Key Source Files and Architecture

Understanding the codebase structure helps when customizing the export process:

  • scripts/export.py — Thin CLI wrapper that imports and executes the export logic.
  • src/mjlab_microduck/export.py — Core implementation containing the main function; handles checkpoint loading, policy reconstruction, normalizer baking, and ONNX serialization.
  • src/mjlab_microduck/tasks/*_env_cfg.py — Task configuration modules defining observation spaces, command slots, and domain randomization settings required to rebuild the policy network.
  • src/mjlab_microduck/publish/ — Utilities for building policy manifests; the ONNX export is reused by the uv run publish command for deployment packaging.

Summary

  • Microduck RL policies are stored as PyTorch checkpoints (.pt files) within WandB runs and must be converted to ONNX for real robot deployment.
  • The scripts/export.py CLI provides the primary interface, delegating to mjlab_microduck.export.main for heavy lifting.
  • The export process bakes the 61-dimensional observation normalizer into the ONNX graph, producing a self-contained model named <TASK_ID>.onnx.
  • The generated ONNX file integrates with the robot runtime via the robotctl policy add command without requiring additional observation preprocessing.

Frequently Asked Questions

What checkpoint format does Microduck RL use?

Microduck RL uses standard PyTorch checkpoint files (.pt) generated by the rsl_rl trainer. These files contain the policy network weights and optimizer states, typically named model_<step>.pt (e.g., model_3000.pt). The export script automatically locates these files within WandB artifacts or local filesystem paths.

How is the observation normalizer handled during export?

The observation normalizer is baked into the ONNX graph as the first layer. During training, the policy learns a running mean and standard deviation for the 61-dimensional observation vector. The exporter retrieves these statistics from the checkpoint and constructs a preprocessing layer that is fused with the policy network before calling torch.onnx.export. This ensures the runtime receives raw sensor data and the model handles normalization internally.

Can I export a policy without using Weights & Biases?

Yes. While the standard workflow uses --wandb-run-path to fetch remote checkpoints, you can export a local checkpoint by providing a local filesystem path to the --checkpoint argument. The export logic detects local paths and bypasses the WandB download, loading the .pt file directly from disk.

How do I verify the exported ONNX model is correct?

The mjlab_microduck.export module performs an automatic validation step after serialization. It runs a single forward pass through both the original PyTorch policy and the exported ONNX model using identical random inputs, comparing output shapes and values to ensure parity. Additionally, you can manually verify the model by loading it with onnxruntime and checking that it produces the expected action dimensions for a dummy observation input.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →