How to Export Trained Policies From Microduck RL to ONNX: Complete Guide

Microduck RL provides a dedicated scripts/export.py command-line tool that converts trained RSL-RL policies into portable ONNX models with baked-in observation normalization and embedded environment metadata.

Exporting trained reinforcement learning policies to ONNX enables seamless deployment on physical robots without requiring the full training stack. Pollen Robotics' Microduck RL repository implements a streamlined export pipeline through the export.py script in the scripts/ directory. This article walks through the complete export process, from checkpoint resolution to metadata embedding, based on the actual source code implementation in pollen-robotics/microduck_rl.


Prerequisites and Environment Setup

Before exporting, ensure your environment matches the repository's locked dependencies. The pyproject.toml specifies exact versions for PyTorch, RSL-RL, and the MJLab simulation framework.


# Install all runtime dependencies

uv sync

The export script requires either:

  • A local checkpoint file (.pt or .pth), or
  • A WandB run path with stored checkpoints

The Export Script: scripts/export.py

The core export functionality lives in scripts/export.py. This script implements a six-step pipeline in its run_export function:

Step 1: Load Environment and Agent Configurations

The script begins by loading task-specific configurations through two helper functions:

  • load_env_cfg (line 51): Loads the MJLab environment configuration for the specified task
  • load_rl_cfg (line 52): Loads the RSL-RL agent hyperparameters and network architecture

These configurations ensure the exported model matches the training environment's observation and action spaces.

Step 2: Resolve the Checkpoint Source

Checkpoint handling spans lines 27–71 and supports two resolution paths:

Source Flag Behavior
Local file --checkpoint-file Load .pt file directly from disk
WandB artifact --wandb-run-path Download checkpoint via WandB API by iteration number or latest

If using WandB, the script authenticates via the wandb Python API and retrieves the requested checkpoint artifact. The --checkpoint flag specifies which iteration to fetch when multiple checkpoints exist.

Step 3: Initialize the Policy and Load Weights

With configurations loaded, the script instantiates an OnPolicyRunner (RSL-RL's standard training orchestrator) and restores the policy state:


# From scripts/export.py lines 22-26

runner.load(checkpoint_path)           # Restores actor-critic weights

policy = runner.get_inference_policy() # Returns wrapped inference-only policy

The get_inference_policy() method returns a wrapped policy optimized for inference: it disables gradient computation and applies any training-time wrappers (such as action clipping or observation filtering).

Step 4: Export to ONNX Format

The actual ONNX conversion occurs through the runner's built-in method:


# From scripts/export.py lines 226-239

runner.export_policy_to_onnx(
    policy, 
    onnx_file_path, 
    target_opset_version=12  # Configurable for compatibility

)

Critical detail: export_policy_to_onnx automatically bakes the observation normalizer into the computation graph. During training, RSL-RL maintains running statistics for observation normalization; these statistics are fused into the ONNX model as constant scale and bias operations. The exported model therefore expects raw, unnormalized observations as input—no preprocessing required at inference time.

Step 5: Attach Environment Metadata

After the raw ONNX file is written, the script embeds additional context for runtime consumption (lines 40–42):

from mjlab.rl.exporter_utils import get_base_metadata, attach_metadata_to_onnx

metadata = get_base_metadata(env_cfg)      # Extract obs layout, command slots, etc.

attach_metadata_to_onnx(onnx_file_path, metadata)

This metadata includes:

  • Observation dimension and semantic layout (61-dimensional vector for Microduck tasks)
  • Command structure and valid ranges
  • Environment-specific constants from the task definition

Storing metadata inside the ONNX file eliminates the need for sidecar configuration files during robot deployment.

Step 6: Finalization

The script prints the absolute path to the generated ONNX file and gracefully shuts down the MJLab environment.


Complete Export Examples

Export From WandB Checkpoint


# Export iteration 3000 from a completed WandB training run

uv run scripts/export.py velocity \
    --onnx-file policies/velocity_3000.onnx \
    --wandb-run-path pollen-robotics/microduck-rl/1a2b3c4d5e6f7g8h \
    --checkpoint 3000

Export From Local Checkpoint


# Export from a manually saved checkpoint file

uv run scripts/export.py velocity \
    --onnx-file policies/velocity_local.onnx \
    --checkpoint-file /path/to/checkpoint_3000.pt

Export Latest Checkpoint Automatically


# Omit --checkpoint to use the highest iteration number available

uv run scripts/export.py velocity \
    --onnx-file policies/velocity_latest.onnx \
    --wandb-run-path pollen-robotics/microduck-rl/1a2b3c4d5e6f7g8h

Loading and Running the Exported ONNX Model

The exported ONNX file is self-contained and can be loaded without the training dependencies. The mjlab.rl.exporter_utils module provides load_onnx_policy for convenient inference:

from mjlab.utils.torch import configure_torch_backends
from mjlab.rl.exporter_utils import load_onnx_policy
import numpy as np

# Optional: optimize PyTorch backend for inference

configure_torch_backends()

# Load the exported policy

policy = load_onnx_policy("policies/velocity_3000.onnx")

# Prepare observation: shape (batch, 61) for Microduck velocity task

# The 61 dimensions comprise: base angular velocity (3), projected gravity (3),

# joint positions (12), joint velocities (12), previous actions (12),

# and command inputs (19) — see src/mjlab_microduck/tasks/mdp.py

obs = np.zeros((1, 61), dtype=np.float32)

# Run inference: returns torch.Tensor of shape (batch, 12) for joint targets

action = policy(obs)
print(action.shape)  # torch.Size([1, 12])

The observation layout is defined in src/mjlab_microduck/tasks/mdp.py and matches the 61-dimensional structure used during training. Because normalization is baked into the ONNX graph, pass raw sensor values directly.


Integration With Robot Runtime

The exported ONNX model integrates directly with the physical robot stack in pollen-robotics/microduck. The scripts/infer_policy.py example demonstrates loading the policy in a simulated environment before hardware deployment:


# Test the exported policy in simulation

uv run scripts/infer_policy.py \
    --onnx-file policies/velocity_3000.onnx \
    --env velocity \
    --num-envs 1

For production deployment, the same ONNX file loads into the robot's real-time control loop via ONNX Runtime or TensorRT, with the embedded metadata providing automatic configuration.


Key Source Files Reference

File Purpose Lines of Interest
scripts/export.py Main export script with run_export implementation 1–250
src/mjlab_microduck/train_cli.py Training entry point producing exportable checkpoints Full file
src/mjlab_microduck/tasks/mdp.py Observation space definition (61-dim) Full file
scripts/infer_policy.py Example ONNX inference loader Full file
src/mjlab_microduck/rl/exporter_utils.py Metadata extraction and attachment utilities get_base_metadata, attach_metadata_to_onnx

Summary

  • Microduck RL exports to ONNX via scripts/export.py, a dedicated command-line tool that wraps RSL-RL's native export functionality with environment-specific metadata handling.

  • Checkpoint resolution supports both local files and WandB artifacts, enabling export from finished training runs or intermediate iterations.

  • Observation normalization is automatically baked into the ONNX graph via export_policy_to_onnx, so the exported model accepts raw sensor inputs directly.

  • Environment metadata is embedded in the ONNX file through attach_metadata_to_onnx, ensuring the robot runtime can automatically configure observation preprocessing and command interfaces.

  • The 61-dimensional observation vector follows the layout defined in src/mjlab_microduck/tasks/mdp.py, consistent across training, export, and deployment.


Frequently Asked Questions

What ONNX opset version does Microduck RL use?

The export script defaults to opset version 12, which balances compatibility with modern runtimes and support for required PyTorch-to-ONNX conversion operations. This is specified in the export_policy_to_onnx call within scripts/export.py. You can modify this by editing the target_opset_version parameter if your deployment target requires a different version.

Can I export policies without WandB integration?

Yes. Use the --checkpoint-file flag to specify a local checkpoint path directly. The WandB path resolution is optional and only required when retrieving checkpoints stored as WandB artifacts. Local exports work entirely offline once dependencies are installed.

Why is my exported policy producing different actions than during training?

This typically indicates a mismatch in observation preprocessing. Verify that you are passing raw, unnormalized observations—the exported ONNX model includes the training normalizer internally. Check that your observation ordering matches the 61-dimensional layout from src/mjlab_microduck/tasks/mdp.py, and confirm the metadata attachment succeeded by inspecting the ONNX file with Netron or the load_onnx_policy loader.

How large are the exported ONNX models?

The exported files are typically 5–15 MB depending on network architecture and PyTorch version. The MLP-based policies used in Microduck RL's velocity and locomotion tasks compress efficiently due to their relatively compact hidden layer structures (default 512×256×128 in the RSL-RL configuration).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →