How to Export Trained Policies From Microduck RL to ONNX: Complete Guide
Microduck RL provides a dedicated scripts/export.py command-line tool that converts trained RSL-RL policies into portable ONNX models with baked-in observation normalization and embedded environment metadata.
Exporting trained reinforcement learning policies to ONNX enables seamless deployment on physical robots without requiring the full training stack. Pollen Robotics' Microduck RL repository implements a streamlined export pipeline through the export.py script in the scripts/ directory. This article walks through the complete export process, from checkpoint resolution to metadata embedding, based on the actual source code implementation in pollen-robotics/microduck_rl.
Prerequisites and Environment Setup
Before exporting, ensure your environment matches the repository's locked dependencies. The pyproject.toml specifies exact versions for PyTorch, RSL-RL, and the MJLab simulation framework.
# Install all runtime dependencies
uv sync
The export script requires either:
- A local checkpoint file (
.ptor.pth), or - A WandB run path with stored checkpoints
The Export Script: scripts/export.py
The core export functionality lives in scripts/export.py. This script implements a six-step pipeline in its run_export function:
Step 1: Load Environment and Agent Configurations
The script begins by loading task-specific configurations through two helper functions:
load_env_cfg(line 51): Loads the MJLab environment configuration for the specified taskload_rl_cfg(line 52): Loads the RSL-RL agent hyperparameters and network architecture
These configurations ensure the exported model matches the training environment's observation and action spaces.
Step 2: Resolve the Checkpoint Source
Checkpoint handling spans lines 27–71 and supports two resolution paths:
| Source | Flag | Behavior |
|---|---|---|
| Local file | --checkpoint-file |
Load .pt file directly from disk |
| WandB artifact | --wandb-run-path |
Download checkpoint via WandB API by iteration number or latest |
If using WandB, the script authenticates via the wandb Python API and retrieves the requested checkpoint artifact. The --checkpoint flag specifies which iteration to fetch when multiple checkpoints exist.
Step 3: Initialize the Policy and Load Weights
With configurations loaded, the script instantiates an OnPolicyRunner (RSL-RL's standard training orchestrator) and restores the policy state:
# From scripts/export.py lines 22-26
runner.load(checkpoint_path) # Restores actor-critic weights
policy = runner.get_inference_policy() # Returns wrapped inference-only policy
The get_inference_policy() method returns a wrapped policy optimized for inference: it disables gradient computation and applies any training-time wrappers (such as action clipping or observation filtering).
Step 4: Export to ONNX Format
The actual ONNX conversion occurs through the runner's built-in method:
# From scripts/export.py lines 226-239
runner.export_policy_to_onnx(
policy,
onnx_file_path,
target_opset_version=12 # Configurable for compatibility
)
Critical detail: export_policy_to_onnx automatically bakes the observation normalizer into the computation graph. During training, RSL-RL maintains running statistics for observation normalization; these statistics are fused into the ONNX model as constant scale and bias operations. The exported model therefore expects raw, unnormalized observations as input—no preprocessing required at inference time.
Step 5: Attach Environment Metadata
After the raw ONNX file is written, the script embeds additional context for runtime consumption (lines 40–42):
from mjlab.rl.exporter_utils import get_base_metadata, attach_metadata_to_onnx
metadata = get_base_metadata(env_cfg) # Extract obs layout, command slots, etc.
attach_metadata_to_onnx(onnx_file_path, metadata)
This metadata includes:
- Observation dimension and semantic layout (61-dimensional vector for Microduck tasks)
- Command structure and valid ranges
- Environment-specific constants from the task definition
Storing metadata inside the ONNX file eliminates the need for sidecar configuration files during robot deployment.
Step 6: Finalization
The script prints the absolute path to the generated ONNX file and gracefully shuts down the MJLab environment.
Complete Export Examples
Export From WandB Checkpoint
# Export iteration 3000 from a completed WandB training run
uv run scripts/export.py velocity \
--onnx-file policies/velocity_3000.onnx \
--wandb-run-path pollen-robotics/microduck-rl/1a2b3c4d5e6f7g8h \
--checkpoint 3000
Export From Local Checkpoint
# Export from a manually saved checkpoint file
uv run scripts/export.py velocity \
--onnx-file policies/velocity_local.onnx \
--checkpoint-file /path/to/checkpoint_3000.pt
Export Latest Checkpoint Automatically
# Omit --checkpoint to use the highest iteration number available
uv run scripts/export.py velocity \
--onnx-file policies/velocity_latest.onnx \
--wandb-run-path pollen-robotics/microduck-rl/1a2b3c4d5e6f7g8h
Loading and Running the Exported ONNX Model
The exported ONNX file is self-contained and can be loaded without the training dependencies. The mjlab.rl.exporter_utils module provides load_onnx_policy for convenient inference:
from mjlab.utils.torch import configure_torch_backends
from mjlab.rl.exporter_utils import load_onnx_policy
import numpy as np
# Optional: optimize PyTorch backend for inference
configure_torch_backends()
# Load the exported policy
policy = load_onnx_policy("policies/velocity_3000.onnx")
# Prepare observation: shape (batch, 61) for Microduck velocity task
# The 61 dimensions comprise: base angular velocity (3), projected gravity (3),
# joint positions (12), joint velocities (12), previous actions (12),
# and command inputs (19) — see src/mjlab_microduck/tasks/mdp.py
obs = np.zeros((1, 61), dtype=np.float32)
# Run inference: returns torch.Tensor of shape (batch, 12) for joint targets
action = policy(obs)
print(action.shape) # torch.Size([1, 12])
The observation layout is defined in src/mjlab_microduck/tasks/mdp.py and matches the 61-dimensional structure used during training. Because normalization is baked into the ONNX graph, pass raw sensor values directly.
Integration With Robot Runtime
The exported ONNX model integrates directly with the physical robot stack in pollen-robotics/microduck. The scripts/infer_policy.py example demonstrates loading the policy in a simulated environment before hardware deployment:
# Test the exported policy in simulation
uv run scripts/infer_policy.py \
--onnx-file policies/velocity_3000.onnx \
--env velocity \
--num-envs 1
For production deployment, the same ONNX file loads into the robot's real-time control loop via ONNX Runtime or TensorRT, with the embedded metadata providing automatic configuration.
Key Source Files Reference
| File | Purpose | Lines of Interest |
|---|---|---|
scripts/export.py |
Main export script with run_export implementation |
1–250 |
src/mjlab_microduck/train_cli.py |
Training entry point producing exportable checkpoints | Full file |
src/mjlab_microduck/tasks/mdp.py |
Observation space definition (61-dim) | Full file |
scripts/infer_policy.py |
Example ONNX inference loader | Full file |
src/mjlab_microduck/rl/exporter_utils.py |
Metadata extraction and attachment utilities | get_base_metadata, attach_metadata_to_onnx |
Summary
-
Microduck RL exports to ONNX via
scripts/export.py, a dedicated command-line tool that wraps RSL-RL's native export functionality with environment-specific metadata handling. -
Checkpoint resolution supports both local files and WandB artifacts, enabling export from finished training runs or intermediate iterations.
-
Observation normalization is automatically baked into the ONNX graph via
export_policy_to_onnx, so the exported model accepts raw sensor inputs directly. -
Environment metadata is embedded in the ONNX file through
attach_metadata_to_onnx, ensuring the robot runtime can automatically configure observation preprocessing and command interfaces. -
The 61-dimensional observation vector follows the layout defined in
src/mjlab_microduck/tasks/mdp.py, consistent across training, export, and deployment.
Frequently Asked Questions
What ONNX opset version does Microduck RL use?
The export script defaults to opset version 12, which balances compatibility with modern runtimes and support for required PyTorch-to-ONNX conversion operations. This is specified in the export_policy_to_onnx call within scripts/export.py. You can modify this by editing the target_opset_version parameter if your deployment target requires a different version.
Can I export policies without WandB integration?
Yes. Use the --checkpoint-file flag to specify a local checkpoint path directly. The WandB path resolution is optional and only required when retrieving checkpoints stored as WandB artifacts. Local exports work entirely offline once dependencies are installed.
Why is my exported policy producing different actions than during training?
This typically indicates a mismatch in observation preprocessing. Verify that you are passing raw, unnormalized observations—the exported ONNX model includes the training normalizer internally. Check that your observation ordering matches the 61-dimensional layout from src/mjlab_microduck/tasks/mdp.py, and confirm the metadata attachment succeeded by inspecting the ONNX file with Netron or the load_onnx_policy loader.
How large are the exported ONNX models?
The exported files are typically 5–15 MB depending on network architecture and PyTorch version. The MLP-based policies used in Microduck RL's velocity and locomotion tasks compress efficiently due to their relatively compact hidden layer structures (default 512×256×128 in the RSL-RL configuration).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →