# How to Export Trained Policies From Microduck RL to ONNX: Complete Guide

> Easily export trained Microduck RL policies to ONNX using the dedicated export script. Convert RSL-RL policies into portable ONNX models with baked-in normalization and metadata.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: how-to-guide
- Published: 2026-09-02

---

**Microduck RL provides a dedicated [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) command-line tool that converts trained RSL-RL policies into portable ONNX models with baked-in observation normalization and embedded environment metadata.**

Exporting trained reinforcement learning policies to ONNX enables seamless deployment on physical robots without requiring the full training stack. Pollen Robotics' Microduck RL repository implements a streamlined export pipeline through the [`export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/export.py) script in the `scripts/` directory. This article walks through the complete export process, from checkpoint resolution to metadata embedding, based on the actual source code implementation in `pollen-robotics/microduck_rl`.

---

## Prerequisites and Environment Setup

Before exporting, ensure your environment matches the repository's locked dependencies. The [`pyproject.toml`](https://github.com/pollen-robotics/microduck_rl/blob/main/pyproject.toml) specifies exact versions for PyTorch, RSL-RL, and the MJLab simulation framework.

```bash

# Install all runtime dependencies

uv sync

```

The export script requires either:
- A **local checkpoint file** (`.pt` or `.pth`), or
- A **WandB run path** with stored checkpoints

---

## The Export Script: [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py)

The core export functionality lives in [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py). This script implements a six-step pipeline in its `run_export` function:

### Step 1: Load Environment and Agent Configurations

The script begins by loading task-specific configurations through two helper functions:

- **`load_env_cfg`** (line 51): Loads the MJLab environment configuration for the specified task
- **`load_rl_cfg`** (line 52): Loads the RSL-RL agent hyperparameters and network architecture

These configurations ensure the exported model matches the training environment's observation and action spaces.

### Step 2: Resolve the Checkpoint Source

Checkpoint handling spans lines 27–71 and supports two resolution paths:

| Source | Flag | Behavior |
|--------|------|----------|
| Local file | `--checkpoint-file` | Load `.pt` file directly from disk |
| WandB artifact | `--wandb-run-path` | Download checkpoint via WandB API by iteration number or latest |

If using WandB, the script authenticates via the `wandb` Python API and retrieves the requested checkpoint artifact. The `--checkpoint` flag specifies which iteration to fetch when multiple checkpoints exist.

### Step 3: Initialize the Policy and Load Weights

With configurations loaded, the script instantiates an `OnPolicyRunner` (RSL-RL's standard training orchestrator) and restores the policy state:

```python

# From scripts/export.py lines 22-26

runner.load(checkpoint_path)           # Restores actor-critic weights

policy = runner.get_inference_policy() # Returns wrapped inference-only policy

```

The `get_inference_policy()` method returns a wrapped policy optimized for inference: it disables gradient computation and applies any training-time wrappers (such as action clipping or observation filtering).

### Step 4: Export to ONNX Format

The actual ONNX conversion occurs through the runner's built-in method:

```python

# From scripts/export.py lines 226-239

runner.export_policy_to_onnx(
    policy, 
    onnx_file_path, 
    target_opset_version=12  # Configurable for compatibility

)

```

**Critical detail:** `export_policy_to_onnx` automatically bakes the observation normalizer into the computation graph. During training, RSL-RL maintains running statistics for observation normalization; these statistics are fused into the ONNX model as constant scale and bias operations. The exported model therefore expects **raw, unnormalized observations** as input—no preprocessing required at inference time.

### Step 5: Attach Environment Metadata

After the raw ONNX file is written, the script embeds additional context for runtime consumption (lines 40–42):

```python
from mjlab.rl.exporter_utils import get_base_metadata, attach_metadata_to_onnx

metadata = get_base_metadata(env_cfg)      # Extract obs layout, command slots, etc.

attach_metadata_to_onnx(onnx_file_path, metadata)

```

This metadata includes:
- Observation dimension and semantic layout (61-dimensional vector for Microduck tasks)
- Command structure and valid ranges
- Environment-specific constants from the task definition

Storing metadata inside the ONNX file eliminates the need for sidecar configuration files during robot deployment.

### Step 6: Finalization

The script prints the absolute path to the generated ONNX file and gracefully shuts down the MJLab environment.

---

## Complete Export Examples

### Export From WandB Checkpoint

```bash

# Export iteration 3000 from a completed WandB training run

uv run scripts/export.py velocity \
    --onnx-file policies/velocity_3000.onnx \
    --wandb-run-path pollen-robotics/microduck-rl/1a2b3c4d5e6f7g8h \
    --checkpoint 3000

```

### Export From Local Checkpoint

```bash

# Export from a manually saved checkpoint file

uv run scripts/export.py velocity \
    --onnx-file policies/velocity_local.onnx \
    --checkpoint-file /path/to/checkpoint_3000.pt

```

### Export Latest Checkpoint Automatically

```bash

# Omit --checkpoint to use the highest iteration number available

uv run scripts/export.py velocity \
    --onnx-file policies/velocity_latest.onnx \
    --wandb-run-path pollen-robotics/microduck-rl/1a2b3c4d5e6f7g8h

```

---

## Loading and Running the Exported ONNX Model

The exported ONNX file is self-contained and can be loaded without the training dependencies. The `mjlab.rl.exporter_utils` module provides `load_onnx_policy` for convenient inference:

```python
from mjlab.utils.torch import configure_torch_backends
from mjlab.rl.exporter_utils import load_onnx_policy
import numpy as np

# Optional: optimize PyTorch backend for inference

configure_torch_backends()

# Load the exported policy

policy = load_onnx_policy("policies/velocity_3000.onnx")

# Prepare observation: shape (batch, 61) for Microduck velocity task

# The 61 dimensions comprise: base angular velocity (3), projected gravity (3),

# joint positions (12), joint velocities (12), previous actions (12),

# and command inputs (19) — see src/mjlab_microduck/tasks/mdp.py

obs = np.zeros((1, 61), dtype=np.float32)

# Run inference: returns torch.Tensor of shape (batch, 12) for joint targets

action = policy(obs)
print(action.shape)  # torch.Size([1, 12])

```

The observation layout is defined in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) and matches the 61-dimensional structure used during training. Because normalization is baked into the ONNX graph, pass raw sensor values directly.

---

## Integration With Robot Runtime

The exported ONNX model integrates directly with the physical robot stack in `pollen-robotics/microduck`. The [`scripts/infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/infer_policy.py) example demonstrates loading the policy in a simulated environment before hardware deployment:

```bash

# Test the exported policy in simulation

uv run scripts/infer_policy.py \
    --onnx-file policies/velocity_3000.onnx \
    --env velocity \
    --num-envs 1

```

For production deployment, the same ONNX file loads into the robot's real-time control loop via ONNX Runtime or TensorRT, with the embedded metadata providing automatic configuration.

---

## Key Source Files Reference

| File | Purpose | Lines of Interest |
|------|---------|-------------------|
| [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) | Main export script with `run_export` implementation | 1–250 |
| [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py) | Training entry point producing exportable checkpoints | Full file |
| [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) | Observation space definition (61-dim) | Full file |
| [`scripts/infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/infer_policy.py) | Example ONNX inference loader | Full file |
| [`src/mjlab_microduck/rl/exporter_utils.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/rl/exporter_utils.py) | Metadata extraction and attachment utilities | `get_base_metadata`, `attach_metadata_to_onnx` |

---

## Summary

- **Microduck RL exports to ONNX via [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py)**, a dedicated command-line tool that wraps RSL-RL's native export functionality with environment-specific metadata handling.

- **Checkpoint resolution supports both local files and WandB artifacts**, enabling export from finished training runs or intermediate iterations.

- **Observation normalization is automatically baked into the ONNX graph** via `export_policy_to_onnx`, so the exported model accepts raw sensor inputs directly.

- **Environment metadata is embedded in the ONNX file** through `attach_metadata_to_onnx`, ensuring the robot runtime can automatically configure observation preprocessing and command interfaces.

- **The 61-dimensional observation vector** follows the layout defined in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py), consistent across training, export, and deployment.

---

## Frequently Asked Questions

### What ONNX opset version does Microduck RL use?

The export script defaults to opset version 12, which balances compatibility with modern runtimes and support for required PyTorch-to-ONNX conversion operations. This is specified in the `export_policy_to_onnx` call within [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py). You can modify this by editing the `target_opset_version` parameter if your deployment target requires a different version.

### Can I export policies without WandB integration?

Yes. Use the `--checkpoint-file` flag to specify a local checkpoint path directly. The WandB path resolution is optional and only required when retrieving checkpoints stored as WandB artifacts. Local exports work entirely offline once dependencies are installed.

### Why is my exported policy producing different actions than during training?

This typically indicates a mismatch in observation preprocessing. Verify that you are passing raw, unnormalized observations—the exported ONNX model includes the training normalizer internally. Check that your observation ordering matches the 61-dimensional layout from [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py), and confirm the metadata attachment succeeded by inspecting the ONNX file with Netron or the `load_onnx_policy` loader.

### How large are the exported ONNX models?

The exported files are typically 5–15 MB depending on network architecture and PyTorch version. The MLP-based policies used in Microduck RL's velocity and locomotion tasks compress efficiently due to their relatively compact hidden layer structures (default 512×256×128 in the RSL-RL configuration).