# How to Train Microduck RL Policies Using PPO: A Complete Developer Guide

> Learn to train Microduck RL policies with PPO using this developer guide. Export checkpoints to ONNX for real-robot deployment and accelerate your robotics projects.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: tutorial
- Published: 2026-09-02

---

**Use the `train_cli` entry point in [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py) to launch PPO training via `rsl_rl`, then export the trained checkpoint to ONNX with [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) for real-robot deployment.**

Microduck RL is Pollen Robotics' reinforcement learning framework for legged robot locomotion, built on the mjlab simulation stack and the `rsl_rl` PPO trainer. This guide walks through the exact commands, configuration files, and code patterns needed to train policies from scratch, validate them with smoke tests, and prepare them for hardware deployment.

## PPO Training Pipeline Overview

The Microduck training workflow follows three stages: **environment configuration**, **PPO optimization**, and **model export**. The `train_cli:main` function in [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py) orchestrates the entire process.

### Environment Selection

Every locomotion task—walking, rolling, spinning—is defined by a config module under `src/mjlab_microduck/tasks/`. Each config exposes a **task ID** that the trainer recognizes.

Available task IDs include:

- `microduck_velocity` — standard velocity tracking walk
- `microduck_roll` — rolling locomotion
- `microduck_spin` — in-place rotation

List all available environments before training:

```bash
uv run list-envs

```

### Launching PPO Training

The `train` command invokes `mjlab_microduck.train_cli:main`, which forwards configuration to `rsl_rl`'s PPO implementation.

**Quick smoke test** (5 iterations, 64 parallel environments):

```bash
uv run train microduck_velocity --env.scene.num-envs 64 --agent.max_iterations 5

```

This catches ~95% of configuration errors in under a minute.

**Full production training** (default 4096 parallel environments):

```bash
uv run train microduck_velocity --env.scene.num-envs 4096

```

### Logging with Weights & Biases

Enable experiment tracking by passing a run path:

```bash
uv run train microduck_velocity --wandb-run-path myteam/microduck/run-1234

```

## PPO Implementation Details

Microduck's PPO trainer inherits from `rsl_rl` with the following characteristics, as implemented in `pollen-robotics/microduck_rl`:

| Feature | Implementation |
|---------|----------------|
| **Policy clipping** | PPO clip-ratio with KL penalty (rsl_rl defaults) |
| **Advantage estimation** | Generalized Advantage Estimation (GAE) with λ-discount |
| **Curriculum** | Managed via `env.event_manager` (see [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py)) |
| **Domain randomization** | Non-accumulating operations (`dr.*` with `operation="add"` or `"scale"`) |
| **Replay buffer** | None — pure on-policy PPO |

The trainer returns exit code `5` for a clean run, as verified by [`tests/test_hf_jobs_flag.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_hf_jobs_flag.py).

## Critical Configuration Constraints

### Observation Layout

All Microduck environments use a **fixed 61-dimensional observation vector**:

- 48 dimensions: proprioception (joint states, IMU, contacts)
- 13 dimensions: command block (velocity targets, yaw rate, etc.)

New environments must preserve this layout. Zero-pad any missing slots rather than resizing.

### Joint Indexing

Always use the helper utilities in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py):

- `_servo_joint_ids()` — returns valid joint indices
- `_servo_joint_pos()` — returns joint positions

Hard-coding joint indices breaks variants with backlash modeling.

### Domain Randomization Rules

DR operations must be **non-accumulative**. Accumulating randomization degrades performance over long training runs. Use explicit `operation="add"` or `operation="scale"` in all `dr.*` config entries.

## Programmatic Training Examples

Launch training from Python for automated experiments:

```python
from mjlab_microduck.train_cli import main as train_main

# Task ID for velocity tracking walk

TASK_ID = "microduck_velocity"

# Smoke test: 5 iterations, 64 parallel envs

exit_code = train_main([
    "train", TASK_ID,
    "--env.scene.num-envs", "64",
    "--agent.max_iterations", "5"
])

assert exit_code == 5, f"Expected clean exit (5), got {exit_code}"

```

Run multiple seeds in a loop:

```python
import subprocess

for seed in range(3):
    subprocess.run([
        "uv", "run", "train", "microduck_velocity",
        "--env.scene.num-envs", "4096",
        "--seed", str(seed),
        "--wandb-run-path", f"myteam/microduck/seed-{seed}"
    ], check=True)

```

## Exporting Trained Policies to ONNX

The **export step is mandatory**. The observation normalizer must be baked into the ONNX file; otherwise the runtime applies mismatched statistics.

Export a checkpoint after training completes:

```bash
uv run scripts/export.py microduck_velocity --wandb-run-path myteam/microduck/run-1234

```

Or invoke programmatically:

```python
import subprocess

subprocess.run([
    "uv", "run", "scripts/export.py",
    "microduck_velocity",
    "--wandb-run-path", "myteam/microduck/run-1234"
], check=True)

```

The output `out.onnx` contains both the policy network and the frozen observation normalizer.

## Visualizing and Validating Policies

Run a trained ONNX policy in simulation:

```bash
uv run scripts/infer_policy.py --walking out.onnx

```

This loads the policy into a MuJoCo rollout for visual verification before hardware deployment.

## Key Source Files

| File | Purpose | Location |
|------|---------|----------|
| [`train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/train_cli.py) | Main CLI entry point launching `rsl_rl` PPO | [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py) |
| `*_env_cfg.py` | Per-task environment configurations | `src/mjlab_microduck/tasks/` |
| [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) | MDP utilities: rewards, joint helpers, curriculum | [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) |
| [`export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/export.py) | ONNX export with baked normalizer | [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) |
| [`infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/infer_policy.py) | Simulated rollout of ONNX policies | [`scripts/infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/infer_policy.py) |
| [`AGENTS.md`](https://github.com/pollen-robotics/microduck_rl/blob/main/AGENTS.md) | Complete command reference and design invariants | Repository root |

## Summary

- **Select environment**: Use `uv run list-envs` to find task IDs, then reference the corresponding `*_env_cfg.py` file.
- **Smoke test first**: Run 5 iterations with 64 environments to catch config errors instantly.
- **Scale to production**: Train with 4096 parallel environments for full policy quality.
- **Preserve obs layout**: Maintain the 61-dimensional observation vector; zero-pad if needed.
- **Use joint helpers**: Call `_servo_joint_ids()` and `_servo_joint_pos()` from [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py), never hard-code indices.
- **Enforce non-accumulative DR**: Set explicit `operation` parameters in all domain randomization configs.
- **Export with normalizer**: Always run [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) before deployment; the normalizer is baked into the ONNX file.

## Frequently Asked Questions

### What exit code indicates successful PPO training in Microduck RL?

Exit code `5` signals a clean training run. This convention is enforced by the unit test in [`tests/test_hf_jobs_flag.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_hf_jobs_flag.py) and returned by `mjlab_microduck.train_cli:main` upon completion.

### Why must I use [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) instead of directly loading the PyTorch checkpoint?

The export script bakes the observation normalizer into the ONNX graph. Loading a raw `.pt` checkpoint without the matching normalizer statistics causes severe policy degradation on hardware. The ONNX file encapsulates both network weights and frozen normalization parameters.

### How do I add a new locomotion task for PPO training?

Create a new `*_env_cfg.py` file in `src/mjlab_microduck/tasks/` that defines your task's MDP configuration, reward terms, and termination conditions. Register the task ID in the environment registry so `list-envs` discovers it. Ensure your observations conform to the 61-dimensional layout and use [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) helpers for joint indexing.

### What is the recommended domain randomization strategy for stable PPO training?

Use non-accumulating operations exclusively. Configure all `dr.*` parameters with explicit `operation="add"` or `operation="scale"` to prevent randomization parameters from compounding over episodes. Accumulating DR causes performance collapse in long training runs, as documented in [`AGENTS.md`](https://github.com/pollen-robotics/microduck_rl/blob/main/AGENTS.md).