How to Train Microduck RL Policies Using PPO: A Complete Developer Guide

Use the train_cli entry point in src/mjlab_microduck/train_cli.py to launch PPO training via rsl_rl, then export the trained checkpoint to ONNX with scripts/export.py for real-robot deployment.

Microduck RL is Pollen Robotics' reinforcement learning framework for legged robot locomotion, built on the mjlab simulation stack and the rsl_rl PPO trainer. This guide walks through the exact commands, configuration files, and code patterns needed to train policies from scratch, validate them with smoke tests, and prepare them for hardware deployment.

PPO Training Pipeline Overview

The Microduck training workflow follows three stages: environment configuration, PPO optimization, and model export. The train_cli:main function in src/mjlab_microduck/train_cli.py orchestrates the entire process.

Environment Selection

Every locomotion task—walking, rolling, spinning—is defined by a config module under src/mjlab_microduck/tasks/. Each config exposes a task ID that the trainer recognizes.

Available task IDs include:

  • microduck_velocity — standard velocity tracking walk
  • microduck_roll — rolling locomotion
  • microduck_spin — in-place rotation

List all available environments before training:

uv run list-envs

Launching PPO Training

The train command invokes mjlab_microduck.train_cli:main, which forwards configuration to rsl_rl's PPO implementation.

Quick smoke test (5 iterations, 64 parallel environments):

uv run train microduck_velocity --env.scene.num-envs 64 --agent.max_iterations 5

This catches ~95% of configuration errors in under a minute.

Full production training (default 4096 parallel environments):

uv run train microduck_velocity --env.scene.num-envs 4096

Logging with Weights & Biases

Enable experiment tracking by passing a run path:

uv run train microduck_velocity --wandb-run-path myteam/microduck/run-1234

PPO Implementation Details

Microduck's PPO trainer inherits from rsl_rl with the following characteristics, as implemented in pollen-robotics/microduck_rl:

Feature Implementation
Policy clipping PPO clip-ratio with KL penalty (rsl_rl defaults)
Advantage estimation Generalized Advantage Estimation (GAE) with λ-discount
Curriculum Managed via env.event_manager (see src/mjlab_microduck/tasks/mdp.py)
Domain randomization Non-accumulating operations (dr.* with operation="add" or "scale")
Replay buffer None — pure on-policy PPO

The trainer returns exit code 5 for a clean run, as verified by tests/test_hf_jobs_flag.py.

Critical Configuration Constraints

Observation Layout

All Microduck environments use a fixed 61-dimensional observation vector:

  • 48 dimensions: proprioception (joint states, IMU, contacts)
  • 13 dimensions: command block (velocity targets, yaw rate, etc.)

New environments must preserve this layout. Zero-pad any missing slots rather than resizing.

Joint Indexing

Always use the helper utilities in src/mjlab_microduck/tasks/mdp.py:

  • _servo_joint_ids() — returns valid joint indices
  • _servo_joint_pos() — returns joint positions

Hard-coding joint indices breaks variants with backlash modeling.

Domain Randomization Rules

DR operations must be non-accumulative. Accumulating randomization degrades performance over long training runs. Use explicit operation="add" or operation="scale" in all dr.* config entries.

Programmatic Training Examples

Launch training from Python for automated experiments:

from mjlab_microduck.train_cli import main as train_main

# Task ID for velocity tracking walk

TASK_ID = "microduck_velocity"

# Smoke test: 5 iterations, 64 parallel envs

exit_code = train_main([
    "train", TASK_ID,
    "--env.scene.num-envs", "64",
    "--agent.max_iterations", "5"
])

assert exit_code == 5, f"Expected clean exit (5), got {exit_code}"

Run multiple seeds in a loop:

import subprocess

for seed in range(3):
    subprocess.run([
        "uv", "run", "train", "microduck_velocity",
        "--env.scene.num-envs", "4096",
        "--seed", str(seed),
        "--wandb-run-path", f"myteam/microduck/seed-{seed}"
    ], check=True)

Exporting Trained Policies to ONNX

The export step is mandatory. The observation normalizer must be baked into the ONNX file; otherwise the runtime applies mismatched statistics.

Export a checkpoint after training completes:

uv run scripts/export.py microduck_velocity --wandb-run-path myteam/microduck/run-1234

Or invoke programmatically:

import subprocess

subprocess.run([
    "uv", "run", "scripts/export.py",
    "microduck_velocity",
    "--wandb-run-path", "myteam/microduck/run-1234"
], check=True)

The output out.onnx contains both the policy network and the frozen observation normalizer.

Visualizing and Validating Policies

Run a trained ONNX policy in simulation:

uv run scripts/infer_policy.py --walking out.onnx

This loads the policy into a MuJoCo rollout for visual verification before hardware deployment.

Key Source Files

File Purpose Location
train_cli.py Main CLI entry point launching rsl_rl PPO src/mjlab_microduck/train_cli.py
*_env_cfg.py Per-task environment configurations src/mjlab_microduck/tasks/
mdp.py MDP utilities: rewards, joint helpers, curriculum src/mjlab_microduck/tasks/mdp.py
export.py ONNX export with baked normalizer scripts/export.py
infer_policy.py Simulated rollout of ONNX policies scripts/infer_policy.py
AGENTS.md Complete command reference and design invariants Repository root

Summary

  • Select environment: Use uv run list-envs to find task IDs, then reference the corresponding *_env_cfg.py file.
  • Smoke test first: Run 5 iterations with 64 environments to catch config errors instantly.
  • Scale to production: Train with 4096 parallel environments for full policy quality.
  • Preserve obs layout: Maintain the 61-dimensional observation vector; zero-pad if needed.
  • Use joint helpers: Call _servo_joint_ids() and _servo_joint_pos() from mdp.py, never hard-code indices.
  • Enforce non-accumulative DR: Set explicit operation parameters in all domain randomization configs.
  • Export with normalizer: Always run scripts/export.py before deployment; the normalizer is baked into the ONNX file.

Frequently Asked Questions

What exit code indicates successful PPO training in Microduck RL?

Exit code 5 signals a clean training run. This convention is enforced by the unit test in tests/test_hf_jobs_flag.py and returned by mjlab_microduck.train_cli:main upon completion.

Why must I use scripts/export.py instead of directly loading the PyTorch checkpoint?

The export script bakes the observation normalizer into the ONNX graph. Loading a raw .pt checkpoint without the matching normalizer statistics causes severe policy degradation on hardware. The ONNX file encapsulates both network weights and frozen normalization parameters.

How do I add a new locomotion task for PPO training?

Create a new *_env_cfg.py file in src/mjlab_microduck/tasks/ that defines your task's MDP configuration, reward terms, and termination conditions. Register the task ID in the environment registry so list-envs discovers it. Ensure your observations conform to the 61-dimensional layout and use mdp.py helpers for joint indexing.

Use non-accumulating operations exclusively. Configure all dr.* parameters with explicit operation="add" or operation="scale" to prevent randomization parameters from compounding over episodes. Accumulating DR causes performance collapse in long training runs, as documented in AGENTS.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →