How to Train Microduck RL Policies Using PPO: A Complete Developer Guide
Use the train_cli entry point in src/mjlab_microduck/train_cli.py to launch PPO training via rsl_rl, then export the trained checkpoint to ONNX with scripts/export.py for real-robot deployment.
Microduck RL is Pollen Robotics' reinforcement learning framework for legged robot locomotion, built on the mjlab simulation stack and the rsl_rl PPO trainer. This guide walks through the exact commands, configuration files, and code patterns needed to train policies from scratch, validate them with smoke tests, and prepare them for hardware deployment.
PPO Training Pipeline Overview
The Microduck training workflow follows three stages: environment configuration, PPO optimization, and model export. The train_cli:main function in src/mjlab_microduck/train_cli.py orchestrates the entire process.
Environment Selection
Every locomotion task—walking, rolling, spinning—is defined by a config module under src/mjlab_microduck/tasks/. Each config exposes a task ID that the trainer recognizes.
Available task IDs include:
microduck_velocity— standard velocity tracking walkmicroduck_roll— rolling locomotionmicroduck_spin— in-place rotation
List all available environments before training:
uv run list-envs
Launching PPO Training
The train command invokes mjlab_microduck.train_cli:main, which forwards configuration to rsl_rl's PPO implementation.
Quick smoke test (5 iterations, 64 parallel environments):
uv run train microduck_velocity --env.scene.num-envs 64 --agent.max_iterations 5
This catches ~95% of configuration errors in under a minute.
Full production training (default 4096 parallel environments):
uv run train microduck_velocity --env.scene.num-envs 4096
Logging with Weights & Biases
Enable experiment tracking by passing a run path:
uv run train microduck_velocity --wandb-run-path myteam/microduck/run-1234
PPO Implementation Details
Microduck's PPO trainer inherits from rsl_rl with the following characteristics, as implemented in pollen-robotics/microduck_rl:
| Feature | Implementation |
|---|---|
| Policy clipping | PPO clip-ratio with KL penalty (rsl_rl defaults) |
| Advantage estimation | Generalized Advantage Estimation (GAE) with λ-discount |
| Curriculum | Managed via env.event_manager (see src/mjlab_microduck/tasks/mdp.py) |
| Domain randomization | Non-accumulating operations (dr.* with operation="add" or "scale") |
| Replay buffer | None — pure on-policy PPO |
The trainer returns exit code 5 for a clean run, as verified by tests/test_hf_jobs_flag.py.
Critical Configuration Constraints
Observation Layout
All Microduck environments use a fixed 61-dimensional observation vector:
- 48 dimensions: proprioception (joint states, IMU, contacts)
- 13 dimensions: command block (velocity targets, yaw rate, etc.)
New environments must preserve this layout. Zero-pad any missing slots rather than resizing.
Joint Indexing
Always use the helper utilities in src/mjlab_microduck/tasks/mdp.py:
_servo_joint_ids()— returns valid joint indices_servo_joint_pos()— returns joint positions
Hard-coding joint indices breaks variants with backlash modeling.
Domain Randomization Rules
DR operations must be non-accumulative. Accumulating randomization degrades performance over long training runs. Use explicit operation="add" or operation="scale" in all dr.* config entries.
Programmatic Training Examples
Launch training from Python for automated experiments:
from mjlab_microduck.train_cli import main as train_main
# Task ID for velocity tracking walk
TASK_ID = "microduck_velocity"
# Smoke test: 5 iterations, 64 parallel envs
exit_code = train_main([
"train", TASK_ID,
"--env.scene.num-envs", "64",
"--agent.max_iterations", "5"
])
assert exit_code == 5, f"Expected clean exit (5), got {exit_code}"
Run multiple seeds in a loop:
import subprocess
for seed in range(3):
subprocess.run([
"uv", "run", "train", "microduck_velocity",
"--env.scene.num-envs", "4096",
"--seed", str(seed),
"--wandb-run-path", f"myteam/microduck/seed-{seed}"
], check=True)
Exporting Trained Policies to ONNX
The export step is mandatory. The observation normalizer must be baked into the ONNX file; otherwise the runtime applies mismatched statistics.
Export a checkpoint after training completes:
uv run scripts/export.py microduck_velocity --wandb-run-path myteam/microduck/run-1234
Or invoke programmatically:
import subprocess
subprocess.run([
"uv", "run", "scripts/export.py",
"microduck_velocity",
"--wandb-run-path", "myteam/microduck/run-1234"
], check=True)
The output out.onnx contains both the policy network and the frozen observation normalizer.
Visualizing and Validating Policies
Run a trained ONNX policy in simulation:
uv run scripts/infer_policy.py --walking out.onnx
This loads the policy into a MuJoCo rollout for visual verification before hardware deployment.
Key Source Files
| File | Purpose | Location |
|---|---|---|
train_cli.py |
Main CLI entry point launching rsl_rl PPO |
src/mjlab_microduck/train_cli.py |
*_env_cfg.py |
Per-task environment configurations | src/mjlab_microduck/tasks/ |
mdp.py |
MDP utilities: rewards, joint helpers, curriculum | src/mjlab_microduck/tasks/mdp.py |
export.py |
ONNX export with baked normalizer | scripts/export.py |
infer_policy.py |
Simulated rollout of ONNX policies | scripts/infer_policy.py |
AGENTS.md |
Complete command reference and design invariants | Repository root |
Summary
- Select environment: Use
uv run list-envsto find task IDs, then reference the corresponding*_env_cfg.pyfile. - Smoke test first: Run 5 iterations with 64 environments to catch config errors instantly.
- Scale to production: Train with 4096 parallel environments for full policy quality.
- Preserve obs layout: Maintain the 61-dimensional observation vector; zero-pad if needed.
- Use joint helpers: Call
_servo_joint_ids()and_servo_joint_pos()frommdp.py, never hard-code indices. - Enforce non-accumulative DR: Set explicit
operationparameters in all domain randomization configs. - Export with normalizer: Always run
scripts/export.pybefore deployment; the normalizer is baked into the ONNX file.
Frequently Asked Questions
What exit code indicates successful PPO training in Microduck RL?
Exit code 5 signals a clean training run. This convention is enforced by the unit test in tests/test_hf_jobs_flag.py and returned by mjlab_microduck.train_cli:main upon completion.
Why must I use scripts/export.py instead of directly loading the PyTorch checkpoint?
The export script bakes the observation normalizer into the ONNX graph. Loading a raw .pt checkpoint without the matching normalizer statistics causes severe policy degradation on hardware. The ONNX file encapsulates both network weights and frozen normalization parameters.
How do I add a new locomotion task for PPO training?
Create a new *_env_cfg.py file in src/mjlab_microduck/tasks/ that defines your task's MDP configuration, reward terms, and termination conditions. Register the task ID in the environment registry so list-envs discovers it. Ensure your observations conform to the 61-dimensional layout and use mdp.py helpers for joint indexing.
What is the recommended domain randomization strategy for stable PPO training?
Use non-accumulating operations exclusively. Configure all dr.* parameters with explicit operation="add" or operation="scale" to prevent randomization parameters from compounding over episodes. Accumulating DR causes performance collapse in long training runs, as documented in AGENTS.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →