How to Train a Reinforcement Learning Policy for the Microduck Bipedal Robot: A Complete PPO Pipeline

The microduck_rl repository provides a complete PPO-based training pipeline built on mjlab (MuJoCo Warp) that enables you to train, export, and deploy policies for the 800g, 14-servo Microduck bipedal robot using configurable velocity-tracking environments and domain randomization.

The pollen-robotics/microduck_rl repository offers a production-ready framework to train a reinforcement learning policy for the Microduck bipedal robot. Built atop mjlab (MuJoCo Warp) and the rsl_rl PPO implementation, this pipeline handles everything from environment configuration to ONNX export for real hardware deployment.

Training Architecture Overview

The system implements a modular architecture centered on the ManagerBasedRlEnvCfg class. In src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py, the function make_microduck_velocity_env_cfg wires together the robot model, terrain, observations, commands, and rewards into a cohesive training environment.

Robot Model Configuration

The MICRODUCK_WALK_ROBOT_CFG in src/mjlab_microduck/robot/microduck_constants.py defines the MJCF specification for the 14-servo robot. This configuration includes:

  • BAM XL330 actuator specifications
  • Passive joint definitions (filtered from observations)
  • Joint limits and sensor configurations

MDP Functions and Rewards

Custom reward terms reside in src/mjlab_microduck/tasks/mdp.py. The implementation provides NaN-safe computation patches and specialized reward functions including:

  • track_linear_velocity and track_angular_velocity for command following
  • head_pose_tracking for gaze direction
  • upright and upright_progress for stability
  • foot_slip (weak penalty) and action_rate_l2 for regularization

Configuring Observations and Actions

The Microduck environment uses a 61-dimensional observation space consisting of 48-dimensional proprioception data plus a 13-dimensional command block. The action space spans 14 dimensions, representing joint-position commands for each servo.

Domain Randomization

Robust sim-to-real transfer relies on extensive domain randomization configured via environment flags in microduck_velocity_env_cfg.py:

  • CoM randomization: Center-of-mass shifts per environment (ENABLE_COM_RANDOMIZATION)
  • Joint friction scaling: Dynamic friction coefficient variation
  • Motor gain scaling: Actuator response randomization
  • Encoder bias: Per-environment constant offsets added to observations
  • IMU orientation randomization: Perturbation of inertial measurement unit readings

Curriculum Learning

Training stability is enhanced through CurriculumTermCfg schedules defined within the environment configuration. These step-wise schedules gradually adjust:

  • Action-rate penalty weights
  • Command velocity ranges
  • Randomization magnitudes

Executing Training Runs

The CLI entry point uv run train <TASK_ID> bootstraps the environment, applies the curriculum, and initiates the PPO loop with trajectory collection, advantage computation, and policy updates.

Smoke Testing Configurations

Before committing to long training runs, validate your configuration with a minimal 5-iteration test:

uv run train velocity --env.scene.num-envs 64 --agent.max_iterations 5

This command creates 64 parallel velocity-tracking environments and stops after five PPO iterations, catching configuration errors early.

Production Training

For full-scale training, allocate sufficient parallel environments and enable logging:

uv run train velocity \
  --env.scene.num-envs 4096 \
  --wandb-run-path myuser/microduck-velocity \
  --hf-jobs

Typical production settings use 4,096 parallel environments, log metrics to Weights & Biases, and optionally upload checkpoints to Hugging Face via the --hf-jobs flag.

Exporting and Deployment

Once training converges, export the policy for hardware deployment using the ONNX exporter.

ONNX Export with Observation Normalization

The export script bakes the observation normalizer directly into the graph and filters passive joints:

uv run scripts/export.py velocity \
  --wandb-run-path myuser/microduck-velocity/run123 \
  --output out.onnx

CPU Inference Testing

Verify the exported policy runs correctly without GPU acceleration:

uv run scripts/infer_policy.py --walking out.onnx

For pure simulation without BAM actuator models, append the --no-bam flag to use the XML PD controller instead.

Summary

  • Environment Setup: Configure via make_microduck_velocity_env_cfg in microduck_velocity_env_cfg.py, which integrates the MICRODUCK_WALK_ROBOT_CFG robot model and custom MDP functions.
  • Observation Space: 61-dimensional vector combining 48 proprioception values with 13 command dimensions; passive joints are automatically filtered.
  • Training Execution: Use uv run train velocity with 64 environments for smoke tests or 4096 for production runs, leveraging PPO from rsl_rl with NaN-safe patches in mdp.py.
  • Domain Randomization: Enable CoM shifts, encoder bias, and friction scaling via environment configuration flags for robust sim-to-real transfer.
  • Deployment: Export trained policies to ONNX format using scripts/export.py, then run inference via scripts/infer_policy.py for hardware deployment.

Frequently Asked Questions

What observation space does the Microduck RL policy use?

The policy consumes a 61-dimensional observation vector comprising 48-dimensional proprioception data (joint positions, velocities, IMU readings) and a 13-dimensional command block. The mdp.py implementation specifically filters out passive joints (prefixed passive_*) from both the observation buffer and the final ONNX export to match the physical robot's active servo count.

How do I prevent numerical instabilities during training?

The repository implements multiple safeguards in src/mjlab_microduck/tasks/mdp.py: a NaN-safe RewardManager.compute wrapper, PPO advantage sanitization, and a custom robot_state_is_nan termination guard. These patches protect the training loop from MuJoCo numerical failures that can occur during aggressive domain randomization.

Can I train the policy without dedicated GPUs?

While the pipeline is optimized for GPU acceleration via mjlab (MuJoCo Warp), you can run inference on CPU-only hosts using uv run scripts/infer_policy.py --walking out.onnx. However, training 4,096 parallel environments effectively requires GPU acceleration to maintain reasonable iteration times.

How do I deploy the trained policy to the physical Microduck robot?

After training, export the policy to ONNX format using scripts/export.py, which embeds the observation normalization statistics directly into the model graph. Transfer the resulting .onnx file to the robot's onboard computer, where it can be executed using the Microduck control stack. The exported model expects the same 61-dimensional observation format used during training.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →