# How to Train a Reinforcement Learning Policy for the Microduck Bipedal Robot: A Complete PPO Pipeline

> Learn how to train a reinforcement learning policy for the Microduck bipedal robot. Explore the complete PPO pipeline for training, exporting, and deploying policies with configurable environments and domain randomization.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: how-to-guide
- Published: 2026-09-08

---

**The microduck_rl repository provides a complete PPO-based training pipeline built on mjlab (MuJoCo Warp) that enables you to train, export, and deploy policies for the 800g, 14-servo Microduck bipedal robot using configurable velocity-tracking environments and domain randomization.**

The pollen-robotics/microduck_rl repository offers a production-ready framework to train a reinforcement learning policy for the Microduck bipedal robot. Built atop **mjlab** (MuJoCo Warp) and the **rsl_rl** PPO implementation, this pipeline handles everything from environment configuration to ONNX export for real hardware deployment.

## Training Architecture Overview

The system implements a modular architecture centered on the `ManagerBasedRlEnvCfg` class. In [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py), the function `make_microduck_velocity_env_cfg` wires together the robot model, terrain, observations, commands, and rewards into a cohesive training environment.

### Robot Model Configuration

The `MICRODUCK_WALK_ROBOT_CFG` in [`src/mjlab_microduck/robot/microduck_constants.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/robot/microduck_constants.py) defines the MJCF specification for the 14-servo robot. This configuration includes:

- **BAM XL330** actuator specifications
- Passive joint definitions (filtered from observations)
- Joint limits and sensor configurations

### MDP Functions and Rewards

Custom reward terms reside in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py). The implementation provides NaN-safe computation patches and specialized reward functions including:

- **`track_linear_velocity`** and **`track_angular_velocity`** for command following
- **`head_pose_tracking`** for gaze direction
- **`upright`** and **`upright_progress`** for stability
- **`foot_slip`** (weak penalty) and **`action_rate_l2`** for regularization

## Configuring Observations and Actions

The Microduck environment uses a **61-dimensional observation space** consisting of 48-dimensional proprioception data plus a 13-dimensional command block. The action space spans **14 dimensions**, representing joint-position commands for each servo.

### Domain Randomization

Robust sim-to-real transfer relies on extensive domain randomization configured via environment flags in [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py):

- **CoM randomization**: Center-of-mass shifts per environment (`ENABLE_COM_RANDOMIZATION`)
- **Joint friction scaling**: Dynamic friction coefficient variation
- **Motor gain scaling**: Actuator response randomization
- **Encoder bias**: Per-environment constant offsets added to observations
- **IMU orientation randomization**: Perturbation of inertial measurement unit readings

### Curriculum Learning

Training stability is enhanced through `CurriculumTermCfg` schedules defined within the environment configuration. These step-wise schedules gradually adjust:

- Action-rate penalty weights
- Command velocity ranges
- Randomization magnitudes

## Executing Training Runs

The CLI entry point `uv run train <TASK_ID>` bootstraps the environment, applies the curriculum, and initiates the PPO loop with trajectory collection, advantage computation, and policy updates.

### Smoke Testing Configurations

Before committing to long training runs, validate your configuration with a minimal 5-iteration test:

```bash
uv run train velocity --env.scene.num-envs 64 --agent.max_iterations 5

```

This command creates 64 parallel velocity-tracking environments and stops after five PPO iterations, catching configuration errors early.

### Production Training

For full-scale training, allocate sufficient parallel environments and enable logging:

```bash
uv run train velocity \
  --env.scene.num-envs 4096 \
  --wandb-run-path myuser/microduck-velocity \
  --hf-jobs

```

Typical production settings use 4,096 parallel environments, log metrics to Weights & Biases, and optionally upload checkpoints to Hugging Face via the `--hf-jobs` flag.

## Exporting and Deployment

Once training converges, export the policy for hardware deployment using the ONNX exporter.

### ONNX Export with Observation Normalization

The export script bakes the observation normalizer directly into the graph and filters passive joints:

```bash
uv run scripts/export.py velocity \
  --wandb-run-path myuser/microduck-velocity/run123 \
  --output out.onnx

```

### CPU Inference Testing

Verify the exported policy runs correctly without GPU acceleration:

```bash
uv run scripts/infer_policy.py --walking out.onnx

```

For pure simulation without BAM actuator models, append the `--no-bam` flag to use the XML PD controller instead.

## Summary

- **Environment Setup**: Configure via `make_microduck_velocity_env_cfg` in [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py), which integrates the `MICRODUCK_WALK_ROBOT_CFG` robot model and custom MDP functions.
- **Observation Space**: 61-dimensional vector combining 48 proprioception values with 13 command dimensions; passive joints are automatically filtered.
- **Training Execution**: Use `uv run train velocity` with 64 environments for smoke tests or 4096 for production runs, leveraging PPO from rsl_rl with NaN-safe patches in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py).
- **Domain Randomization**: Enable CoM shifts, encoder bias, and friction scaling via environment configuration flags for robust sim-to-real transfer.
- **Deployment**: Export trained policies to ONNX format using [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py), then run inference via [`scripts/infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/infer_policy.py) for hardware deployment.

## Frequently Asked Questions

### What observation space does the Microduck RL policy use?

The policy consumes a 61-dimensional observation vector comprising 48-dimensional proprioception data (joint positions, velocities, IMU readings) and a 13-dimensional command block. The [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) implementation specifically filters out passive joints (prefixed `passive_*`) from both the observation buffer and the final ONNX export to match the physical robot's active servo count.

### How do I prevent numerical instabilities during training?

The repository implements multiple safeguards in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py): a NaN-safe `RewardManager.compute` wrapper, PPO advantage sanitization, and a custom `robot_state_is_nan` termination guard. These patches protect the training loop from MuJoCo numerical failures that can occur during aggressive domain randomization.

### Can I train the policy without dedicated GPUs?

While the pipeline is optimized for GPU acceleration via mjlab (MuJoCo Warp), you can run inference on CPU-only hosts using `uv run scripts/infer_policy.py --walking out.onnx`. However, training 4,096 parallel environments effectively requires GPU acceleration to maintain reasonable iteration times.

### How do I deploy the trained policy to the physical Microduck robot?

After training, export the policy to ONNX format using [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py), which embeds the observation normalization statistics directly into the model graph. Transfer the resulting `.onnx` file to the robot's onboard computer, where it can be executed using the Microduck control stack. The exported model expects the same 61-dimensional observation format used during training.