# What Is the Microduck RL Training Repository Based On?

> Explore the Microduck RL training repository. Discover its foundation in MuJoCo Warp, rsl_rl, and a custom BAM actuator model for sim-to-real transfer in legged robotics.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: getting-started
- Published: 2026-09-08

---

**The Microduck RL training repository is built on MuJoCo Warp for physics simulation, the rsl_rl library for Proximal Policy Optimization (PPO), and a custom Battery-Actuated-Motor (BAM) actuator model to close the sim-to-real gap for the 800 g bipedal Microduck robot.**

The `pollen-robotics/microduck_rl` repository provides a complete simulation-to-real pipeline for training control policies of the 25 cm Microduck biped. Its architecture layers specialized robotics middleware on top of modern GPU-accelerated physics to deliver robust, deployable locomotion policies.

## Core Architecture Layers

The repository is structured around three foundational components that handle physics, learning, and hardware modeling.

### Simulation Engine: MuJoCo Warp via mjlab

At the lowest layer, the repository uses **MuJoCo Warp**—the next-generation GPU-accelerated version of MuJoCo—accessed through the **mjlab** framework. This combination supplies the low-level physics simulation, scene management, and RL-environment scaffolding. According to the project README, mjlab provides the underlying infrastructure that manages parallel environment execution and scene composition.

### Policy Learning: PPO from rsl_rl

For policy optimization, the repository implements **Proximal Policy Optimization (PPO)** from the **rsl_rl** library. The integration includes safety patches to handle numerical instability during training. Specifically, in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) (lines 30–40), the `RewardManager` is patched with a `_nan_safe_reward_compute` wrapper that guards against **NaN** and **Inf** values in reward calculations, preventing training crashes from physics edge cases.

### Actuator Modeling: BAM with Domain Randomization

To bridge the simulation-to-reality gap, the repository models the **BAM** (Battery-Actuated-Motor) actuator—a voltage-controlled XL330 servo with load-dependent friction. The `FrictionDRBamActuator` class in [`src/mjlab_microduck/actuator/friction_dr_bam.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/actuator/friction_dr_bam.py) implements this high-fidelity model, incorporating:
- Battery voltage and voltage sag randomization
- Command delay simulation  
- Friction magnitude randomization

This domain randomization ensures policies remain robust when transferred to the physical robot's power electronics.

## Additional Technical Pillars

Beyond the core triad, several architectural decisions enable seamless deployment.

### Backlash Simulation and Gear Play

The repository simulates mechanical backlash through a configurable ±1° gear-play model. Implemented in [`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py), this wrapper inserts passive hinge joints to mimic gearbox looseness while preserving the original observation and action dimensions. This design allows the same policy to run in both ideal and backlash-affected environments without architectural changes.

### Fixed Observation Contract

All tasks share a fixed **61-dimensional** observation vector comprising 48 proprioception channels plus command slots. This standardization enables hot-swapping policies at runtime and ensures consistent interfaces for the ONNX export pipeline described in the repository conventions.

### ONNX Export with Baked Normalization

The export pipeline in [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) converts trained policies to ONNX format while baking the observation normalizer directly into the computation graph. This guarantees that deployment runtimes receive correctly scaled inputs without requiring external preprocessing.

## Training Workflow and CLI Tools

The repository provides thin orchestration wrappers around mjlab's CLI:

- **[`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py)** – Entry point for `uv run train` commands that forward to mjlab's training script
- **[`src/mjlab_microduck/hf_jobs.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/hf_jobs.py)** – Helper for submitting distributed training jobs to Hugging Face infrastructure

Typical workflows use these commands:

```bash

# Train a walking policy with 4096 parallel environments

uv run train Mjlab-Velocity-Flat-MicroDuck --env.scene.num-envs 4096

# Visualize a trained policy

uv run play Mjlab-Velocity-Flat-MicroDuck --wandb-run-path myteam/microduck/abc123

# Export to ONNX with integrated normalizer

uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck --wandb-run-path myteam/microduck/abc123

# Run CPU-only inference for debugging

uv run scripts/infer_policy.py --walking output.onnx

```

## Summary

- **MuJoCo Warp** via mjlab provides the GPU-accelerated physics backbone for parallel environment simulation
- **rsl_rl PPO** implementation includes NaN/Inf safety guards in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py)
- **BAM actuator model** in [`src/mjlab_microduck/actuator/friction_dr_bam.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/actuator/friction_dr_bam.py) simulates voltage-controlled servos with realistic friction and power electronics
- Domain randomization covers battery characteristics, delays, and friction magnitudes to ensure sim-to-real transfer
- **Backlash simulation** via [`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py) adds configurable mechanical gear play
- **61-dimensional** fixed observation space enables policy hot-swapping and standardized ONNX export
- CLI tools and Hugging Face integration streamline local and cloud training workflows

## Frequently Asked Questions

### What physics engine does the Microduck RL repository use?

The repository uses **MuJoCo Warp**, the next-generation GPU-accelerated version of MuJoCo, accessed through the **mjlab** framework. According to the pollen-robotics/microduck_rl source code, this combination handles scene management, parallel environment execution, and low-level physics stepping.

### How does the repository handle the sim-to-real gap?

The codebase addresses the sim-to-real gap through three mechanisms: the **BAM actuator model** (`FrictionDRBamActuator`) that accurately simulates voltage-controlled servo dynamics with load-dependent friction; comprehensive **domain randomization** of battery voltage, sag, and command delays; and **backlash simulation** that models mechanical gear play up to ±1°.

### What is the observation space size in Microduck RL?

All tasks implement a fixed **61-dimensional** observation vector consisting of 48 proprioception channels plus command slots. This fixed-size contract allows runtime policy swapping and consistent ONNX export across different training tasks.

### How do I export a trained policy for deployment?

Use the export script located at [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py), which converts the trained PPO policy to ONNX format while baking the observation normalizer into the graph. This ensures the deployed model receives correctly scaled inputs without external preprocessing.