What Is the Microduck RL Training Repository Based On?

The Microduck RL training repository is built on MuJoCo Warp for physics simulation, the rsl_rl library for Proximal Policy Optimization (PPO), and a custom Battery-Actuated-Motor (BAM) actuator model to close the sim-to-real gap for the 800 g bipedal Microduck robot.

The pollen-robotics/microduck_rl repository provides a complete simulation-to-real pipeline for training control policies of the 25 cm Microduck biped. Its architecture layers specialized robotics middleware on top of modern GPU-accelerated physics to deliver robust, deployable locomotion policies.

Core Architecture Layers

The repository is structured around three foundational components that handle physics, learning, and hardware modeling.

Simulation Engine: MuJoCo Warp via mjlab

At the lowest layer, the repository uses MuJoCo Warp—the next-generation GPU-accelerated version of MuJoCo—accessed through the mjlab framework. This combination supplies the low-level physics simulation, scene management, and RL-environment scaffolding. According to the project README, mjlab provides the underlying infrastructure that manages parallel environment execution and scene composition.

Policy Learning: PPO from rsl_rl

For policy optimization, the repository implements Proximal Policy Optimization (PPO) from the rsl_rl library. The integration includes safety patches to handle numerical instability during training. Specifically, in src/mjlab_microduck/tasks/mdp.py (lines 30–40), the RewardManager is patched with a _nan_safe_reward_compute wrapper that guards against NaN and Inf values in reward calculations, preventing training crashes from physics edge cases.

Actuator Modeling: BAM with Domain Randomization

To bridge the simulation-to-reality gap, the repository models the BAM (Battery-Actuated-Motor) actuator—a voltage-controlled XL330 servo with load-dependent friction. The FrictionDRBamActuator class in src/mjlab_microduck/actuator/friction_dr_bam.py implements this high-fidelity model, incorporating:

  • Battery voltage and voltage sag randomization
  • Command delay simulation
  • Friction magnitude randomization

This domain randomization ensures policies remain robust when transferred to the physical robot's power electronics.

Additional Technical Pillars

Beyond the core triad, several architectural decisions enable seamless deployment.

Backlash Simulation and Gear Play

The repository simulates mechanical backlash through a configurable ±1° gear-play model. Implemented in src/mjlab_microduck/tasks/backlash.py, this wrapper inserts passive hinge joints to mimic gearbox looseness while preserving the original observation and action dimensions. This design allows the same policy to run in both ideal and backlash-affected environments without architectural changes.

Fixed Observation Contract

All tasks share a fixed 61-dimensional observation vector comprising 48 proprioception channels plus command slots. This standardization enables hot-swapping policies at runtime and ensures consistent interfaces for the ONNX export pipeline described in the repository conventions.

ONNX Export with Baked Normalization

The export pipeline in scripts/export.py converts trained policies to ONNX format while baking the observation normalizer directly into the computation graph. This guarantees that deployment runtimes receive correctly scaled inputs without requiring external preprocessing.

Training Workflow and CLI Tools

The repository provides thin orchestration wrappers around mjlab's CLI:

Typical workflows use these commands:


# Train a walking policy with 4096 parallel environments

uv run train Mjlab-Velocity-Flat-MicroDuck --env.scene.num-envs 4096

# Visualize a trained policy

uv run play Mjlab-Velocity-Flat-MicroDuck --wandb-run-path myteam/microduck/abc123

# Export to ONNX with integrated normalizer

uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck --wandb-run-path myteam/microduck/abc123

# Run CPU-only inference for debugging

uv run scripts/infer_policy.py --walking output.onnx

Summary

  • MuJoCo Warp via mjlab provides the GPU-accelerated physics backbone for parallel environment simulation
  • rsl_rl PPO implementation includes NaN/Inf safety guards in src/mjlab_microduck/tasks/mdp.py
  • BAM actuator model in src/mjlab_microduck/actuator/friction_dr_bam.py simulates voltage-controlled servos with realistic friction and power electronics
  • Domain randomization covers battery characteristics, delays, and friction magnitudes to ensure sim-to-real transfer
  • Backlash simulation via src/mjlab_microduck/tasks/backlash.py adds configurable mechanical gear play
  • 61-dimensional fixed observation space enables policy hot-swapping and standardized ONNX export
  • CLI tools and Hugging Face integration streamline local and cloud training workflows

Frequently Asked Questions

What physics engine does the Microduck RL repository use?

The repository uses MuJoCo Warp, the next-generation GPU-accelerated version of MuJoCo, accessed through the mjlab framework. According to the pollen-robotics/microduck_rl source code, this combination handles scene management, parallel environment execution, and low-level physics stepping.

How does the repository handle the sim-to-real gap?

The codebase addresses the sim-to-real gap through three mechanisms: the BAM actuator model (FrictionDRBamActuator) that accurately simulates voltage-controlled servo dynamics with load-dependent friction; comprehensive domain randomization of battery voltage, sag, and command delays; and backlash simulation that models mechanical gear play up to ±1°.

What is the observation space size in Microduck RL?

All tasks implement a fixed 61-dimensional observation vector consisting of 48 proprioception channels plus command slots. This fixed-size contract allows runtime policy swapping and consistent ONNX export across different training tasks.

How do I export a trained policy for deployment?

Use the export script located at scripts/export.py, which converts the trained PPO policy to ONNX format while baking the observation normalizer into the graph. This ensures the deployed model receives correctly scaled inputs without external preprocessing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →