What Is the Microduck RL Training Repository Based On?
The Microduck RL training repository is built on MuJoCo Warp for physics simulation, the rsl_rl library for Proximal Policy Optimization (PPO), and a custom Battery-Actuated-Motor (BAM) actuator model to close the sim-to-real gap for the 800 g bipedal Microduck robot.
The pollen-robotics/microduck_rl repository provides a complete simulation-to-real pipeline for training control policies of the 25 cm Microduck biped. Its architecture layers specialized robotics middleware on top of modern GPU-accelerated physics to deliver robust, deployable locomotion policies.
Core Architecture Layers
The repository is structured around three foundational components that handle physics, learning, and hardware modeling.
Simulation Engine: MuJoCo Warp via mjlab
At the lowest layer, the repository uses MuJoCo Warp—the next-generation GPU-accelerated version of MuJoCo—accessed through the mjlab framework. This combination supplies the low-level physics simulation, scene management, and RL-environment scaffolding. According to the project README, mjlab provides the underlying infrastructure that manages parallel environment execution and scene composition.
Policy Learning: PPO from rsl_rl
For policy optimization, the repository implements Proximal Policy Optimization (PPO) from the rsl_rl library. The integration includes safety patches to handle numerical instability during training. Specifically, in src/mjlab_microduck/tasks/mdp.py (lines 30–40), the RewardManager is patched with a _nan_safe_reward_compute wrapper that guards against NaN and Inf values in reward calculations, preventing training crashes from physics edge cases.
Actuator Modeling: BAM with Domain Randomization
To bridge the simulation-to-reality gap, the repository models the BAM (Battery-Actuated-Motor) actuator—a voltage-controlled XL330 servo with load-dependent friction. The FrictionDRBamActuator class in src/mjlab_microduck/actuator/friction_dr_bam.py implements this high-fidelity model, incorporating:
- Battery voltage and voltage sag randomization
- Command delay simulation
- Friction magnitude randomization
This domain randomization ensures policies remain robust when transferred to the physical robot's power electronics.
Additional Technical Pillars
Beyond the core triad, several architectural decisions enable seamless deployment.
Backlash Simulation and Gear Play
The repository simulates mechanical backlash through a configurable ±1° gear-play model. Implemented in src/mjlab_microduck/tasks/backlash.py, this wrapper inserts passive hinge joints to mimic gearbox looseness while preserving the original observation and action dimensions. This design allows the same policy to run in both ideal and backlash-affected environments without architectural changes.
Fixed Observation Contract
All tasks share a fixed 61-dimensional observation vector comprising 48 proprioception channels plus command slots. This standardization enables hot-swapping policies at runtime and ensures consistent interfaces for the ONNX export pipeline described in the repository conventions.
ONNX Export with Baked Normalization
The export pipeline in scripts/export.py converts trained policies to ONNX format while baking the observation normalizer directly into the computation graph. This guarantees that deployment runtimes receive correctly scaled inputs without requiring external preprocessing.
Training Workflow and CLI Tools
The repository provides thin orchestration wrappers around mjlab's CLI:
src/mjlab_microduck/train_cli.py– Entry point foruv run traincommands that forward to mjlab's training scriptsrc/mjlab_microduck/hf_jobs.py– Helper for submitting distributed training jobs to Hugging Face infrastructure
Typical workflows use these commands:
# Train a walking policy with 4096 parallel environments
uv run train Mjlab-Velocity-Flat-MicroDuck --env.scene.num-envs 4096
# Visualize a trained policy
uv run play Mjlab-Velocity-Flat-MicroDuck --wandb-run-path myteam/microduck/abc123
# Export to ONNX with integrated normalizer
uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck --wandb-run-path myteam/microduck/abc123
# Run CPU-only inference for debugging
uv run scripts/infer_policy.py --walking output.onnx
Summary
- MuJoCo Warp via mjlab provides the GPU-accelerated physics backbone for parallel environment simulation
- rsl_rl PPO implementation includes NaN/Inf safety guards in
src/mjlab_microduck/tasks/mdp.py - BAM actuator model in
src/mjlab_microduck/actuator/friction_dr_bam.pysimulates voltage-controlled servos with realistic friction and power electronics - Domain randomization covers battery characteristics, delays, and friction magnitudes to ensure sim-to-real transfer
- Backlash simulation via
src/mjlab_microduck/tasks/backlash.pyadds configurable mechanical gear play - 61-dimensional fixed observation space enables policy hot-swapping and standardized ONNX export
- CLI tools and Hugging Face integration streamline local and cloud training workflows
Frequently Asked Questions
What physics engine does the Microduck RL repository use?
The repository uses MuJoCo Warp, the next-generation GPU-accelerated version of MuJoCo, accessed through the mjlab framework. According to the pollen-robotics/microduck_rl source code, this combination handles scene management, parallel environment execution, and low-level physics stepping.
How does the repository handle the sim-to-real gap?
The codebase addresses the sim-to-real gap through three mechanisms: the BAM actuator model (FrictionDRBamActuator) that accurately simulates voltage-controlled servo dynamics with load-dependent friction; comprehensive domain randomization of battery voltage, sag, and command delays; and backlash simulation that models mechanical gear play up to ±1°.
What is the observation space size in Microduck RL?
All tasks implement a fixed 61-dimensional observation vector consisting of 48 proprioception channels plus command slots. This fixed-size contract allows runtime policy swapping and consistent ONNX export across different training tasks.
How do I export a trained policy for deployment?
Use the export script located at scripts/export.py, which converts the trained PPO policy to ONNX format while baking the observation normalizer into the graph. This ensures the deployed model receives correctly scaled inputs without external preprocessing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →