How Microduck RL Achieves Sim‑to‑Real Transfer: A Complete Technical Breakdown

Microduck RL achieves reliable sim‑to‑real transfer through a rigorously engineered pipeline combining physics‑accurate actuator modelling, comprehensive domain randomisation, fixed observation contracts, and continuous validation against hardware.

The pollen-robotics/microduck_rl repository implements a closed‑loop system that trains reinforcement learning policies in MuJoCo simulation and deploys them directly to physical XL330‑based robots with minimal performance degradation. This guide examines the exact mechanisms—from friction modelling to ONNX export—that make this transfer possible.


Core Principle: Model Everything That Varies

Traditional sim‑to‑real approaches often fail because simulation omits physical phenomena present on hardware. Microduck RL inverts this: every source of real‑world variance is either explicitly modelled or randomised during training. According to the source code, this strategy appears in the comment blocks of environment configs: # ── Domain randomisation (matched to the velocity env for sim2real parity) ────.


BAM Actuator: Physics‑Accurate Servo Simulation

The foundation of Microduck RL's transfer capability is the BAM (Bipedal Actuator Model) actuator implementation in src/mjlab_microduck/actuator/friction_dr_bam.py.

Custom Friction Model

Unlike MuJoCo's generic dof_frictionloss, the BAM actuator computes its own friction combining three physical effects:

  • Coulomb friction – constant resistance independent of velocity
  • Stribeck friction – velocity‑dependent transition from static to kinetic friction
  • Load‑dependent friction – variation with joint torque

The FrictionDRBamActuator class wraps this model and exposes a friction_scale parameter that varies per environment:


# From friction_dr_bam.py - FrictionDRBamActuator class

# friction_scale is sampled per episode to mimic real stiction variation

Backlash Encoding

Real servos exhibit dead‑zone behavior from gearbox backlash. The BacklashEncoderBamActuator (also in friction_dr_bam.py) reads encoder positions through passive *_backlash joints defined in the robot MJCF, reproducing the exact encoder behavior seen on hardware.


Domain Randomisation: Training for Robustness

Every training episode samples fresh physical parameters. This non‑accumulating domain randomisation forces policies to tolerate the exact uncertainties encountered on the real robot.

Randomised Parameters

Parameter Implementation
friction_scale randomize_bam_friction event in tasks/mdp.py
Encoder bias Added to joint position observations
IMU misalignment Rotation perturbations to orientation readings
Payload mass Variation in base link mass properties
Centre of mass Perturbation to link inertial properties

The event registration occurs in src/mjlab_microduck/tasks/mdp.py, with DR configuration spread across environment cfg files like microduck_velocity_env_cfg.py.


The 61‑Dimensional Observation Contract

Microduck RL enforces a fixed observation layout that never changes between simulation and deployment.

Layout Specification

  • 48 dimensions – proprioceptive observations (joint positions, velocities, IMU data)
  • 13 dimensions – command block (velocity targets, gait parameters, etc.)

Every environment, including test‑bench variants, pads unused command slots and maintains identical ordering. This invariant is documented in AGENTS.md and implemented in microduck_velocity_env_cfg.py.

The fixed layout guarantees that ONNX exports match runtime expectations exactly—no runtime reshaping or conditional logic required.


Normaliser Baking: Eliminating Runtime Discrepancy

During export in scripts/export.py, the observation normaliser learned during training is baked directly into the ONNX graph:


RunningMeanStd statistics → ONNX nodes → fused inference graph

This eliminates a common sim‑real failure mode: different scaling constants between training and deployment. The same scaling applied in simulation is now inseparable from the policy itself.

Export a trained checkpoint:

uv run scripts/export.py <TASK_ID> \
    --wandb-run-path <entity>/<project>/<run_id> \
    --output policy.onnx

Test‑Bench Validation: Quantifying the Sim‑Real Gap

The scripts/testbench_sim2real.py script provides continuous validation that the pipeline maintains parity between simulation and hardware.

Three Operating Modes

  1. Simulation testing – deterministic MuJoCo with exact BAM M6 actuator:
uv run python scripts/testbench_sim2real.py \
    --mode sim --sim-backend bam \
    --onnx policy.onnx --out sim.npz
  1. Hardware testing – real XL330 test rig:
uv run python scripts/testbench_sim2real.py \
    --mode real --onnx policy.onnx \
    --out real.npz --port /dev/ttyUSB0 --motor-id 1
  1. Comparison mode – statistical analysis and plotting:
uv run python scripts/testbench_sim2real.py \
    --compare sim.npz real.npz --out-plot comparison.png

The script reports MAE and RMS errors for trajectory, velocity, and action statistics. According to the source analysis, typical joint error remains under a few degrees—satisfying the project's sim‑real transfer requirement.


Shared Regularisers: Curriculum Consistency

Action‑rate penalties, CoM regularisation, and joint‑torque‑rate limits are identical across all environments. This ensures that sim2real regularisers trained in the velocity base env transfer to derived tasks like microduck_standup_env_cfg.py and microduck_spin_env_cfg.py.

The comment markers in these cfg files explicitly link DR settings to the velocity environment, maintaining parity across the task family.


Deployment Runtime: scripts/infer_policy.py

The inference script demonstrates loading and executing the baked ONNX policy at 50 Hz—matching the robot's control frequency:

uv run python scripts/infer_policy.py --onnx policy.onnx

This mirrors the runtime in the pollen-robotics/microduck repository, which loads the identical ONNX file without additional conversion.


The Closed-Loop Pipeline

Stage Component Output
Training PPO via rsl_rl on velocity env with BAM + DR Policy checkpoint
Export scripts/export.py with normaliser baking policy.onnx
Verification testbench_sim2real.py sim vs. real comparison Error metrics & plots
Deployment Microduck robot runtime Physical execution

Summary

  • BAM actuator in friction_dr_bam.py replaces generic MuJoCo friction with voltage‑controlled XL330 physics including Coulomb, Stribeck, and load‑dependent effects.

  • Domain randomisation via randomize_bam_friction and related events in tasks/mdp.py randomises friction scale, encoder bias, IMU alignment, and inertial properties per episode.

  • 61‑D observation contract enforced across all environments guarantees ONNX export compatibility with runtime expectations.

  • Normaliser baking during export eliminates scaling mismatches between simulation and hardware.

  • Test‑bench validation in testbench_sim2real.py provides quantitative proof of sim‑real parity through comparative rollouts and statistical error analysis.

  • Shared curricula across task variants ensure regularisation consistency for transfer learning.


Frequently Asked Questions

What makes the BAM actuator more accurate than standard MuJoCo actuators?

Standard MuJoCo actuators use a simple dof_frictionloss parameter. The BAM actuator in friction_dr_bam.py implements voltage‑controlled dynamics with physically accurate Coulomb, Stribeck, and load‑dependent friction models. It also models encoder backlash through passive joints, capturing the dead‑zone behavior of real servo gearboxes.

How does domain randomisation prevent overfitting to simulation?

The randomize_bam_friction event and related DR mechanisms sample fresh physical parameters every episode without accumulation. This forces the policy to maintain performance across a distribution of dynamics rather than optimising for a single simulation instance. The randomisation targets specific hardware variations: stiction, encoder bias, IMU noise, and payload changes.

Why is the 61‑dimensional observation layout fixed?

Fixed dimensions ensure that the exported ONNX graph has a static input signature. When environments pad unused command slots identically, the runtime in pollen-robotics/microduck can load any policy without reconfiguration. AGENTS.md documents this as a repository invariant.

What error metrics indicate successful sim‑to‑real transfer?

The testbench_sim2real.py script computes MAE and RMS errors for joint trajectories, velocities, and action sequences when comparing simulated and hardware rollouts. According to the source analysis, successful transfer typically shows joint errors under a few degrees. Continuous integration runs this comparison to detect pipeline regressions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →