# How Microduck RL Achieves Sim‑to‑Real Transfer: A Complete Technical Breakdown

> Discover how Microduck RL achieves reliable sim2real transfer. Learn about physics accurate modelling, domain randomisation, and continuous hardware validation.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-02

---

**Microduck RL achieves reliable sim‑to‑real transfer through a rigorously engineered pipeline combining physics‑accurate actuator modelling, comprehensive domain randomisation, fixed observation contracts, and continuous validation against hardware.**

The `pollen-robotics/microduck_rl` repository implements a closed‑loop system that trains reinforcement learning policies in MuJoCo simulation and deploys them directly to physical XL330‑based robots with minimal performance degradation. This guide examines the exact mechanisms—from friction modelling to ONNX export—that make this transfer possible.

---

## Core Principle: Model Everything That Varies

Traditional sim‑to‑real approaches often fail because simulation omits physical phenomena present on hardware. Microduck RL inverts this: every source of real‑world variance is either explicitly modelled or randomised during training. According to the source code, this strategy appears in the comment blocks of environment configs: `# ── Domain randomisation (matched to the velocity env for sim2real parity) ────`.

---

## BAM Actuator: Physics‑Accurate Servo Simulation

The foundation of Microduck RL's transfer capability is the **BAM (Bipedal Actuator Model)** actuator implementation in [`src/mjlab_microduck/actuator/friction_dr_bam.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/actuator/friction_dr_bam.py).

### Custom Friction Model

Unlike MuJoCo's generic `dof_frictionloss`, the BAM actuator computes its own friction combining three physical effects:

- **Coulomb friction** – constant resistance independent of velocity
- **Stribeck friction** – velocity‑dependent transition from static to kinetic friction
- **Load‑dependent friction** – variation with joint torque

The `FrictionDRBamActuator` class wraps this model and exposes a `friction_scale` parameter that varies per environment:

```python

# From friction_dr_bam.py - FrictionDRBamActuator class

# friction_scale is sampled per episode to mimic real stiction variation

```

### Backlash Encoding

Real servos exhibit dead‑zone behavior from gearbox backlash. The `BacklashEncoderBamActuator` (also in [`friction_dr_bam.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/friction_dr_bam.py)) reads encoder positions *through* passive `*_backlash` joints defined in the robot MJCF, reproducing the exact encoder behavior seen on hardware.

---

## Domain Randomisation: Training for Robustness

Every training episode samples fresh physical parameters. This **non‑accumulating domain randomisation** forces policies to tolerate the exact uncertainties encountered on the real robot.

### Randomised Parameters

| Parameter | Implementation |
|-----------|---------------|
| `friction_scale` | `randomize_bam_friction` event in [`tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tasks/mdp.py) |
| Encoder bias | Added to joint position observations |
| IMU misalignment | Rotation perturbations to orientation readings |
| Payload mass | Variation in base link mass properties |
| Centre of mass | Perturbation to link inertial properties |

The event registration occurs in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py), with DR configuration spread across environment cfg files like [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py).

---

## The 61‑Dimensional Observation Contract

Microduck RL enforces a **fixed observation layout** that never changes between simulation and deployment.

### Layout Specification

- **48 dimensions** – proprioceptive observations (joint positions, velocities, IMU data)
- **13 dimensions** – command block (velocity targets, gait parameters, etc.)

Every environment, including test‑bench variants, pads unused command slots and maintains identical ordering. This invariant is documented in [`AGENTS.md`](https://github.com/pollen-robotics/microduck_rl/blob/main/AGENTS.md) and implemented in [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py).

The fixed layout guarantees that ONNX exports match runtime expectations exactly—no runtime reshaping or conditional logic required.

---

## Normaliser Baking: Eliminating Runtime Discrepancy

During export in [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py), the observation normaliser learned during training is **baked directly into the ONNX graph**:

```

RunningMeanStd statistics → ONNX nodes → fused inference graph

```

This eliminates a common sim‑real failure mode: different scaling constants between training and deployment. The same scaling applied in simulation is now inseparable from the policy itself.

Export a trained checkpoint:

```bash
uv run scripts/export.py <TASK_ID> \
    --wandb-run-path <entity>/<project>/<run_id> \
    --output policy.onnx

```

---

## Test‑Bench Validation: Quantifying the Sim‑Real Gap

The [`scripts/testbench_sim2real.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/testbench_sim2real.py) script provides **continuous validation** that the pipeline maintains parity between simulation and hardware.

### Three Operating Modes

1. **Simulation testing** – deterministic MuJoCo with exact BAM M6 actuator:

```bash
uv run python scripts/testbench_sim2real.py \
    --mode sim --sim-backend bam \
    --onnx policy.onnx --out sim.npz

```

2. **Hardware testing** – real XL330 test rig:

```bash
uv run python scripts/testbench_sim2real.py \
    --mode real --onnx policy.onnx \
    --out real.npz --port /dev/ttyUSB0 --motor-id 1

```

3. **Comparison mode** – statistical analysis and plotting:

```bash
uv run python scripts/testbench_sim2real.py \
    --compare sim.npz real.npz --out-plot comparison.png

```

The script reports **MAE and RMS errors** for trajectory, velocity, and action statistics. According to the source analysis, typical joint error remains under a few degrees—satisfying the project's sim‑real transfer requirement.

---

## Shared Regularisers: Curriculum Consistency

Action‑rate penalties, CoM regularisation, and joint‑torque‑rate limits are **identical across all environments**. This ensures that sim2real regularisers trained in the velocity base env transfer to derived tasks like [`microduck_standup_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_standup_env_cfg.py) and [`microduck_spin_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_spin_env_cfg.py).

The comment markers in these cfg files explicitly link DR settings to the velocity environment, maintaining parity across the task family.

---

## Deployment Runtime: [`scripts/infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/infer_policy.py)

The inference script demonstrates loading and executing the baked ONNX policy at **50 Hz**—matching the robot's control frequency:

```bash
uv run python scripts/infer_policy.py --onnx policy.onnx

```

This mirrors the runtime in the `pollen-robotics/microduck` repository, which loads the identical ONNX file without additional conversion.

---

## The Closed-Loop Pipeline

| Stage | Component | Output |
|-------|-----------|--------|
| Training | PPO via `rsl_rl` on velocity env with BAM + DR | Policy checkpoint |
| Export | [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) with normaliser baking | `policy.onnx` |
| Verification | [`testbench_sim2real.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/testbench_sim2real.py) sim vs. real comparison | Error metrics & plots |
| Deployment | Microduck robot runtime | Physical execution |

---

## Summary

- **BAM actuator** in [`friction_dr_bam.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/friction_dr_bam.py) replaces generic MuJoCo friction with voltage‑controlled XL330 physics including Coulomb, Stribeck, and load‑dependent effects.

- **Domain randomisation** via `randomize_bam_friction` and related events in [`tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tasks/mdp.py) randomises friction scale, encoder bias, IMU alignment, and inertial properties per episode.

- **61‑D observation contract** enforced across all environments guarantees ONNX export compatibility with runtime expectations.

- **Normaliser baking** during export eliminates scaling mismatches between simulation and hardware.

- **Test‑bench validation** in [`testbench_sim2real.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/testbench_sim2real.py) provides quantitative proof of sim‑real parity through comparative rollouts and statistical error analysis.

- **Shared curricula** across task variants ensure regularisation consistency for transfer learning.

---

## Frequently Asked Questions

### What makes the BAM actuator more accurate than standard MuJoCo actuators?

Standard MuJoCo actuators use a simple `dof_frictionloss` parameter. The BAM actuator in [`friction_dr_bam.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/friction_dr_bam.py) implements voltage‑controlled dynamics with physically accurate Coulomb, Stribeck, and load‑dependent friction models. It also models encoder backlash through passive joints, capturing the dead‑zone behavior of real servo gearboxes.

### How does domain randomisation prevent overfitting to simulation?

The `randomize_bam_friction` event and related DR mechanisms sample fresh physical parameters every episode without accumulation. This forces the policy to maintain performance across a distribution of dynamics rather than optimising for a single simulation instance. The randomisation targets specific hardware variations: stiction, encoder bias, IMU noise, and payload changes.

### Why is the 61‑dimensional observation layout fixed?

Fixed dimensions ensure that the exported ONNX graph has a static input signature. When environments pad unused command slots identically, the runtime in `pollen-robotics/microduck` can load any policy without reconfiguration. [`AGENTS.md`](https://github.com/pollen-robotics/microduck_rl/blob/main/AGENTS.md) documents this as a repository invariant.

### What error metrics indicate successful sim‑to‑real transfer?

The [`testbench_sim2real.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/testbench_sim2real.py) script computes MAE and RMS errors for joint trajectories, velocities, and action sequences when comparing simulated and hardware rollouts. According to the source analysis, successful transfer typically shows joint errors under a few degrees. Continuous integration runs this comparison to detect pipeline regressions.