Microduck RL Domain Randomization Strategy: How Non‑Accumulating Perturbations Enable Sim‑to‑Real Transfer

Microduck RL uses a non‑accumulating, per‑episode domain randomization strategy that resets all physical parameters to nominal values before applying fresh random samples at the start of every training episode.

This approach prevents drift that would corrupt the simulation state over time, keeping the physics realistic while exposing policies to a wide distribution of hardware conditions. The strategy is implemented across actuator friction, mass/inertia, center‑of‑mass offsets, and motor gains in the pollen‑robotics/microduck_rl repository.

What Is Domain Randomization in Microduck RL?

Domain randomization is a sim‑to‑real technique that varies physical parameters during training to make policies robust to mismatches between simulation and reality. In Microduck RL, this strategy is deliberately non‑accumulating: every episode begins with a clean slate.

The design philosophy, documented in AGENTS.md, states:

Domain randomization must not accumulate across resets.

This invariant ensures that random perturbations do not compound across episodes, which would cause the simulated robot to drift unrealistically far from its hardware counterpart.

How the Non‑Accumulating Pattern Works

Each randomization function follows a three‑step template implemented in src/mjlab_microduck/tasks/mdp.py:

  1. Cache nominal values on first call
  2. Restore original parameters at episode reset
  3. Apply fresh random sample to the reset state

This pattern appears consistently across all randomization types. The mjlab 1.3.0 dr.* operations with operation="add" or operation="scale" enforce this behavior natively; custom randomizers implement it explicitly.

BAM Actuator Friction Randomization

The FrictionDRBamActuator class in src/mjlab_microduck/actuator/friction_dr_bam.py (lines 14–16) exposes a per‑environment friction_scale parameter. This scalar multiplies the actuator's velocity‑independent friction budget (Coulomb, Stribeck, and load components).


# From src/mjlab_microduck/actuator/friction_dr_bam.py

class FrictionDRBamActuator(BAMActuator):
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.friction_scale = None  # Per-env scale, set by DR event

The randomization event is registered in task configurations and executed via randomize_bam_friction in src/mjlab_microduck/tasks/mdp.py (lines 18–25):


# Example: Adding friction randomization to a task configuration

from mjlab_microduck.actuator.friction_dr_bam import FrictionDRBamActuator

def make_microduck_velocity_env_cfg():
    cfg = make_microduck_velocity_env_cfg_base()
    cfg.startup_events.append(
        dict(
            name="randomize_friction",
            fn="randomize_bam_friction",
            args=dict(scale_range=(0.8, 1.2)),  # 80%–120% of nominal

        )
    )
    return cfg

The set_friction_scale method resets friction to nominal before applying each new scale, guaranteeing no accumulation.

Mass and Inertia Randomization

Body mass and inertia are randomized together by a single scale factor per environment. The implementation in randomize_mass_and_inertia (src/mjlab_microduck/tasks/mdp.py, lines 44–92) caches original values and restores them before each new sample:


# Example: Randomizing mass and inertia for specific body IDs

from mjlab_microduck.tasks.mdp import randomize_mass_and_inertia

def apply_mass_randomization(env, env_ids):
    randomize_mass_and_inertia(
        env,
        env_ids,
        scale_range=(0.9, 1.1),  # ±10% variation

        asset_cfg=SceneEntityCfg(name="microduck", body_ids=slice(0, 10)),
    )

Key implementation detail: The function stores original_mass and original_inertia on first invocation, then uses these cached values as the baseline for all subsequent randomizations—not the previously randomized values.

Center‑of‑Mass Offset Randomization

CoM offsets follow the identical non‑accumulating pattern in randomize_com_offset (same file). Original CoM positions are cached and restored before each episode's uniform random sample, preventing positional drift that would misalign the robot's dynamics.

Voltage and Motor Gain Randomization

The BAM actuator natively supports per‑environment gain scaling through kp_scale and kd_scale parameters. These are initialized in src/mjlab_microduck/actuator/friction_dr_bam.py and sampled fresh each episode through the generic randomize_* event mechanism, adhering to the same reset‑then‑apply discipline.

Sensor Bias and Observation Noise

Observation noise and sensor biases are injected through the observation pipeline. These are re‑sampled per episode and never accumulated, ensuring that the policy learns to filter transient sensing errors rather than adapt to persistent, unrealistic bias.

Episode Reset Execution Flow

The event manager automatically triggers all registered randomizations when env.reset() is called:


# Inside the training loop

env.reset()  # Triggers in sequence:

             #   1. randomize_bam_friction (fresh friction_scale sample)

             #   2. randomize_mass_and_inertia (fresh mass/inertia scale)

             #   3. randomize_com_offset (fresh CoM offset)

             #   4. Additional registered randomizers...

Each function independently follows the cache‑restore‑sample protocol, ensuring orthogonal randomization of all physical parameters.

File Structure and Key Components

File Purpose
src/mjlab_microduck/actuator/friction_dr_bam.py FrictionDRBamActuator with per‑env friction_scale hook (lines 14–16); BacklashEncoderBamActuator subclass
src/mjlab_microduck/tasks/mdp.py Core randomization functions: randomize_bam_friction, randomize_mass_and_inertia, randomize_com_offset
AGENTS.md Design invariants including the non‑accumulation rule
tests/test_slope_curriculum.py Regression tests verifying randomization correctness and NaN‑free execution

Summary

  • Non‑accumulating design is the defining characteristic of Microduck RL's domain randomization strategy—every episode starts from nominal parameters.
  • Per‑episode sampling of friction scales, mass/inertia, CoM offsets, and motor gains creates a broad distribution of physical conditions without drift.
  • Explicit reset logic in all custom randomizers (randomize_mass_and_inertia, randomize_com_offset) complements native mjlab 1.3.0 dr.* operations.
  • Sim‑to‑real reliability emerges from training on realistic, non‑degraded physics variations that approximate actual hardware tolerances.

Frequently Asked Questions

How does Microduck RL prevent domain randomization from accumulating across episodes?

Every randomization function caches nominal values on first use, restores them at each episode reset, then applies a fresh random sample. This reset‑then‑apply pattern is enforced in src/mjlab_microduck/tasks/mdp.py for mass, inertia, and CoM randomizations, and natively by mjlab 1.3.0's dr.* operations for built‑in parameters.

What physical parameters does Microduck RL randomize?

The system randomizes BAM actuator friction (via friction_scale), body mass and inertia (joint scale factor), center‑of‑mass offsets, motor gains (kp_scale, kd_scale), and sensor noise/bias. All follow the non‑accumulating per‑episode pattern.

Why is non‑accumulating randomization important for sim‑to‑real transfer?

Accumulating randomizations would cause simulated parameters to drift unboundedly from hardware specifications, producing policies optimized for unrealistic physics. By resetting to nominal values each episode, Microduck RL keeps the simulation grounded in physically plausible configurations while still exposing the policy to varied conditions.

Where are the domain randomization functions defined in the codebase?

The core implementations reside in src/mjlab_microduck/tasks/mdp.py (functions randomize_bam_friction, randomize_mass_and_inertia, randomize_com_offset). The actuator‑level support is in src/mjlab_microduck/actuator/friction_dr_bam.py. Design principles are documented in AGENTS.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →