# What is Armature Randomization in Microduck RL? A Domain Randomization Deep-Dive

> Explore armature randomization in Microduck RL, a domain randomization technique that scales rotor inertia to train robust policies against motor variations.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-02

---

**Armature randomization is a domain-randomization technique in Microduck RL that varies the reflected rotor inertia of every joint at the start of each episode by scaling it ±10%, forcing policies to learn robust behaviors that tolerate real-world motor variations.**

The `microduck_rl` repository by Pollen Robotics implements this technique to bridge the sim-to-real gap for their Microduck quadruped robot. Each joint uses a Dynamixel XL330 servo—a "BAM" actuator—whose rotor inertia contributes to the joint's **armature** term in the dynamics. By randomizing this value, the training distribution covers the manufacturing tolerances and temperature-dependent variations encountered on physical hardware.

## How Armature Randomization Works in Microduck RL

### The Physics Behind Armature

In MuJoCo-based simulation, **armature** represents the reflected inertia of the motor rotor as seen at the joint output. For the Microduck robot's BAM actuators, this is non-negligible and affects how the joint responds to torque commands. The armature term appears in the equations of motion as an additional inertia that must be overcome during acceleration.

### The Randomization Mechanism

Microduck RL implements armature randomization through MuJoCo-lab's domain-randomization system. The key implementation resides in [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) at lines 500-512, where the `randomize_armature` event term is registered.

```python

# Configuration constants from microduck_velocity_env_cfg.py

ENABLE_ARMATURE_RANDOMIZATION = True
ARMATURE_RANDOMIZATION_RANGE = (0.9, 1.1)   # 10% variation in each direction

# Event registration executed on every environment reset

cfg.events["randomize_armature"] = EventTermCfg(
    func=dr.joint_armature,
    mode="reset",
    params={
        "asset_cfg": SceneEntityCfg("robot", joint_names=(r".*",)),
        "operation": "scale",
        "ranges": ARMATURE_RANDOMIZATION_RANGE,
    },
)

```

The `dr.joint_armature` helper function (defined in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py)) performs a **non-accumulating** scale operation: it reads the compile-time default armature value for each joint and multiplies it by a fresh random sample from `ARMATURE_RANDOMIZATION_RANGE` on every reset. Because MuJoCo-lab's randomization ops reset to defaults each time, the armature values are independently re-sampled per episode without compounding.

### Joint Selection and Propagation

The `SceneEntityCfg("robot", joint_names=(r".*",))` selector uses a regex pattern to target **all** joints in the robot's kinematic tree. The scaled armature value propagates directly to the BAM actuator's internal `dof_armature` property—the actuator does not zero this value, so the policy experiences the altered inertia throughout the episode.

## Why Armature Randomization Improves Policy Robustness

### Sim-to-Real Transfer

Real Dynamixel servos exhibit unit-to-unit variation in rotor inertia due to manufacturing tolerances. Additionally, temperature changes during operation affect magnetic properties and bearing friction, effectively altering the inertial response. Armature randomization ensures the policy encounters this variation during training, preventing catastrophic failure when deployed on hardware.

### Prevention of Dynamics Overfitting

Without domain randomization, policies can memorize the precise inertial properties of the simulation model—what researchers call **dynamics overfitting**. This produces brittle controllers that exploit unmodeled resonances or timing peculiarities of the specific simulation parameters. By presenting a distribution of armature values, Microduck RL forces the policy to learn control strategies that generalize across the plausible range of physical hardware.

### Episode-Level Variability vs. Step-Level Noise

The `mode="reset"` configuration is deliberate: armature changes occur **once per episode**, not every simulation step. This choice reflects the physical reality that motor inertia is a slow-varying parameter (changing with temperature over minutes), not a high-frequency noise source. Step-level randomization would create unrealistic dynamics and hinder learning stability.

## Implementation Across Microduck RL Environments

The armature randomization pattern appears consistently across task configurations in the repository. The velocity-tracking environment ([`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py)) serves as the primary reference implementation, but the same technique extends to specialized tasks.

| Environment File | Task Purpose | Armature Config Lines |
|------------------|------------|----------------------|
| [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py) | Forward velocity tracking | 500-512 |
| [`microduck_roller_crouch_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_roller_crouch_env_cfg.py) | Roller-aided crouching | Similar pattern |
| Other `microduck_*_env_cfg.py` files | Specialized behaviors | Reused configuration |

This consistency ensures that all Microduck RL policies benefit from the same robustness-injection technique regardless of the specific task objective.

## Comparison with Other Domain Randomization Techniques

Microduck RL employs armature randomization alongside other physics perturbations. Understanding how they interact clarifies its specific contribution:

- **Mass randomization**: Varies link masses (different `dr.*` function)—affects gravitational and inertial terms in the equations of motion
- **Friction randomization**: Modulates contact and joint friction coefficients—affects energy dissipation
- **Armature randomization**: Specifically targets **motor-level inertia**—affects the acceleration response to torque commands without changing steady-state behavior

The distinction matters for control: armature primarily impacts high-frequency dynamics and torque ripple response, whereas mass and friction dominate low-frequency, quasi-static regimes.

## Verifying Armature Randomization in Your Setup

To confirm armature randomization is active in a Microduck RL training run, inspect the logged environment configuration or add diagnostic printing:

```python

# Diagnostic: verify armature values per joint after reset

from mjlab_microduck.tasks.mdp import dr

# Inside environment step or callback

for joint_idx in range(env.num_joints):
    armature = env.sim.data.dof_armature[joint_idx]
    print(f"Joint {joint_idx}: armature = {armature:.6f}")

```

Values should vary episode-to-episode within `(default × 0.9, default × 1.1)` rather than remaining fixed at compile-time defaults.

## Summary

- **Armature randomization** in Microduck RL varies the reflected rotor inertia of Dynamixel XL330 joints by ±10% per episode using `dr.joint_armature` with `operation="scale"`
- The implementation in [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py) registers an event term with `mode="reset"` for non-accumulating, episode-level resampling
-This technique improves **sim-to-real transfer** by covering manufacturing tolerances and temperature-dependent motor variations
- Armature affects **high-frequency acceleration dynamics** distinct from mass or friction randomization
- The pattern is **reused across environment configurations** for consistent robustness training

## Frequently Asked Questions

### How does armature randomization differ from mass randomization in Microduck RL?

Mass randomization varies the inertial properties of robot links (their mass, center of mass, and inertia tensor) using `dr.body_mass` or `dr.body_inertia`, affecting how the robot responds to gravity and external forces. Armature randomization specifically targets the **motor rotor inertia** reflected at the joint output using `dr.joint_armature`. This distinction matters because armature primarily changes the joint's acceleration response to torque commands—a high-frequency, control-relevant effect—whereas mass randomization dominates the slower, gravity-compensated dynamics.

### Why is armature randomization set to ±10% rather than a larger range?

The ±10% range (`ARMATURE_RANDOMIZATION_RANGE = (0.9, 1.1)`) is calibrated to match observed variation in physical Dynamixel XL330 servos without including physically implausible values. Excessive armature variation can destabilize the simulation by creating dynamics outside the actuator's torque bandwidth or violating model assumptions in the MuJoCo integrator. Pollen Robotics likely validated this range against hardware characterizations to ensure the training distribution remains representative while still providing meaningful diversity.

### Can I disable armature randomization for ablation studies?

Yes—set `ENABLE_ARMATURE_RANDOMIZATION = False` in your environment configuration file before building the `EventTermCfg`. When disabled, the `randomize_armature` key should either be omitted from `cfg.events` or the scaling range set to `(1.0, 1.0)` for deterministic dynamics. This is useful for isolating the contribution of armature randomization to sim-to-real transfer performance or debugging control instability during training.

### Does armature randomization affect the learned policy's energy efficiency?

Indirectly, yes. Policies trained with armature randomization must be conservative in their torque commands to remain stable across the range of possible inertial responses. This typically results in **smoother, less aggressive control strategies** compared to policies trained on fixed armature, which can exploit precise model knowledge for marginally higher performance at the cost of robustness. The energy tradeoff is generally favorable: slightly higher nominal consumption yields dramatically better reliability on hardware.