What is Armature Randomization in Microduck RL? A Domain Randomization Deep-Dive
Armature randomization is a domain-randomization technique in Microduck RL that varies the reflected rotor inertia of every joint at the start of each episode by scaling it ±10%, forcing policies to learn robust behaviors that tolerate real-world motor variations.
The microduck_rl repository by Pollen Robotics implements this technique to bridge the sim-to-real gap for their Microduck quadruped robot. Each joint uses a Dynamixel XL330 servo—a "BAM" actuator—whose rotor inertia contributes to the joint's armature term in the dynamics. By randomizing this value, the training distribution covers the manufacturing tolerances and temperature-dependent variations encountered on physical hardware.
How Armature Randomization Works in Microduck RL
The Physics Behind Armature
In MuJoCo-based simulation, armature represents the reflected inertia of the motor rotor as seen at the joint output. For the Microduck robot's BAM actuators, this is non-negligible and affects how the joint responds to torque commands. The armature term appears in the equations of motion as an additional inertia that must be overcome during acceleration.
The Randomization Mechanism
Microduck RL implements armature randomization through MuJoCo-lab's domain-randomization system. The key implementation resides in src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py at lines 500-512, where the randomize_armature event term is registered.
# Configuration constants from microduck_velocity_env_cfg.py
ENABLE_ARMATURE_RANDOMIZATION = True
ARMATURE_RANDOMIZATION_RANGE = (0.9, 1.1) # 10% variation in each direction
# Event registration executed on every environment reset
cfg.events["randomize_armature"] = EventTermCfg(
func=dr.joint_armature,
mode="reset",
params={
"asset_cfg": SceneEntityCfg("robot", joint_names=(r".*",)),
"operation": "scale",
"ranges": ARMATURE_RANDOMIZATION_RANGE,
},
)
The dr.joint_armature helper function (defined in src/mjlab_microduck/tasks/mdp.py) performs a non-accumulating scale operation: it reads the compile-time default armature value for each joint and multiplies it by a fresh random sample from ARMATURE_RANDOMIZATION_RANGE on every reset. Because MuJoCo-lab's randomization ops reset to defaults each time, the armature values are independently re-sampled per episode without compounding.
Joint Selection and Propagation
The SceneEntityCfg("robot", joint_names=(r".*",)) selector uses a regex pattern to target all joints in the robot's kinematic tree. The scaled armature value propagates directly to the BAM actuator's internal dof_armature property—the actuator does not zero this value, so the policy experiences the altered inertia throughout the episode.
Why Armature Randomization Improves Policy Robustness
Sim-to-Real Transfer
Real Dynamixel servos exhibit unit-to-unit variation in rotor inertia due to manufacturing tolerances. Additionally, temperature changes during operation affect magnetic properties and bearing friction, effectively altering the inertial response. Armature randomization ensures the policy encounters this variation during training, preventing catastrophic failure when deployed on hardware.
Prevention of Dynamics Overfitting
Without domain randomization, policies can memorize the precise inertial properties of the simulation model—what researchers call dynamics overfitting. This produces brittle controllers that exploit unmodeled resonances or timing peculiarities of the specific simulation parameters. By presenting a distribution of armature values, Microduck RL forces the policy to learn control strategies that generalize across the plausible range of physical hardware.
Episode-Level Variability vs. Step-Level Noise
The mode="reset" configuration is deliberate: armature changes occur once per episode, not every simulation step. This choice reflects the physical reality that motor inertia is a slow-varying parameter (changing with temperature over minutes), not a high-frequency noise source. Step-level randomization would create unrealistic dynamics and hinder learning stability.
Implementation Across Microduck RL Environments
The armature randomization pattern appears consistently across task configurations in the repository. The velocity-tracking environment (microduck_velocity_env_cfg.py) serves as the primary reference implementation, but the same technique extends to specialized tasks.
| Environment File | Task Purpose | Armature Config Lines |
|---|---|---|
microduck_velocity_env_cfg.py |
Forward velocity tracking | 500-512 |
microduck_roller_crouch_env_cfg.py |
Roller-aided crouching | Similar pattern |
Other microduck_*_env_cfg.py files |
Specialized behaviors | Reused configuration |
This consistency ensures that all Microduck RL policies benefit from the same robustness-injection technique regardless of the specific task objective.
Comparison with Other Domain Randomization Techniques
Microduck RL employs armature randomization alongside other physics perturbations. Understanding how they interact clarifies its specific contribution:
- Mass randomization: Varies link masses (different
dr.*function)—affects gravitational and inertial terms in the equations of motion - Friction randomization: Modulates contact and joint friction coefficients—affects energy dissipation
- Armature randomization: Specifically targets motor-level inertia—affects the acceleration response to torque commands without changing steady-state behavior
The distinction matters for control: armature primarily impacts high-frequency dynamics and torque ripple response, whereas mass and friction dominate low-frequency, quasi-static regimes.
Verifying Armature Randomization in Your Setup
To confirm armature randomization is active in a Microduck RL training run, inspect the logged environment configuration or add diagnostic printing:
# Diagnostic: verify armature values per joint after reset
from mjlab_microduck.tasks.mdp import dr
# Inside environment step or callback
for joint_idx in range(env.num_joints):
armature = env.sim.data.dof_armature[joint_idx]
print(f"Joint {joint_idx}: armature = {armature:.6f}")
Values should vary episode-to-episode within (default × 0.9, default × 1.1) rather than remaining fixed at compile-time defaults.
Summary
- Armature randomization in Microduck RL varies the reflected rotor inertia of Dynamixel XL330 joints by ±10% per episode using
dr.joint_armaturewithoperation="scale" - The implementation in
microduck_velocity_env_cfg.pyregisters an event term withmode="reset"for non-accumulating, episode-level resampling -This technique improves sim-to-real transfer by covering manufacturing tolerances and temperature-dependent motor variations - Armature affects high-frequency acceleration dynamics distinct from mass or friction randomization
- The pattern is reused across environment configurations for consistent robustness training
Frequently Asked Questions
How does armature randomization differ from mass randomization in Microduck RL?
Mass randomization varies the inertial properties of robot links (their mass, center of mass, and inertia tensor) using dr.body_mass or dr.body_inertia, affecting how the robot responds to gravity and external forces. Armature randomization specifically targets the motor rotor inertia reflected at the joint output using dr.joint_armature. This distinction matters because armature primarily changes the joint's acceleration response to torque commands—a high-frequency, control-relevant effect—whereas mass randomization dominates the slower, gravity-compensated dynamics.
Why is armature randomization set to ±10% rather than a larger range?
The ±10% range (ARMATURE_RANDOMIZATION_RANGE = (0.9, 1.1)) is calibrated to match observed variation in physical Dynamixel XL330 servos without including physically implausible values. Excessive armature variation can destabilize the simulation by creating dynamics outside the actuator's torque bandwidth or violating model assumptions in the MuJoCo integrator. Pollen Robotics likely validated this range against hardware characterizations to ensure the training distribution remains representative while still providing meaningful diversity.
Can I disable armature randomization for ablation studies?
Yes—set ENABLE_ARMATURE_RANDOMIZATION = False in your environment configuration file before building the EventTermCfg. When disabled, the randomize_armature key should either be omitted from cfg.events or the scaling range set to (1.0, 1.0) for deterministic dynamics. This is useful for isolating the contribution of armature randomization to sim-to-real transfer performance or debugging control instability during training.
Does armature randomization affect the learned policy's energy efficiency?
Indirectly, yes. Policies trained with armature randomization must be conservative in their torque commands to remain stable across the range of possible inertial responses. This typically results in smoother, less aggressive control strategies compared to policies trained on fixed armature, which can exploit precise model knowledge for marginally higher performance at the cost of robustness. The energy tradeoff is generally favorable: slightly higher nominal consumption yields dramatically better reliability on hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →