Domain Randomization Techniques Applied to the Microduck Actuator Model
The Microduck RL framework implements five non-accumulating domain randomization techniques applied to the actuator model—including motor PD gain scaling, joint friction scaling, joint damping variation, armature randomization, and mass/inertia perturbation—to ensure robust sim-to-real transfer for the BAM-based robot.
The pollen-robotics/microduck_rl repository provides a high-fidelity simulation environment for the Microduck robot built on the BAM (Biologically-inspired Actuator Model) actuator. To minimize the reality gap between MuJoCo simulation and physical hardware, the framework applies specific domain randomization techniques applied to the actuator model during environment resets. These modifications target physical parameters such as friction budgets, motor gains, and inertial properties, with all randomizations designed as non-accumulating events that reset to nominal values before each new sample.
Motor PD Gain Randomization
The primary method for randomizing control dynamics targets the proportional and derivative gains of the servo motors. In src/mjlab_microduck/tasks/mdp.py, the function randomize_delayed_actuator_gains implements per-environment scaling of KP (proportional) and KD (derivative) gains for all BAM actuators.
During each environment reset, the system performs two critical steps to ensure non-accumulating behavior:
- Reset to nominal: The actuator's
reset_gainsmethod restores baseline KP/KD values - Apply new scale: Fresh random samples are drawn from user-provided
kp_rangeandkd_range, then applied viaset_gains
def randomize_delayed_actuator_gains(env, env_ids, kp_range, kd_range, asset_cfg=_DEFAULT_ASSET_CFG, operation="scale"):
asset = env.scene[asset_cfg.name]
for actuator in asset.actuators:
if not isinstance(actuator, BamActuator):
continue
n_joints = len(actuator.ctrl_ids)
kp_samples = torch.rand(len(env_ids), n_joints, device=env.device) * (kp_range[1] - kp_range[0]) + kp_range[0]
kd_samples = torch.rand(len(env_ids), n_joints, device=env.device) * (kd_range[1] - kd_range[0]) + kd_range[0]
actuator.reset_gains(env_ids) # restore nominal
actuator.set_gains(env_ids,
kp_scale=kp_samples.mean(dim=1, keepdim=True),
kd_scale=kd_samples.mean(dim=1, keepdim=True))
This approach prevents gain drift across episodes while exposing the policy to varied stiffness and damping characteristics.
Joint Friction Scaling
To simulate varying lubricant conditions and mechanical wear, the framework implements a specialized actuator class that randomizes the velocity-independent friction budget. The FrictionDRBamActuator class in src/mjlab_microduck/actuator/friction_dr_bam.py extends the base BAM actuator with a scalable friction multiplier.
The randomize_bam_friction function in src/mjlab_microduck/tasks/mdp.py manages this process:
- Storage: Each environment maintains a
friction_scaletensor (initialized to 1.0) - Randomization: Samples from
scale_rangeare applied viaset_friction_scale - Computation: The
_compute_friction_budgetmethod multiplies the base Coulomb, Stribeck, and load-dependent friction by the current scale factor
class FrictionDRBamActuator(BamActuator):
def initialize(self, mj_model, model, data, device):
super().initialize(mj_model, model, data, device)
self.friction_scale = torch.ones_like(self.kp_scale)
self.default_friction_scale = self.friction_scale.clone()
def _compute_friction_budget(self, motor_torque, external_torque, stribeck_coeff):
base = super()._compute_friction_budget(motor_torque, external_torque, stribeck_coeff)
fs = getattr(self, "friction_scale", None)
return base if fs is None else base * fs
This technique specifically targets the velocity-independent friction components, distinct from viscous damping effects.
Joint Damping and Armature Randomization
The framework includes two additional physical parameter randomizations that affect actuator dynamics through different mechanisms:
Damping Variation (No-Op for BAM)
The randomize_dof_field_scaled function can scale the MuJoCo dof_damping field, which represents viscous friction from lubricants and temperature variations. However, under the BAM actuator model, this field is explicitly zeroed in edit_spec, making this randomization a no-op for BAM-based environments while maintaining compatibility with other actuator types.
Armature Randomization
Unlike damping, the dof_armature field—representing the rotor's effective inertia—is preserved in BAM actuators. The dr.joint_armature event scales this value by factors drawn from ARMATURE_RANDOMIZATION_RANGE. This directly modifies the actuator's reflected inertia, affecting acceleration responses and high-frequency dynamics.
Mass and Inertia Perturbations
While not applied directly to the actuator object, the randomize_mass_and_inertia function in src/mjlab_microduck/tasks/mdp.py performs uniform scaling of robot body masses and inertial properties. These changes indirectly impact the actuator model by:
- Altering gravitational and inertial loads seen by the joint
- Changing the effective friction requirements under varying payload conditions
- Modifying the torque demands across the operating envelope
Configuring Domain Randomization Events
All randomization techniques are wired into task configurations through EventTermCfg instances. The configuration file src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py demonstrates how to enable specific randomizations using boolean flags.
if ENABLE_KP_RANDOMIZATION or ENABLE_KD_RANDOMIZATION:
cfg.events["randomize_motor_gains"] = EventTermCfg(
func=microduck_mdp.randomize_delayed_actuator_gains,
mode="reset",
params={
"asset_cfg": SceneEntityCfg("robot"),
"operation": "scale",
"kp_range": KP_RANDOMIZATION_RANGE,
"kd_range": KD_RANDOMIZATION_RANGE,
},
)
if ENABLE_JOINT_FRICTION_RANDOMIZATION:
cfg.events["randomize_joint_friction"] = EventTermCfg(
func=microduck_mdp.randomize_bam_friction,
mode="reset",
params={
"asset_cfg": SceneEntityCfg("robot"),
"scale_range": JOINT_FRICTION_RANDOMIZATION_RANGE,
},
)
The mode="reset" parameter ensures these events trigger during environment resets, maintaining the non-accumulating property essential for stable long-duration training runs.
Summary
- Motor PD Gain Scaling: Randomizes proportional and derivative gains via
randomize_delayed_actuator_gainsinsrc/mjlab_microduck/tasks/mdp.py, preventing accumulation by resetting to nominal values before sampling. - Joint Friction Scaling: Modifies velocity-independent friction budgets through
FrictionDRBamActuatorandrandomize_bam_friction, accounting for Coulomb and Stribeck effects. - Joint Damping: Implemented via
randomize_dof_field_scaledbut functionally disabled for BAM actuators as the field is zeroed in the actuator specification. - Armature Randomization: Scales rotor inertia through
dr.joint_armature, directly affecting BAM dynamics since the armature field is preserved. - Mass and Inertia: Indirectly affects actuators by varying mechanical loads through
randomize_mass_and_inertia.
These domain randomization techniques applied to the actuator model ensure that policies trained in simulation generalize effectively to the physical Microduck robot.
Frequently Asked Questions
How does friction randomization differ from standard MuJoCo damping?
Friction randomization in Microduck targets velocity-independent friction (Coulomb and Stribeck effects) through the FrictionDRBamActuator class, whereas standard MuJoCo damping represents velocity-dependent viscous friction. The BAM actuator explicitly manages its own friction budget internally, rendering the native dof_damping field inactive while the custom friction scaling remains fully operational.
Why are the randomization events considered "non-accumulating"?
Each randomization event first resets the parameter to its nominal value before applying a new random sample. For example, reset_gains restores baseline KP/KD values before set_gains applies new scales, and reset_friction_scale restores the factor to 1.0 before sampling. This design prevents error accumulation across thousands of environment resets, ensuring stable training distributions.
Can these techniques be applied to non-BAM actuators?
While the friction and gain randomization functions specifically check for BamActuator instances, the infrastructure supports other actuator types. The joint damping randomization (randomize_dof_field_scaled) and armature randomization operate directly on MuJoCo fields and would affect any actuator model using those fields. However, the specialized FrictionDRBamActuator features require inheritance from the BAM class.
Which parameters have the largest impact on sim-to-real transfer?
According to the implementation, motor PD gain scaling and joint friction scaling receive the most configuration attention, with dedicated enable flags (ENABLE_KP_RANDOMIZATION, ENABLE_JOINT_FRICTION_RANDOMIZATION) and default range parameters. The friction model particularly addresses hardware variability in lubrication and joint wear, while gain randomization covers controller calibration differences between simulation and physical hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →