How Microduck RL Models Backlash in Its Actuator Architecture: A Complete Technical Guide
Microduck RL simulates gearbox backlash by injecting unactuated passive hinges into the MJCF robot description and routing encoder feedback through these joints to match real hardware behavior.
The Microduck RL framework from Pollen Robotics implements a physics-based backlash model that captures the mechanical play between servo motors and output shafts. Rather than approximating backlash with mathematical functions, the simulation adds explicit degrees of freedom that replicate how encoders actually measure position on the physical robot.
Backlash Joint Architecture
The core mechanism relies on dual-joint construction: every actuated servo joint receives a companion passive hinge that represents gear play.
Passive Joint Injection
The script src/mjlab_microduck/robot/microduck/add_backlash.py transforms standard Microduck MJCF descriptions by inserting passive_<joint>_backlash hinges. Each backlash hinge:
- Rotates on the same axis as its parent servo joint
- Has a ±1° range constraining the maximum mechanical play
- Belongs to the
backlashclass for styling and property management
This construction mirrors physical gearboxes where the motor shaft and output shaft can diverge slightly before load transmission occurs.
Backlash-Enabled Robot Variants
Three MJCF specifications expose this configuration, defined in src/mjlab_microduck/robot/microduck_constants.py:
| Constant | MJCF File | Use Case |
|---|---|---|
MICRODUCK_ALLCOLLISIONS_BACKLASH_XML |
robot_allcollisions_backlash.xml |
Full collision simulation |
MICRODUCK_WALK_BACKLASH_XML |
robot_walk_backlash.xml |
Locomotion training |
MICRODUCK_ALLCOLLISIONS_ROLLERS_BACKLASH_XML |
robot_allcollisions_rollers_backlash.xml |
Roller-skate variant |
BacklashEncoderBamActuatorCfg: Hardware-Matching Actuator Behavior
The BacklashEncoderBamActuatorCfg (defined in microduck_constants.py) extends the standard BAM actuator to report measurements through the backlash hinge rather than reading the servo joint directly.
This design choice reflects a critical hardware detail: on the physical Microduck robot, encoders mount on the output side of the gear train. The simulated actuator therefore returns:
encoder_reading = servo_joint_position + backlash_joint_offset
The actuator otherwise maintains identical control dynamics to its non-backlash counterpart, ensuring fair comparisons across model variants.
Observation Functions: Computing Through-Backlash Readings
The observation pipeline in src/mjlab_microduck/tasks/mdp.py implements backlash-aware position and velocity measurements using joint index caching for performance.
Position Observation with Backlash
def joint_pos_rel_backlash(env, asset, asset_cfg):
main_ids, bl_ids, mask = _backlash_encoder_ids(env, asset, asset_cfg)
return (env.data.qpos[main_ids] + env.data.qpos[bl_ids]) - default_pos
Velocity Observation with Backlash
def joint_vel_rel_backlash(env, asset, asset_cfg):
main_ids, bl_ids, mask = _backlash_encoder_ids(env, asset, asset_cfg)
return env.data.qvel[main_ids] + env.data.qvel[bl_ids]
ID Caching with _backlash_encoder_ids
The helper function _backlash_encoder_ids (also in mdp.py) performs one-time discovery of:
main_ids: Indices of servo joints inenv.data.qposbl_ids: Indices of correspondingpassive_*_backlashjointsmask: Boolean array enabling the same code to run on models without backlash
This caching eliminates runtime string matching during training rollouts.
Task-Level Integration: The make_backlash_variant Function
Converting any Microduck environment to its backlash counterpart requires coordinated changes across robot model, observations, and reward structure. The make_backlash_variant function in src/mjlab_microduck/tasks/backlash.py handles this transformation:
- Robot swap: Replaces the base MJCF with its backlash-enabled version
- Observation replacement: Substitutes
joint_pos→joint_pos_rel_backlashandjoint_vel→joint_vel_rel_backlash - Reward scoping: Restricts
dof_pos_limitsrewards to servo joints only, allowing backlash hinges to rest at their ±1° limits without penalty
The resulting environment maintains an identical 61-dimensional observation space to the base version, enabling direct policy transfer comparisons.
Creating a Backlash-Enabled Environment
from mjlab_microduck.tasks import make_backlash_variant
from mjlab_microduck.tasks import make_microduck_velocity_env_cfg
# Base configuration without backlash
base_cfg = make_microduck_velocity_env_cfg(play=False)
# Convert to backlash variant
backlash_cfg = make_backlash_variant(base_cfg, asset_cfg=None)
The asset_cfg parameter accepts optional asset-specific overrides; passing None uses default mappings.
Physical Fidelity: Why This Approach Matters
The Microduck RL backlash model achieves high sim-to-real transfer through mechanism-level accuracy:
- Energy conserving: Passive hinges use MuJoCo's default damping, letting backlash settle naturally under load
- Observation-realistic: Encoder readings include the cumulative effect of gear play, matching how controllers actually perceive state
- Policy-agnostic: The 61-D observation space remains unchanged, so existing training infrastructures require no modification
This stands in contrast to approaches that model backlash as noise or hysteresis functions in observation space, which can introduce unrealistic phase relationships between command and measured position.
Summary
- Backlash joints are unactuated
passive_<joint>_backlashhinges with ±1° range, injected byadd_backlash.py - Three MJCF variants expose backlash:
robot_allcollisions_backlash.xml,robot_walk_backlash.xml, and the rollers variant BacklashEncoderBamActuatorCfgroutes encoder feedback through the backlash hinge to match physical encoder placement- Observation functions in
mdp.pysum servo and backlash joint states using cached index mappings make_backlash_variantconverts any base environment while preserving observation space dimensions
Frequently Asked Questions
How does the BacklashEncoderBamActuatorCfg differ from the standard BAM actuator?
The only difference is encoder feedback routing. According to microduck_constants.py, BacklashEncoderBamActuatorCfg inherits all BAM actuator parameters but reports joint position and velocity as the sum of servo joint and backlash joint states. This matches hardware where encoders measure output shaft position rather than motor shaft position.
Can policies trained without backlash transfer to backlash-enabled environments?
Direct transfer is possible but suboptimal. The observation spaces are dimensionally identical (61-D), but the dynamics differ because backlash introduces free play between command and initial response. Pollen Robotics recommends either training on the backlash variant or using domain randomization across both conditions for robust policies.
What determines the ±1° backlash range limit?
The ±1° range in add_backlash.py represents the physical gear play measured on the actual Microduck hardware. This value is hardcoded in the MJCF generation script and applies uniformly across all joints. The small range ensures backlash effects remain subtle—typical of high-quality harmonic or planetary drives—while still capturing the essential nonlinearity for sim-to-real transfer.
How does _backlash_encoder_ids handle robots without backlash joints?
It returns an empty mask for non-backlash models. The function detects missing passive_*_backlash joints by name pattern matching and sets mask to all False in that case. Observation functions then effectively ignore the backlash term (summing zero or filtering via the mask), allowing the same codebase to serve both model variants without branching logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →