How Backlash Is Modeled in the Microduck Robot for Sim2real Transfer
Bold summary: Microduck models backlash through three coupled components: passive hinge joints injected into the MJCF robot description, a custom BacklashEncoderBamActuator that reads encoders through the play, and environment wrapper functions that expose encoder-relative observations to the policy without breaking existing training code.
The sim2real transfer gap for low-cost servo motors is dominated by mechanical backlash—the dead zone in gear trains where motor motion doesn't immediately transmit to the output link. In the pollen-robotics/microduck_rl repository, this phenomenon is modeled explicitly rather than learned implicitly, allowing reinforcement learning policies to experience realistic sensor readings before deployment to physical hardware. This article breaks down the implementation across joint injection, actuator design, and environment configuration.
Backlash Joint Injection with add_backlash.py
The foundation of backlash modeling is structural. The stand-alone script src/mjlab_microduck/robot/microduck/add_backlash.py modifies MuJoCo MJCF robot files to insert unactuated hinge joints representing mechanical play.
For every joint marked class="chosen_actuator", the parser (lines 4–13, 31–55) inserts:
- A
<joint class="backlash">namedpassive_<joint>_backlashimmediately after the original joint - A
<default class="backlash">block limiting hinge travel to ±½ × backlash_deg (default 2° total play, yielding ±1° limits)
The passive_ prefix is critical. Uniform regex patterns like ^(?!passive_).* throughout the codebase automatically exclude these joints from actuator selection, observation gathering, and reward computation—no downstream code changes required.
Generate a backlash-enabled robot description:
python3 src/mjlab_microduck/robot/microduck/add_backlash.py \
robot_groundcontact_backlash.xml --backlash-deg 2.0
Backlash-Aware Actuator: BacklashEncoderBamActuator
Physical servos read their encoders after the gear train, measuring motor position plus accumulated play. The class BacklashEncoderBamActuator (defined in src/mjlab_microduck/actuator/friction_dr_bam.py, lines 64–78) replicates this behavior.
Inheriting from FrictionDRBamActuator, it discovers corresponding backlash joint IDs during initialization. At each control step, it computes:
effective_position = cmd.pos + qpos[backlash_joint]
This encoder-through-backlash reading then feeds into the standard BAM (Battery-Actuator Model) voltage law. The policy's commanded position is transformed by real physics rather than ideal kinematics.
Configuration follows the standard actuator pattern:
from mjlab_microduck.actuator.friction_dr_bam import BacklashEncoderBamActuatorCfg
act_cfg = BacklashEncoderBamActuatorCfg()
act = act_cfg.build(
entity=robot_entity,
target_ids=servo_joint_ids,
target_names=servo_joint_names
)
Environment Configuration with make_backlash_variant
The training pipeline must expose encoder-relative observations without breaking dimensionality assumptions. The helper make_backlash_variant in src/mjlab_microduck/tasks/backlash.py (lines 42–73) performs three swaps:
- Robot configuration: Replaces base config with
MICRODUCK_BACKLASH_ROBOT_CFG(defined inmicroduck_constants.py) - Observation terms: Substitutes
joint_pos_rel_backlashandjoint_vel_rel_backlashfor standard joint observations - Reward selectors: Restricts soft-limit penalties to original servo joints, ignoring passive backlash joints
The resulting environment maintains identical observation dimensionality (61 dimensions) so pretrained policies or standard architectures require no modification.
from mjlab_microduck.tasks.microduck_velocity_env_cfg import make_microduck_velocity_env_cfg
from mjlab_microduck.tasks.backlash import make_backlash_variant
cfg = make_microduck_velocity_env_cfg()
cfg_backlash = make_backlash_variant(cfg) # Ready for training
Complete Training Pipeline
The three components integrate into a unified workflow:
| Step | Command / Code | Output |
|---|---|---|
| 1. Generate robot | python add_backlash.py robot.xml --backlash-deg 2.0 |
MJCF with passive_*_backlash joints |
| 2. Configure environment | make_backlash_variant(cfg) |
Backlash-aware training config |
| 3. Train policy | Standard uv run train ... |
Policy conditioned on realistic encoder noise |
Because passive joints are excluded by regex, existing observation terms, reward functions, and logging utilities function unchanged. This non-invasive design preserves code stability while closing the sim2real gap.
Key Implementation Files
src/mjlab_microduck/robot/microduck/add_backlash.py— MJCF joint injectionsrc/mjlab_microduck/actuator/friction_dr_bam.py—BacklashEncoderBamActuatorimplementationsrc/mjlab_microduck/tasks/backlash.py— Environment wrapper and observation wiringsrc/mjlab_microduck/robot/microduck_constants.py—MICRODUCK_BACKLASH_ROBOT_CFGdefinition
Summary
- Structural realism: Backlash joints are explicitly added to the MuJoCo model with tunable play limits
- Sensor fidelity:
BacklashEncoderBamActuatorreproduces physical servo encoder behavior by reading through the play joint - Training compatibility:
make_backlash_variantmaintains observation dimensionality while swapping to encoder-relative terms - Code isolation: The
passive_prefix convention prevents backlash joints from interfering with existing regex-based selection logic
Frequently Asked Questions
How much backlash does the default Microduck configuration model?
The default configuration applies 2° total play (±1° limits per joint), matching the gear train backlash observed in the physical Dynamixel servos used on the hardware platform. This value can be adjusted via the --backlash-deg argument to add_backlash.py.
Why insert passive joints instead of adding noise to observations?
Passive joints model backlash as stateful physics rather than instantaneous noise. The play accumulates based on motion history and external forces, matching the hysteretic behavior of real gear trains. Observation noise would lose this memory effect and produce inferior sim2real transfer.
Does adding backlash joints increase training time?
Yes, typically 20–40% more simulation steps are required for equivalent policy convergence because the effective action space becomes discontinuous. However, policies trained with backlash generalize to hardware without fine-tuning, eliminating costly real-world adaptation iterations.
Can I use BacklashEncoderBamActuator with other robot configurations?
Yes, provided the robot MJCF contains joints following the passive_<joint>_backlash naming convention. The actuator discovers backlash joint IDs automatically by matching parent body names, making it reusable across different robot morphologies that share the same injection workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →