# How Backlash Is Simulated in the Microduck RL Training Environments

> Learn how backlash is simulated in Microduck RL training environments. Discover techniques for passive hinge joints, backlash-aware observations, and combined joint states for accurate encoder readings.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: internals
- Published: 2026-09-01

---

**Backlash is simulated in the Microduck RL environments by injecting passive hinge joints into the MJCF robot model, swapping environment configurations to use backlash-aware observations, and computing combined joint states that replicate physical encoder readings.**

The `pollen-robotics/microduck_rl` repository implements a three-stage pipeline to model mechanical gear play for reinforcement learning training. Understanding how backlash is simulated in the Microduck RL training environments ensures policies trained in simulation transfer correctly to the physical robot, accounting for the tiny angular lag between a servo's output gear and the attached link.

## Robot Model Augmentation with Backlash Joints

The first stage modifies the robot's physical description to include unactuated joints representing gear play.

### Injecting Passive Joints via add_backlash.py

The post-import script [`src/mjlab_microduck/robot/microduck/add_backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/robot/microduck/add_backlash.py) parses the exported MJCF robot description and inserts an extra **passive** hinge joint for every actuated servo joint. Each new joint follows the naming convention `passive_<joint>_backlash`, remains unactuated, and is limited to a symmetric range of **±1°** that represents the physical gear play.

The script also injects a `<default class="backlash">` block that configures these joints with stiff limit constraints, small damping values, and appropriate armature settings. This ensures the passive joints exhibit realistic physical behavior when the servo attempts to change direction or hold position against external forces.

## Environment Configuration for Backlash Variants

Once the model contains backlash joints, the environment configuration must be updated to utilize them.

### The make_backlash_variant() Helper

Located in [`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py), the `make_backlash_variant()` function takes any existing Microduck environment configuration and transforms it into a backlash-aware variant. This helper performs three critical operations:

- **Swaps the robot description** to use `MICRODUCK_BACKLASH_ROBOT_CFG` (or the velocity-specific variant) instead of the standard configuration.
- **Replaces observation terms** with backlash-aware versions: `joint_pos_rel_backlash` and `joint_vel_rel_backlash`.
- **Scopes soft-limit penalties** to servo joints only, ensuring the newly added backlash hinges—which sit at their ±1° limits—do not generate spurious penalties during training.

This configuration layer ensures that existing task logic remains unchanged while the simulation now accounts for mechanical play.

## Backlash-Aware Observations and Reward Logic

The final stage implements the mathematical modeling that combines servo and backlash states into realistic encoder readings.

### Computing Encoder-Through-Backlash States

The core simulation logic resides in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py). A cached helper function `_backlash_encoder_ids` discovers, for each servo joint, whether a matching `passive_*_backlash` joint exists and returns three tensors: the indices of the main (servo) joints, the indices of the backlash hinges, and a binary mask (1.0 where a backlash joint is present).

The observation functions `joint_pos_rel_backlash` and `joint_vel_rel_backlash` then compute **qpos** and **qvel** as the sum of the servo joint value and the backlash joint value (scaled by the mask). This reproduces the real encoder reading that sits **after** the gear play, meaning the policy observes the combined angle and velocity exactly as a physical robot would.

The implementation uses regex conventions such as `^(?!passive_).*` throughout the environment (rewards, domain randomization, and curricula) to automatically ignore the added joints where they should not contribute, while ensuring they affect only the encoder-based observations.

## Practical Implementation Example

To create a velocity-tracking task with backlash simulation:

```python
from mjlab_microduck.tasks.backlash import make_backlash_variant
from mjlab_microduck.tasks.microduck_velocity_env_cfg import make_microduck_velocity_env_cfg

# Create base configuration

base_cfg = make_microduck_velocity_env_cfg()

# Transform to backlash variant

backlash_cfg = make_backlash_variant(base_cfg)

# This swaps the robot model and observation functions automatically

# The configuration now uses:

# - Robot with injected passive joints

# - joint_pos_rel_backlash instead of standard joint positions

# - joint_vel_rel_backlash instead of standard joint velocities

```

During rollout, the observation term returns the combined state:

```python

# Access the backlash-aware observation

obs = env.observations["actor"]["joint_pos"].func(env, biased=False)

# obs contains (qpos[servo] + qpos[backlash]) - default servo position

```

## Summary

- **Model Augmentation**: The [`add_backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/add_backlash.py) script injects passive hinge joints with a ±1° range into the MJCF robot description to model physical gear play.
- **Configuration Layer**: The `make_backlash_variant()` helper in [`tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tasks/backlash.py) creates task variants that use backlash-aware robot descriptions and observation functions.
- **Observation Logic**: Functions `joint_pos_rel_backlash` and `joint_vel_rel_backlash` in [`tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tasks/mdp.py) compute combined joint states that replicate real encoder readings through the gear play.
- **Selective Filtering**: Regex patterns like `^(?!passive_).*` ensure backlash joints are excluded from reward penalties while remaining active in observations.

## Frequently Asked Questions

### What is the default angular range for backlash joints in Microduck RL?

The default angular range is **±1°** (symmetric), representing the typical gear play between the servo output shaft and the attached link. This limit is configured in the `<default class="backlash">` block within the augmented MJCF robot description.

### How does the simulation prevent backlash joints from affecting reward penalties?

The environment scopes soft-limit penalties to servo joints only by using regex filtering. The pattern `^(?!passive_).*` excludes any joint with the `passive_` prefix—including the `passive_*_backlash` joints—from penalty calculations, ensuring these passive joints do not generate spurious costs when they sit at their ±1° limits.

### Can I use backlash simulation with any Microduck task configuration?

Yes. The `make_backlash_variant()` function in [`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py) is designed to wrap any existing environment configuration. It automatically swaps the robot configuration for the backlash-augmented variant and replaces the standard observation terms with their backlash-aware counterparts.

### Where are the backlash-augmented robot XML files defined?

The augmented robot files are referenced by JSON descriptors located at `src/mjlab_microduck/robot/microduck/config_mjcf_*_backlash.json`. These configuration files point to the XML robot files that have been processed by [`add_backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/add_backlash.py) to contain the injected passive joints.