# How Observation and Action Dimensions Change with Backlash Simulation in Microduck RL

> Discover how backlash simulation in Microduck RL preserves 14D observation and action spaces by filtering unactuated DOFs and computing encoder-relative state vectors.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-01

---

**Backlash simulation adds internal passive hinges to each servo joint but preserves the 14-dimensional observation and action spaces by filtering unactuated degrees of freedom and computing encoder-relative state vectors.**

The `pollen-robotics/microduck_rl` repository implements a 14-servo quadruped robot in MuJoCo for reinforcement learning. When enabling **backlash simulation**, the physics model inserts unactuated passive hinges to mimic mechanical play in the drivetrain, yet the RL policy continues to receive a 61-dimensional observation vector and output 14-dimensional actions. This article explains exactly how the codebase maintains these fixed dimensions despite the additional joints, referencing the specific source files and functions responsible for the transformation.

## The Backlash Joint Architecture

When a backlash robot model is activated, the simulation inserts **two extra passive hinges** in series with each of the 14 servo joints. These hinges represent mechanical backlash and are completely **unactuated**—they move only in response to external forces and servo motion.

Although this increases the total joint count in the MuJoCo model, these additions are strictly internal. The policy interface remains unchanged because the codebase explicitly separates actuated servos from passive mechanical elements.

## Preserving Observation Dimensions with Encoder-Relative Functions

The observation space maintains its shape through specialized observation functions that aggregate servo and backlash states into single encoder readings. In **[`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py)**, the `make_backlash_variant` function (lines 15‑18) performs two critical operations:

1. Swaps the robot XML for the backlash-enabled version.
2. Rewrites the observation terms **`joint_pos`** and **`joint_vel`** to use **`joint_pos_rel_backlash`** and **`joint_vel_rel_backlash`** instead.

These replacement functions, defined in **[`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py)**, return the *encoder view* of each joint by summing the servo position with the current backlash angle. Consequently, the observation tensor retains a shape of `(num_envs, 14)` for joint positions and velocities, keeping the total observation dimension at **61** (48 proprioceptive features plus 13 command features).

```python
from mjlab_microduck.tasks.microduck_velocity_env_cfg import make_microduck_velocity_env_cfg
from mjlab_microduck.tasks.backlash import make_backlash_variant

# Base configuration without backlash

base_cfg = make_microduck_velocity_env_cfg(play=False, rough=False)

# Convert to backlash variant—observations automatically use *_rel_backlash* functions

backlash_cfg = make_backlash_variant(base_cfg)
env = backlash_cfg.make()
env.reset()

# Verify observation shape remains 14-dimensional for joint states

obs = env.observation_manager.active_terms["actor"]
print("joint_pos shape:", obs["joint_pos"].func(env).shape)  # (num_envs, 14)

```

## Maintaining 14-Dimensional Action Control

The action space remains **14-dimensional** because the policy is restricted to commanding only the servo joints. The filtering logic resides in **[`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py)**, where the helper **`_servo_joint_ids`** uses the regular expression `^(?!passive_).*` to exclude all passive joints from the action interface.

The constant **`mdp._SERVO_JOINTS_ONLY`** enforces this constraint during action application, ensuring that the 28 additional passive hinges (two per servo) receive zero torque commands and do not appear in the action tensor.

```python
from mjlab_microduck.tasks.mdp import _servo_joint_ids

robot = env.scene["robot"]
servo_ids = _servo_joint_ids(env, robot)
print("Servo joint ids:", servo_ids)  # array([0, 1, 2, ..., 13]) — only 14 controllable joints

# Confirm action scale matches servo count only

act = env.action_manager.get_term("joint_pos")
print("Action shape:", act._scale.shape)  # (1, 14)

```

## ONNX Export and Sim-to-Real Compatibility

During sim-to-real export, the codebase ensures the ONNX model metadata matches the 14-dimensional action space. The function **`mdp._get_base_metadata_no_passive`** patches the exporter to exclude passive joints from the metadata dictionary. This guarantees that exported `action_scale` and `joint_names` arrays contain exactly 14 entries, preventing dimension mismatches when deploying to the physical robot.

```python
from mjlab_microduck.scripts.export import export_policy

# Export patches metadata to exclude passive joints automatically

export_policy(
    task_id="Mjlab-Velocity-Backlash-MicroDuck",
    wandb_run_path="pollen-robotics/microduck_rl/<run-id>",
    output_path="out_backlash.onnx",
)

# Resulting ONNX contains joint_names with 14 entries only

```

## Summary

- **Backlash simulation** adds two passive hinges per servo internally without changing the external API.
- **Observations** stay at 61 dimensions (including 14 joint positions/velocities) by using encoder-relative aggregation functions that sum servo and backlash angles.
- **Actions** remain 14-dimensional through regex-based filtering in `_servo_joint_ids` that excludes all `passive_` prefixed joints.
- **ONNX export** compatibility is preserved via metadata patching that removes passive joints from the action scaling parameters.

## Frequently Asked Questions

### Does backlash simulation increase the observation vector size?

No. The observation vector remains **61-dimensional** (48 proprioceptive states plus 13 command features). While the raw physics simulation contains additional joints, the `joint_pos_rel_backlash` and `joint_vel_rel_backlash` functions collapse these into 14 encoder-equivalent values, maintaining the original observation shape required by the policy network.

### How many passive joints are added per servo?

The backlash model inserts **two unactuated passive hinges** in series with each of the 14 servos. These hinges simulate mechanical play but are excluded from both observation and action spaces by the filtering logic in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py).

### Why does the action space stay 14-dimensional despite extra joints?

The policy controls only actuated servo joints. The helper function `_servo_joint_ids` uses the negative lookahead regex `^(?!passive_).*` to strip passive joints from the action interface, and the constant `_SERVO_JOINTS_ONLY` ensures the action manager applies torques exclusively to the 14 controllable degrees of freedom.

### Where is the observation term rewriting handled?

The transformation occurs in **[`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py)**, specifically within `make_backlash_variant` (lines 15‑18). This function swaps the robot XML and reconfigures the observation manager to use the backlash-aware position and velocity functions, ensuring seamless integration with the existing 14-dimensional observation pipeline.