# How to Train a Microduck RL Policy with Backlash Simulation: Complete Guide

> Learn to train a Microduck RL policy with backlash simulation. Follow our guide to select tasks, run training, and export your policy for hardware deployment. Get started today!

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Train a Microduck RL policy with backlash simulation by selecting a task ID containing `-Backlash-`, running `uv run train <TASK_ID>`, and optionally exporting to ONNX for hardware deployment.**

The `microduck_rl` repository by Pollen Robotics provides backlash-enabled reinforcement learning variants that model the physical robot's ±1° gear-play hinge and encoder behavior. This guide covers the exact workflow, source code architecture, and implementation details needed to train policies that transfer directly to real hardware.

## How Backlash Simulation Works in Microduck RL

Microduck's backlash variants replicate physical actuator behavior by introducing passive joints that simulate gear play. According to the source code in [`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py), the `make_backlash_variant` function performs four key transformations:

1. **Robot model swap** – Replaces the standard robot with a backlash-equipped XML (`MICRODUCK_WALK_BACKLASH_ROBOT_CFG`, `MICRODUCK_ROLLERS_BACKLASH_ROBOT_CFG`, or ground-contact variant)
2. **Observation remapping** – Points joint position and velocity to `joint_pos_rel_backlash` and `joint_vel_rel_backlash` functions that return `qpos[servo] + qpos[backlash]`
3. **Soft-limit exclusion** – Restricts `dof_pos_limits` reward penalties to servo joints only
4. **Pose reward adjustment** – Updates regex patterns to ignore backlash joints in tracking rewards

The critical architectural guarantee: **the observation space remains 61-dimensional and the action space stays 14-dimensional**. This means identical training code, hyperparameters, and deployment pipelines work for both standard and backlash policies.

## Selecting the Correct Backlash Task ID

Task registration occurs in [`src/mjlab_microduck/tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/__init__.py). The `_BACKLASH_TASKS` tuple enumerates all base tasks with their backlash counterparts. Available task IDs follow a predictable pattern:

| Base Task | Backlash Variant |
|-----------|----------------|
| `Mjlab-Velocity-Flat-MicroDuck` | `Mjlab-Velocity-Flat-Backlash-MicroDuck` |
| `Mjlab-Velocity-Rough-MicroDuck` | `Mjlab-Velocity-Rough-Backlash-MicroDuck` |
| `Mjlab-VelStand-Flat-MicroDuck` | `Mjlab-VelStand-Flat-Backlash-MicroDuck` |
| `Mjlab-StandUp-Flat-MicroDuck` | `Mjlab-StandUp-Flat-Backlash-MicroDuck` |

List all available backlash tasks:

```bash
uv run list-envs | grep Backlash

```

## Step-by-Step Training Workflow

### 1. Environment Setup

Install dependencies with exact version resolution:

```bash
uv sync

```

This resolves CUDA-specific PyTorch wheels and all MJLab dependencies.

### 2. Smoke Test (Recommended)

Validate configuration before committing GPU hours:

```bash
uv run train Mjlab-Velocity-Flat-Backlash-MicroDuck \
    --env.scene.num-envs 64 \
    --agent.max_iterations 5

```

### 3. Full Training Run

Execute the primary training command via [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py):

```bash
uv run train Mjlab-Velocity-Flat-Backlash-MicroDuck \
    --env.scene.num-envs 4096 \
    --agent.max_iterations 50000

```

Key parameters:
- `--env.scene.num-envs` – Parallel environments (tune to GPU memory)
- `--agent.max_iterations` – Total PPO update count
- `--hf-jobs` – Submit to Hugging Face Jobs instead of local execution

The training CLI is a thin wrapper that intercepts `--hf-jobs` before importing task modules, then forwards to MJLab's core `train` script.

### 4. Policy Export for Deployment

Convert the trained checkpoint to ONNX format:

```bash
uv run scripts/export.py Mjlab-Velocity-Flat-Backlash-MicroDuck \
    --wandb-run-path <entity>/<project>/<run_id>

```

The exported model requires no preprocessing changes—deploy with standard `robotctl policy add` commands.

## Key Source File Reference

| File | Purpose |
|------|---------|
| [`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py) | `make_backlash_variant()` – core transformation logic |
| [`src/mjlab_microduck/tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/__init__.py) | Task registration, `_BACKLASH_TASKS` tuple, `MicroduckOnPolicyRunner` |
| [`src/mjlab_microduck/robot/microduck_constants.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/robot/microduck_constants.py) | Backlash robot configurations |
| [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) | Example base environment definition |
| [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py) | Entry point for `uv run train` |
| [`AGENTS.md`](https://github.com/pollen-robotics/microduck_rl/blob/main/AGENTS.md) | Command reference and best practices |

## Curriculum and Reward Configuration

Backlash tasks **inherit all curriculum definitions** from their base configurations. The `curriculum` sections in files like [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py) apply unchanged. No additional tuning is required for:

- Domain randomization events
- Curriculum thresholds
- Reward term weightings
- Termination conditions

The `make_backlash_variant` function in [`backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/backlash.py) specifically preserves these sections while only modifying robot-specific elements.

## Summary

- **Identify** your task by inserting `-Backlash-` into any standard Microduck task ID
- **Train** with `uv run train <TASK_ID>` using identical hyperparameters to non-backlash variants
- **Export** via [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) for hardware deployment without model changes
- **Verify** behavior matches physical robot encoders that read through gear play

The backlash simulation architecture maintains full API compatibility while adding physical fidelity—enabling zero-friction sim-to-real transfer.

## Frequently Asked Questions

### What hardware differences does backlash simulation capture?

The backlash model adds a passive hinge joint (`passive_<joint>_backlash`) with ±1° free play and remaps encoder observations to sum servo and backlash positions. This matches physical encoders that measure output shaft position rather than motor position, as implemented in `microduck_mdp.joint_pos_rel_backlash`.

### Can I use the same hyperparameters for backlash and standard training?

Yes. The observation space (61-D) and action space (14-D) remain identical. The `MicroduckOnPolicyRunner` class in [`tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tasks/__init__.py) handles serialization without configuration changes. All PPO hyperparameters, network architectures, and reward weights transfer directly.

### How do I know if my task is using the correct backlash robot configuration?

Check [`src/mjlab_microduck/robot/microduck_constants.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/robot/microduck_constants.py) for three variants: `MICRODUCK_WALK_BACKLASH_ROBOT_CFG` (legged locomotion), `MICRODUCK_ROLLERS_BACKLASH_ROBOT_CFG` (wheeled), and ground-contact versions. The `make_backlash_variant` function automatically selects based on the base task's robot family.