How to Train a Microduck RL Policy with Backlash Simulation: Complete Guide

Train a Microduck RL policy with backlash simulation by selecting a task ID containing -Backlash-, running uv run train <TASK_ID>, and optionally exporting to ONNX for hardware deployment.

The microduck_rl repository by Pollen Robotics provides backlash-enabled reinforcement learning variants that model the physical robot's ±1° gear-play hinge and encoder behavior. This guide covers the exact workflow, source code architecture, and implementation details needed to train policies that transfer directly to real hardware.

How Backlash Simulation Works in Microduck RL

Microduck's backlash variants replicate physical actuator behavior by introducing passive joints that simulate gear play. According to the source code in src/mjlab_microduck/tasks/backlash.py, the make_backlash_variant function performs four key transformations:

  1. Robot model swap – Replaces the standard robot with a backlash-equipped XML (MICRODUCK_WALK_BACKLASH_ROBOT_CFG, MICRODUCK_ROLLERS_BACKLASH_ROBOT_CFG, or ground-contact variant)
  2. Observation remapping – Points joint position and velocity to joint_pos_rel_backlash and joint_vel_rel_backlash functions that return qpos[servo] + qpos[backlash]
  3. Soft-limit exclusion – Restricts dof_pos_limits reward penalties to servo joints only
  4. Pose reward adjustment – Updates regex patterns to ignore backlash joints in tracking rewards

The critical architectural guarantee: the observation space remains 61-dimensional and the action space stays 14-dimensional. This means identical training code, hyperparameters, and deployment pipelines work for both standard and backlash policies.

Selecting the Correct Backlash Task ID

Task registration occurs in src/mjlab_microduck/tasks/__init__.py. The _BACKLASH_TASKS tuple enumerates all base tasks with their backlash counterparts. Available task IDs follow a predictable pattern:

Base Task Backlash Variant
Mjlab-Velocity-Flat-MicroDuck Mjlab-Velocity-Flat-Backlash-MicroDuck
Mjlab-Velocity-Rough-MicroDuck Mjlab-Velocity-Rough-Backlash-MicroDuck
Mjlab-VelStand-Flat-MicroDuck Mjlab-VelStand-Flat-Backlash-MicroDuck
Mjlab-StandUp-Flat-MicroDuck Mjlab-StandUp-Flat-Backlash-MicroDuck

List all available backlash tasks:

uv run list-envs | grep Backlash

Step-by-Step Training Workflow

1. Environment Setup

Install dependencies with exact version resolution:

uv sync

This resolves CUDA-specific PyTorch wheels and all MJLab dependencies.

Validate configuration before committing GPU hours:

uv run train Mjlab-Velocity-Flat-Backlash-MicroDuck \
    --env.scene.num-envs 64 \
    --agent.max_iterations 5

3. Full Training Run

Execute the primary training command via src/mjlab_microduck/train_cli.py:

uv run train Mjlab-Velocity-Flat-Backlash-MicroDuck \
    --env.scene.num-envs 4096 \
    --agent.max_iterations 50000

Key parameters:

  • --env.scene.num-envs – Parallel environments (tune to GPU memory)
  • --agent.max_iterations – Total PPO update count
  • --hf-jobs – Submit to Hugging Face Jobs instead of local execution

The training CLI is a thin wrapper that intercepts --hf-jobs before importing task modules, then forwards to MJLab's core train script.

4. Policy Export for Deployment

Convert the trained checkpoint to ONNX format:

uv run scripts/export.py Mjlab-Velocity-Flat-Backlash-MicroDuck \
    --wandb-run-path <entity>/<project>/<run_id>

The exported model requires no preprocessing changes—deploy with standard robotctl policy add commands.

Key Source File Reference

File Purpose
src/mjlab_microduck/tasks/backlash.py make_backlash_variant() – core transformation logic
src/mjlab_microduck/tasks/__init__.py Task registration, _BACKLASH_TASKS tuple, MicroduckOnPolicyRunner
src/mjlab_microduck/robot/microduck_constants.py Backlash robot configurations
src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py Example base environment definition
src/mjlab_microduck/train_cli.py Entry point for uv run train
AGENTS.md Command reference and best practices

Curriculum and Reward Configuration

Backlash tasks inherit all curriculum definitions from their base configurations. The curriculum sections in files like microduck_velocity_env_cfg.py apply unchanged. No additional tuning is required for:

  • Domain randomization events
  • Curriculum thresholds
  • Reward term weightings
  • Termination conditions

The make_backlash_variant function in backlash.py specifically preserves these sections while only modifying robot-specific elements.

Summary

  • Identify your task by inserting -Backlash- into any standard Microduck task ID
  • Train with uv run train <TASK_ID> using identical hyperparameters to non-backlash variants
  • Export via scripts/export.py for hardware deployment without model changes
  • Verify behavior matches physical robot encoders that read through gear play

The backlash simulation architecture maintains full API compatibility while adding physical fidelity—enabling zero-friction sim-to-real transfer.

Frequently Asked Questions

What hardware differences does backlash simulation capture?

The backlash model adds a passive hinge joint (passive_<joint>_backlash) with ±1° free play and remaps encoder observations to sum servo and backlash positions. This matches physical encoders that measure output shaft position rather than motor position, as implemented in microduck_mdp.joint_pos_rel_backlash.

Can I use the same hyperparameters for backlash and standard training?

Yes. The observation space (61-D) and action space (14-D) remain identical. The MicroduckOnPolicyRunner class in tasks/__init__.py handles serialization without configuration changes. All PPO hyperparameters, network architectures, and reward weights transfer directly.

How do I know if my task is using the correct backlash robot configuration?

Check src/mjlab_microduck/robot/microduck_constants.py for three variants: MICRODUCK_WALK_BACKLASH_ROBOT_CFG (legged locomotion), MICRODUCK_ROLLERS_BACKLASH_ROBOT_CFG (wheeled), and ground-contact versions. The make_backlash_variant function automatically selects based on the base task's robot family.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →