# How to Set Up the Microduck RL Training Framework: Complete Installation and Usage Guide

> Master Microduck RL training with this comprehensive guide. Install and deploy ONNX policies using UV commands and a CUDA GPU. Start your reinforcement learning journey now.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: how-to-guide
- Published: 2026-09-02

---

**The Microduck RL training framework requires UV for dependency management, a CUDA-enabled GPU for MuJoCo Warp simulation, and four main commands—`uv sync`, `uv run train`, `uv run play`, and `uv run scripts/export.py`—to move from installation to deployment-ready ONNX policies.**

Setting up the Microduck RL training framework enables you to train reinforcement learning policies for robotic control using MuJoCo Warp and PPO before deploying them to real hardware. The `pollen-robotics/microduck_rl` repository provides a complete simulation-to-real pipeline with automatic observation normalization and ONNX export capabilities. This guide walks through the exact steps validated against the source code, including critical configuration details for ARM architectures and GPU memory tuning.

## Prerequisites and System Requirements

Before installing the Microduck RL framework, ensure your system meets the hardware and software dependencies defined in the repository's README.

**CUDA-enabled GPU**: Training requires NVIDIA CUDA because **MuJoCo Warp** ([`mujoco.warp`](https://github.com/google-deepmind/mujoco_warp)) executes physics simulations entirely on the GPU. CPU-only training is not supported for the RL loop.

**UV Package Manager**: The project uses Astral's UV instead of pip or conda. Install UV via the official installer:

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh

```

ARM64 users (Jetson, Apple Silicon with docker) must set an extended HTTP timeout before the first dependency sync to prevent wheel download failures:

```bash
export UV_HTTP_TIMEOUT=600

```

## Installation and Environment Setup

Clone the repository and synchronize dependencies using the lockfile provided in the repository root.

```bash
git clone https://github.com/pollen-robotics/microduck_rl.git
cd microduck_rl
uv sync

```

The `uv sync` command installs the exact CUDA-specific wheel versions pinned in the lockfile. On aarch64 platforms, this step is critical because UV pulls CUDA-enabled PyTorch wheels rather than the default CPU-only variants.

## Training Your First Policy

The framework provides a thin CLI shim at [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py) that forwards commands to the underlying mjlab training engine while intercepting platform-specific flags like `--hf-jobs`.

### Smoke Test Configuration

Verify your installation with a minimal 5-iteration run using 64 parallel environments. This catches CUDA memory errors and configuration mismatches quickly:

```bash
uv run train Mjlab-Velocity-Flat-MicroDuck \
  --env.scene.num-envs 64 \
  --agent.max_iterations 5

```

### Full Distributed Training

For complete training, increase the environment count to the default 4096 parallel simulations (adjust based on your GPU memory). The training entry point resides in [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py), which handles argument parsing and checkpoint initialization:

```bash
uv run train Mjlab-Velocity-Flat-MicroDuck \
  --env.scene.num-envs 4096

```

Checkpoint handling is built into mjlab’s train CLI. To resume an interrupted run, simply re-run the command with the same output directory; the trainer automatically detects and loads the latest checkpoint from the wandb run path.

## Visualization and Export

After convergence, visualize the policy behavior and export it for real-robot deployment.

### Interactive Playback

Use the `play` command to render the trained policy in the MuJoCo viewer:

```bash
uv run play Mjlab-Velocity-Flat-MicroDuck \
  --wandb-run-path <entity/project/run_id>

```

### ONNX Export with Baked Normalizer

Real-robot deployment requires the observation normalizer to be embedded directly in the model graph. The raw checkpoint cannot be used directly because the runtime expects a fixed 61-dimensional observation vector with normalization applied.

Run the export script located at [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py):

```bash
uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck \
  --wandb-run-path <entity/project/run_id>

```

This bakes the running mean and standard deviation into the ONNX graph, ensuring deterministic inference without external normalization logic.

## Deployment and CPU Inference

Run the exported policy without GPU acceleration using [`scripts/infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/infer_policy.py). This script loads the ONNX model and steps a CPU-only MuJoCo simulation, useful for testing control logic on laptops or edge devices without CUDA:

```bash
uv run scripts/infer_policy.py --walking output.onnx

```

## Advanced Configuration Options

### Backlash Variants

The framework supports training with actuator backlash simulation. Instantiate backlash variants by prepending `-Backlash-` to any task ID. The wrapper logic lives in [`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py):

```bash
uv run train Mjlab-Velocity-Backlash-Flat-MicroDuck \
  --env.scene.num-envs 4096

```

### Hugging Face Jobs

If you lack local GPU resources, submit training jobs to Hugging Face's compute cluster by adding the `--hf-jobs` flag. The interception logic is implemented in [`src/mjlab_microduck/train_hook.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_hook.py):

```bash
uv run train Mjlab-Velocity-Flat-MicroDuck --hf-jobs

```

### BAM Actuator Model

For domain-randomized friction modeling, the repository includes a BAM (Backlash and Friction) actuator implementation in [`src/mjlab_microduck/actuator/friction_dr_bam.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/actuator/friction_dr_bam.py). This module is automatically selected when using backlash-variant environments.

## Summary

- **Install UV** and set `UV_HTTP_TIMEOUT=600` on ARM machines before running `uv sync` to ensure CUDA wheels download correctly
- **Training requires a CUDA GPU**; use smoke tests with 64 environments to validate setup before launching full 4096-env jobs via [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py)
- **Export policies** using [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py) to bake the observation normalizer into the ONNX graph; raw checkpoints break the 61-dim observation contract
- **Deploy on CPU** using [`scripts/infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/infer_policy.py) for simulation testing without GPU hardware
- **Enable cloud training** with the `--hf-jobs` flag handled by [`src/mjlab_microduck/train_hook.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_hook.py) when local compute is unavailable

## Frequently Asked Questions

### What hardware is required to run Microduck RL training?

You need an NVIDIA GPU with CUDA support. MuJoCo Warp executes physics entirely on the GPU, making CUDA a hard requirement for the training loop. However, once you export the policy to ONNX using [`scripts/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/export.py), you can run inference on CPU-only machines using [`scripts/infer_policy.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/scripts/infer_policy.py).

### Why does dependency sync fail on my ARM device?

The PyTorch CUDA wheels for aarch64 are large and require extended download timeouts. Set `export UV_HTTP_TIMEOUT=600` before running `uv sync` to prevent HTTP timeouts during the first dependency resolution. This is documented in the README lines 26-28 and is critical for Jetson or ARM server deployments.

### How do I resume training from a checkpoint?

Re-run the original `uv run train` command with the same wandb run path. The mjlab training engine automatically detects existing checkpoints in the run directory and resumes from the latest saved state. Do not modify the `--agent.max_iterations` parameter unless you want to extend the total training budget beyond the original configuration.

### Can I train without a local GPU?

Yes. Add the `--hf-jobs` flag to any train command. The [`src/mjlab_microduck/train_hook.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_hook.py) module intercepts this flag and submits the job to Hugging Face's compute cluster. This is useful for training large 4096-environment batches when local GPU memory is insufficient.