How to Set Up the Microduck RL Training Framework: Complete Installation and Usage Guide

The Microduck RL training framework requires UV for dependency management, a CUDA-enabled GPU for MuJoCo Warp simulation, and four main commands—uv sync, uv run train, uv run play, and uv run scripts/export.py—to move from installation to deployment-ready ONNX policies.

Setting up the Microduck RL training framework enables you to train reinforcement learning policies for robotic control using MuJoCo Warp and PPO before deploying them to real hardware. The pollen-robotics/microduck_rl repository provides a complete simulation-to-real pipeline with automatic observation normalization and ONNX export capabilities. This guide walks through the exact steps validated against the source code, including critical configuration details for ARM architectures and GPU memory tuning.

Prerequisites and System Requirements

Before installing the Microduck RL framework, ensure your system meets the hardware and software dependencies defined in the repository's README.

CUDA-enabled GPU: Training requires NVIDIA CUDA because MuJoCo Warp (mujoco.warp) executes physics simulations entirely on the GPU. CPU-only training is not supported for the RL loop.

UV Package Manager: The project uses Astral's UV instead of pip or conda. Install UV via the official installer:

curl -LsSf https://astral.sh/uv/install.sh | sh

ARM64 users (Jetson, Apple Silicon with docker) must set an extended HTTP timeout before the first dependency sync to prevent wheel download failures:

export UV_HTTP_TIMEOUT=600

Installation and Environment Setup

Clone the repository and synchronize dependencies using the lockfile provided in the repository root.

git clone https://github.com/pollen-robotics/microduck_rl.git
cd microduck_rl
uv sync

The uv sync command installs the exact CUDA-specific wheel versions pinned in the lockfile. On aarch64 platforms, this step is critical because UV pulls CUDA-enabled PyTorch wheels rather than the default CPU-only variants.

Training Your First Policy

The framework provides a thin CLI shim at src/mjlab_microduck/train_cli.py that forwards commands to the underlying mjlab training engine while intercepting platform-specific flags like --hf-jobs.

Smoke Test Configuration

Verify your installation with a minimal 5-iteration run using 64 parallel environments. This catches CUDA memory errors and configuration mismatches quickly:

uv run train Mjlab-Velocity-Flat-MicroDuck \
  --env.scene.num-envs 64 \
  --agent.max_iterations 5

Full Distributed Training

For complete training, increase the environment count to the default 4096 parallel simulations (adjust based on your GPU memory). The training entry point resides in src/mjlab_microduck/train_cli.py, which handles argument parsing and checkpoint initialization:

uv run train Mjlab-Velocity-Flat-MicroDuck \
  --env.scene.num-envs 4096

Checkpoint handling is built into mjlab’s train CLI. To resume an interrupted run, simply re-run the command with the same output directory; the trainer automatically detects and loads the latest checkpoint from the wandb run path.

Visualization and Export

After convergence, visualize the policy behavior and export it for real-robot deployment.

Interactive Playback

Use the play command to render the trained policy in the MuJoCo viewer:

uv run play Mjlab-Velocity-Flat-MicroDuck \
  --wandb-run-path <entity/project/run_id>

ONNX Export with Baked Normalizer

Real-robot deployment requires the observation normalizer to be embedded directly in the model graph. The raw checkpoint cannot be used directly because the runtime expects a fixed 61-dimensional observation vector with normalization applied.

Run the export script located at scripts/export.py:

uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck \
  --wandb-run-path <entity/project/run_id>

This bakes the running mean and standard deviation into the ONNX graph, ensuring deterministic inference without external normalization logic.

Deployment and CPU Inference

Run the exported policy without GPU acceleration using scripts/infer_policy.py. This script loads the ONNX model and steps a CPU-only MuJoCo simulation, useful for testing control logic on laptops or edge devices without CUDA:

uv run scripts/infer_policy.py --walking output.onnx

Advanced Configuration Options

Backlash Variants

The framework supports training with actuator backlash simulation. Instantiate backlash variants by prepending -Backlash- to any task ID. The wrapper logic lives in src/mjlab_microduck/tasks/backlash.py:

uv run train Mjlab-Velocity-Backlash-Flat-MicroDuck \
  --env.scene.num-envs 4096

Hugging Face Jobs

If you lack local GPU resources, submit training jobs to Hugging Face's compute cluster by adding the --hf-jobs flag. The interception logic is implemented in src/mjlab_microduck/train_hook.py:

uv run train Mjlab-Velocity-Flat-MicroDuck --hf-jobs

BAM Actuator Model

For domain-randomized friction modeling, the repository includes a BAM (Backlash and Friction) actuator implementation in src/mjlab_microduck/actuator/friction_dr_bam.py. This module is automatically selected when using backlash-variant environments.

Summary

  • Install UV and set UV_HTTP_TIMEOUT=600 on ARM machines before running uv sync to ensure CUDA wheels download correctly
  • Training requires a CUDA GPU; use smoke tests with 64 environments to validate setup before launching full 4096-env jobs via src/mjlab_microduck/train_cli.py
  • Export policies using scripts/export.py to bake the observation normalizer into the ONNX graph; raw checkpoints break the 61-dim observation contract
  • Deploy on CPU using scripts/infer_policy.py for simulation testing without GPU hardware
  • Enable cloud training with the --hf-jobs flag handled by src/mjlab_microduck/train_hook.py when local compute is unavailable

Frequently Asked Questions

What hardware is required to run Microduck RL training?

You need an NVIDIA GPU with CUDA support. MuJoCo Warp executes physics entirely on the GPU, making CUDA a hard requirement for the training loop. However, once you export the policy to ONNX using scripts/export.py, you can run inference on CPU-only machines using scripts/infer_policy.py.

Why does dependency sync fail on my ARM device?

The PyTorch CUDA wheels for aarch64 are large and require extended download timeouts. Set export UV_HTTP_TIMEOUT=600 before running uv sync to prevent HTTP timeouts during the first dependency resolution. This is documented in the README lines 26-28 and is critical for Jetson or ARM server deployments.

How do I resume training from a checkpoint?

Re-run the original uv run train command with the same wandb run path. The mjlab training engine automatically detects existing checkpoints in the run directory and resumes from the latest saved state. Do not modify the --agent.max_iterations parameter unless you want to extend the total training budget beyond the original configuration.

Can I train without a local GPU?

Yes. Add the --hf-jobs flag to any train command. The src/mjlab_microduck/train_hook.py module intercepts this flag and submits the job to Hugging Face's compute cluster. This is useful for training large 4096-environment batches when local GPU memory is insufficient.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →