Core Components of the Microduck RL Architecture: A Complete Technical Guide

The Microduck RL architecture comprises six tightly-coupled modules: Environment Configurations that define task-specific scenes, MDP Functions that implement reward logic and NaN safety patches, custom BAM Actuator models with friction domain randomization, a Training CLI that wraps mjlab's PPO trainer, an ONNX Export pipeline that bakes observation normalizers, and Utility Scripts for inference and sim-to-real validation.

The pollen-robotics/microduck_rl repository provides a specialized reinforcement learning framework built atop mjlab and rsl_rl specifically designed for quadruped robot control. This architecture follows a modular design pattern that separates scene definition, Markov Decision Process (MDP) logic, actuator physics, and deployment pipelines. Understanding these core components is essential for researchers implementing sim-to-real transfer or extending the framework with new locomotion tasks.

Environment Configurations

Task definitions in Microduck RL are encapsulated in dedicated *_env_cfg.py files that assemble the complete simulation scene. Each configuration—such as microduck_velocity_env_cfg.py for walking or microduck_standup_env_cfg.py for recovery behaviors—registers the robot model, domain randomization parameters, observation noise layouts, and command slots.

These configuration files serve as the single source of truth for the training scene. They define the 61-dimensional observation space contract and specify which terrain properties, friction ranges, and initial conditions the policy must handle. When instantiating an environment, the configuration factory functions return a structured object compatible with mjlab's ManagerBasedRlEnv.

from src.mjlab_microduck.tasks.microduck_velocity_env_cfg import make_microduck_velocity_env_cfg

cfg = make_microduck_velocity_env_cfg(play=False, rough=False)

from mjlab.envs.manager_based_rl_env import ManagerBasedRlEnv
env = ManagerBasedRlEnv(cfg)

MDP Functions and Rewards

The heart of the learning problem lives in src/mjlab_microduck/tasks/mdp.py. This module encapsulates the Markov Decision Process by defining observations, actions, reward terms, and termination conditions. It also patches the reward manager for NaN safety, ensuring training stability even when simulation instabilities occur.

Key utilities include _servo_joint_ids for joint indexing, reset_with_forward_velocity for initialization, and task-specific reward terms such as upright_progress for posture shaping and leg_action_rate_l2 for regularization penalties. The module implements fall detection logic and recovery bonuses that encourage robust policies.

from src.mjlab_microduck.tasks.mdp import upright_progress, leg_action_rate_l2

# Compute upright shaping reward

upright_reward = upright_progress(env)

# Add regularization penalty to minimize jerky movements

leg_rate_penalty = leg_action_rate_l2(env)

Custom Actuator Models

Microduck RL uses a custom BAM (Backlash-Aware Module) actuator defined in src/mjlab_microduck/actuator/friction_dr_bam.py. The FrictionDRBamActuator class implements per-environment friction randomization through a friction_scale parameter, while BacklashEncoderBamActuator models encoder feedback transmitted through mechanical backlash.

These actuator models provide realistic torque-control dynamics critical for sim-to-real transfer. The domain randomization capabilities include friction scaling and backlash modeling, ensuring policies trained in simulation remain robust when deployed on physical hardware with imperfect actuators.

Training and Deployment Pipeline

Training CLI

The entry point uv run train ... forwards to mjlab's trainer via src/mjlab_microduck/train_cli.py. This script exists primarily to maintain a stable console-script name while delegating the actual PPO training logic to the underlying mjlab framework. It handles optional HuggingFace job submission and loads the selected environment configuration based on task identifiers.

ONNX Export

After training completes, scripts/export.py invokes the mjlab exporter with a patched get_base_metadata function. This patch filters out passive_* joints and bakes the observation normalizer directly into the exported ONNX file, ensuring the deployed policy matches the 61-dimensional observation contract expected by the real robot runtime.


# Export trained policy with baked normalizer

uv run scripts/export.py <TASK_ID> --wandb-run-path <entity/project/run_id>

# Produces out.onnx ready for deployment

Utility and Validation Scripts

The repository includes specialized utilities for debugging and validation:

  • scripts/infer_policy.py: Runs CPU-only inference on exported ONNX models, loading the policy and stepping the environment to print episode statistics without requiring GPU training hardware.
  • scripts/testbench_sim2real.py: Validates policy behavior by comparing simulation rollouts against real-robot data, identifying discrepancies in dynamics or sensor responses.
  • view_slope_terrain.py: Visualizes terrain geometry for debugging environment setups.

# Run inference on exported model

uv run scripts/infer_policy.py --walking out.onnx

Summary

The Microduck RL architecture provides a complete pipeline from simulation to physical deployment:

  • Environment Configurations define task-specific scenes and observation contracts in *_env_cfg.py files
  • MDP Functions implement reward logic, termination conditions, and NaN safety guards in tasks/mdp.py
  • BAM Actuator Models provide realistic physics with friction domain randomization in actuator/friction_dr_bam.py
  • Training CLI offers a stable entry point to mjlab's PPO trainer through train_cli.py
  • ONNX Export produces deployment-ready policies with baked normalizers and filtered passive joints via scripts/export.py
  • Utility Scripts enable rapid debugging, inference, and sim-to-real validation without full training runs

Frequently Asked Questions

What is the observation dimension in Microduck RL policies?

The framework enforces a 61-dimensional observation contract throughout the architecture. This dimensionality is maintained in environment configurations, baked into exported ONNX models, and expected by the real robot runtime. The observation vector includes proprioceptive data, command inputs, and terrain features.

How does Microduck RL handle domain randomization?

Domain randomization occurs at multiple levels: Environment configurations inject observation noise and terrain variations, while the FrictionDRBamActuator in friction_dr_bam.py applies per-environment friction scaling through the friction_scale parameter. Backlash modeling through BacklashEncoderBamActuator further randomizes sensor feedback to simulate mechanical imperfections.

What safety mechanisms prevent training instabilities?

The tasks/mdp.py module patches the reward manager to guard against NaN values that can destabilize PPO training. Additionally, termination conditions detect falls and unsafe configurations, triggering episode resets before gradients corrupt the policy. The upright shaping reward also encourages stable postures that avoid simulation instabilities.

How is the trained policy exported for real robot deployment?

The scripts/export.py utility processes the trained checkpoint using a patched get_base_metadata function that removes passive joints from the joint list and bakes the observation normalizer statistics directly into the ONNX graph. This ensures the exported out.onnx file requires no external normalization files and matches the exact observation format expected by the embedded robot controller.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →