Core MDP Functions in Microduck RL: Architecture and Key Files Explained

The MDP functions in Microduck RL are centralized in src/mjlab_microduck/tasks/mdp.py, which houses all reward terms, state-reset utilities, and joint-index helpers used to define the robot's learning objectives.

Microduck RL is an open-source reinforcement learning framework for robotics developed by Pollen Robotics. All Markov Decision Process (MDP) logic—including reward shaping, penalty calculations, and episode initialization—resides in a single, well-organized Python module that serves as the definitive source for the learning problem definition.

Central MDP Module (src/mjlab_microduck/tasks/mdp.py)

The file src/mjlab_microduck/tasks/mdp.py contains the complete implementation of Microduck RL's MDP functions. This module patches the reward manager for NaN-safe operations and provides specialized utilities for the 14-servo robot configuration.

Reward and Penalty Functions

The module implements a comprehensive collection of reward terms that guide the robot's behavior. Key functions include:

  • leg_action_rate_l2 – Applies an L2 penalty on leg action changes to encourage smooth movements
  • neck_action_rate_l2 – Penalizes rapid neck movements for stable head positioning
  • body_upright_linear – Provides linear rewards for maintaining upright posture
  • fallen_state_penalty – Applies a per-step tax while the robot is in a fallen state
  • recovery_success – Signals successful recovery from fallen states
  • wheel_glide_reward – Rewards efficient wheel-based locomotion

These functions are designed to work with the canonical 14-servo layout, ensuring consistent reward calculation across different task configurations.

State-Reset Utilities

Environment initialization relies on several reset helpers that prepare the robot for training episodes:

  • reset_with_forward_velocity – Initializes the robot with random forward velocity within specified ranges
  • reset_action_history – Clears tracked action sequences for temporal penalty calculations
  • reset_rolling_entry – Prepares wheel spin states for rolling tasks

These utilities are called from environment callbacks to standardize episode starts across different task families.

Joint-Index Abstractions

The module provides joint-index helpers that abstract away passive joints:

  • _servo_joint_ids – Returns indices for active servo joints only
  • _servo_joint_pos – Retrieves positions for the 14-servo canonical layout

These helpers ensure every reward term operates on the correct joint subset, preventing passive joint interference in reward calculations.

Potential-Based Shaping

Dense reward gradients are provided through shaping functions:

  • upright_progress – Tracks posture improvement toward vertical orientation
  • height_progress – Measures elevation changes toward target heights
  • standing_composite_score – Calculates a composite metric combining height, upright alignment, and pose accuracy

Supporting Files and Task Registration

While mdp.py contains the core logic, several surrounding files integrate these functions into the training pipeline.

Task Registration (tasks/__init__.py)

The src/mjlab_microduck/tasks/__init__.py file registers each task family with the mjlab system. It imports symbols from mdp.py and exposes them to the task registry, making the MDP functions available to environment configurations without direct file imports.

Environment Configurations (*_env_cfg.py)

Individual task configurations reside in files like microduck_velocity_env_cfg.py and microduck_roller_standup_env_cfg.py. These files import specific MDP functions to build task-specific reward managers:

from mjlab_microduck.tasks import mdp

reward_terms = [
    mdp.leg_action_rate_l2,          # L2 penalty on leg action changes

    mdp.neck_action_rate_l2,         # L2 penalty on neck action changes

    mdp.body_upright_linear,         # Linear upright reward

    mdp.fallen_state_penalty,        # Per-step tax while fallen

]

Specialized Wrappers (backlash.py and symmetry.py)

Two additional files extend MDP functionality:

Practical Implementation Examples

Initialize environments with forward velocity using the reset utilities:

def post_reset(env, env_ids):
    mdp.reset_with_forward_velocity(
        env,
        env_ids,
        velocity_range=(0.3, 0.8),           # random forward speed in m/s

        fraction_stages=[{"step":0, "fraction":0.8},
                         {"step":2000*24, "fraction":0.0}],
    )

Calculate composite standing scores for dense reward signals:

score = mdp.standing_composite_score(
    env,
    target_height=0.115,
    height_std=0.01,
    upright_std=0.1,
    pose_std=0.05,
    joint_indices=list(range(14)),   # all servo joints

)

Summary

  • src/mjlab_microduck/tasks/mdp.py serves as the single source of truth for all MDP functions in Microduck RL
  • The module contains NaN-safe reward manager patches, joint-index abstractions for 14-servo layouts, and comprehensive state-reset utilities
  • Reward terms include action rate penalties, upright bonuses, and potential-based shaping functions for dense gradients
  • Environment configurations import these functions through tasks/__init__.py while specialized wrappers in backlash.py and symmetry.py extend functionality
  • All reset utilities support velocity initialization, action history management, and wheel spin preparation

Frequently Asked Questions

Where are reward functions defined in Microduck RL?

All reward functions are defined in src/mjlab_microduck/tasks/mdp.py, including leg and neck action rate penalties, upright rewards, and fallen state penalties. This centralization ensures consistent reward logic across all task configurations.

How does Microduck RL handle joint indexing for rewards?

The MDP module provides helper functions _servo_joint_ids and _servo_joint_pos that filter passive joints and return indices for the canonical 14-servo layout. These helpers ensure reward calculations only consider active joints relevant to the learning objective.

What utilities exist for resetting environments in Microduck RL?

The framework provides reset_with_forward_velocity for velocity initialization, reset_action_history for clearing temporal buffers, and reset_rolling_entry for wheel spin preparation. These are called from environment post-reset callbacks to standardize episode starts.

How are MDP functions registered with the training pipeline?

The src/mjlab_microduck/tasks/__init__.py file registers task families and imports symbols from mdp.py, making them available to environment configuration files. Individual task configs then import specific functions from the mdp module to construct their reward managers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →