Core MDP Functions in Microduck RL: Architecture and Key Files Explained
The MDP functions in Microduck RL are centralized in src/mjlab_microduck/tasks/mdp.py, which houses all reward terms, state-reset utilities, and joint-index helpers used to define the robot's learning objectives.
Microduck RL is an open-source reinforcement learning framework for robotics developed by Pollen Robotics. All Markov Decision Process (MDP) logic—including reward shaping, penalty calculations, and episode initialization—resides in a single, well-organized Python module that serves as the definitive source for the learning problem definition.
Central MDP Module (src/mjlab_microduck/tasks/mdp.py)
The file src/mjlab_microduck/tasks/mdp.py contains the complete implementation of Microduck RL's MDP functions. This module patches the reward manager for NaN-safe operations and provides specialized utilities for the 14-servo robot configuration.
Reward and Penalty Functions
The module implements a comprehensive collection of reward terms that guide the robot's behavior. Key functions include:
leg_action_rate_l2– Applies an L2 penalty on leg action changes to encourage smooth movementsneck_action_rate_l2– Penalizes rapid neck movements for stable head positioningbody_upright_linear– Provides linear rewards for maintaining upright posturefallen_state_penalty– Applies a per-step tax while the robot is in a fallen staterecovery_success– Signals successful recovery from fallen stateswheel_glide_reward– Rewards efficient wheel-based locomotion
These functions are designed to work with the canonical 14-servo layout, ensuring consistent reward calculation across different task configurations.
State-Reset Utilities
Environment initialization relies on several reset helpers that prepare the robot for training episodes:
reset_with_forward_velocity– Initializes the robot with random forward velocity within specified rangesreset_action_history– Clears tracked action sequences for temporal penalty calculationsreset_rolling_entry– Prepares wheel spin states for rolling tasks
These utilities are called from environment callbacks to standardize episode starts across different task families.
Joint-Index Abstractions
The module provides joint-index helpers that abstract away passive joints:
_servo_joint_ids– Returns indices for active servo joints only_servo_joint_pos– Retrieves positions for the 14-servo canonical layout
These helpers ensure every reward term operates on the correct joint subset, preventing passive joint interference in reward calculations.
Potential-Based Shaping
Dense reward gradients are provided through shaping functions:
upright_progress– Tracks posture improvement toward vertical orientationheight_progress– Measures elevation changes toward target heightsstanding_composite_score– Calculates a composite metric combining height, upright alignment, and pose accuracy
Supporting Files and Task Registration
While mdp.py contains the core logic, several surrounding files integrate these functions into the training pipeline.
Task Registration (tasks/__init__.py)
The src/mjlab_microduck/tasks/__init__.py file registers each task family with the mjlab system. It imports symbols from mdp.py and exposes them to the task registry, making the MDP functions available to environment configurations without direct file imports.
Environment Configurations (*_env_cfg.py)
Individual task configurations reside in files like microduck_velocity_env_cfg.py and microduck_roller_standup_env_cfg.py. These files import specific MDP functions to build task-specific reward managers:
from mjlab_microduck.tasks import mdp
reward_terms = [
mdp.leg_action_rate_l2, # L2 penalty on leg action changes
mdp.neck_action_rate_l2, # L2 penalty on neck action changes
mdp.body_upright_linear, # Linear upright reward
mdp.fallen_state_penalty, # Per-step tax while fallen
]
Specialized Wrappers (backlash.py and symmetry.py)
Two additional files extend MDP functionality:
src/mjlab_microduck/tasks/backlash.py– Wraps environment configurations into "-Backlash-" variants while preserving the same MDP functions frommdp.pysrc/mjlab_microduck/tasks/symmetry.py– Provides symmetry-related utilities used by specific MDP reward terms to encourage balanced robot behavior
Practical Implementation Examples
Initialize environments with forward velocity using the reset utilities:
def post_reset(env, env_ids):
mdp.reset_with_forward_velocity(
env,
env_ids,
velocity_range=(0.3, 0.8), # random forward speed in m/s
fraction_stages=[{"step":0, "fraction":0.8},
{"step":2000*24, "fraction":0.0}],
)
Calculate composite standing scores for dense reward signals:
score = mdp.standing_composite_score(
env,
target_height=0.115,
height_std=0.01,
upright_std=0.1,
pose_std=0.05,
joint_indices=list(range(14)), # all servo joints
)
Summary
src/mjlab_microduck/tasks/mdp.pyserves as the single source of truth for all MDP functions in Microduck RL- The module contains NaN-safe reward manager patches, joint-index abstractions for 14-servo layouts, and comprehensive state-reset utilities
- Reward terms include action rate penalties, upright bonuses, and potential-based shaping functions for dense gradients
- Environment configurations import these functions through
tasks/__init__.pywhile specialized wrappers inbacklash.pyandsymmetry.pyextend functionality - All reset utilities support velocity initialization, action history management, and wheel spin preparation
Frequently Asked Questions
Where are reward functions defined in Microduck RL?
All reward functions are defined in src/mjlab_microduck/tasks/mdp.py, including leg and neck action rate penalties, upright rewards, and fallen state penalties. This centralization ensures consistent reward logic across all task configurations.
How does Microduck RL handle joint indexing for rewards?
The MDP module provides helper functions _servo_joint_ids and _servo_joint_pos that filter passive joints and return indices for the canonical 14-servo layout. These helpers ensure reward calculations only consider active joints relevant to the learning objective.
What utilities exist for resetting environments in Microduck RL?
The framework provides reset_with_forward_velocity for velocity initialization, reset_action_history for clearing temporal buffers, and reset_rolling_entry for wheel spin preparation. These are called from environment post-reset callbacks to standardize episode starts.
How are MDP functions registered with the training pipeline?
The src/mjlab_microduck/tasks/__init__.py file registers task families and imports symbols from mdp.py, making them available to environment configuration files. Individual task configs then import specific functions from the mdp module to construct their reward managers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →