# Core MDP Functions in Microduck RL: Architecture and Key Files Explained

> Explore the core MDP functions in Microduck RL, including reward terms, state-reset utilities, and joint-index helpers, all located in src/mjlab_microduck/tasks/mdp.py for efficient robot learning.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: architecture
- Published: 2026-09-08

---

**The MDP functions in Microduck RL are centralized in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py), which houses all reward terms, state-reset utilities, and joint-index helpers used to define the robot's learning objectives.**

Microduck RL is an open-source reinforcement learning framework for robotics developed by Pollen Robotics. All Markov Decision Process (MDP) logic—including reward shaping, penalty calculations, and episode initialization—resides in a single, well-organized Python module that serves as the definitive source for the learning problem definition.

## Central MDP Module ([`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py))

The file [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) contains the complete implementation of Microduck RL's MDP functions. This module patches the reward manager for NaN-safe operations and provides specialized utilities for the 14-servo robot configuration.

### Reward and Penalty Functions

The module implements a comprehensive collection of reward terms that guide the robot's behavior. Key functions include:

- **`leg_action_rate_l2`** – Applies an L2 penalty on leg action changes to encourage smooth movements
- **`neck_action_rate_l2`** – Penalizes rapid neck movements for stable head positioning
- **`body_upright_linear`** – Provides linear rewards for maintaining upright posture
- **`fallen_state_penalty`** – Applies a per-step tax while the robot is in a fallen state
- **`recovery_success`** – Signals successful recovery from fallen states
- **`wheel_glide_reward`** – Rewards efficient wheel-based locomotion

These functions are designed to work with the canonical 14-servo layout, ensuring consistent reward calculation across different task configurations.

### State-Reset Utilities

Environment initialization relies on several reset helpers that prepare the robot for training episodes:

- **`reset_with_forward_velocity`** – Initializes the robot with random forward velocity within specified ranges
- **`reset_action_history`** – Clears tracked action sequences for temporal penalty calculations
- **`reset_rolling_entry`** – Prepares wheel spin states for rolling tasks

These utilities are called from environment callbacks to standardize episode starts across different task families.

### Joint-Index Abstractions

The module provides joint-index helpers that abstract away passive joints:

- **`_servo_joint_ids`** – Returns indices for active servo joints only
- **`_servo_joint_pos`** – Retrieves positions for the 14-servo canonical layout

These helpers ensure every reward term operates on the correct joint subset, preventing passive joint interference in reward calculations.

### Potential-Based Shaping

Dense reward gradients are provided through shaping functions:

- **`upright_progress`** – Tracks posture improvement toward vertical orientation
- **`height_progress`** – Measures elevation changes toward target heights
- **`standing_composite_score`** – Calculates a composite metric combining height, upright alignment, and pose accuracy

## Supporting Files and Task Registration

While [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) contains the core logic, several surrounding files integrate these functions into the training pipeline.

### Task Registration ([`tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tasks/__init__.py))

The [`src/mjlab_microduck/tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/__init__.py) file registers each task family with the mjlab system. It imports symbols from [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) and exposes them to the task registry, making the MDP functions available to environment configurations without direct file imports.

### Environment Configurations (`*_env_cfg.py`)

Individual task configurations reside in files like [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py) and [`microduck_roller_standup_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_roller_standup_env_cfg.py). These files import specific MDP functions to build task-specific reward managers:

```python
from mjlab_microduck.tasks import mdp

reward_terms = [
    mdp.leg_action_rate_l2,          # L2 penalty on leg action changes

    mdp.neck_action_rate_l2,         # L2 penalty on neck action changes

    mdp.body_upright_linear,         # Linear upright reward

    mdp.fallen_state_penalty,        # Per-step tax while fallen

]

```

### Specialized Wrappers ([`backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/backlash.py) and [`symmetry.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/symmetry.py))

Two additional files extend MDP functionality:

- **[`src/mjlab_microduck/tasks/backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/backlash.py)** – Wraps environment configurations into "-Backlash-" variants while preserving the same MDP functions from [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py)
- **[`src/mjlab_microduck/tasks/symmetry.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/symmetry.py)** – Provides symmetry-related utilities used by specific MDP reward terms to encourage balanced robot behavior

## Practical Implementation Examples

Initialize environments with forward velocity using the reset utilities:

```python
def post_reset(env, env_ids):
    mdp.reset_with_forward_velocity(
        env,
        env_ids,
        velocity_range=(0.3, 0.8),           # random forward speed in m/s

        fraction_stages=[{"step":0, "fraction":0.8},
                         {"step":2000*24, "fraction":0.0}],
    )

```

Calculate composite standing scores for dense reward signals:

```python
score = mdp.standing_composite_score(
    env,
    target_height=0.115,
    height_std=0.01,
    upright_std=0.1,
    pose_std=0.05,
    joint_indices=list(range(14)),   # all servo joints

)

```

## Summary

- **[`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py)** serves as the single source of truth for all MDP functions in Microduck RL
- The module contains NaN-safe reward manager patches, joint-index abstractions for 14-servo layouts, and comprehensive state-reset utilities
- Reward terms include action rate penalties, upright bonuses, and potential-based shaping functions for dense gradients
- Environment configurations import these functions through [`tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tasks/__init__.py) while specialized wrappers in [`backlash.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/backlash.py) and [`symmetry.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/symmetry.py) extend functionality
- All reset utilities support velocity initialization, action history management, and wheel spin preparation

## Frequently Asked Questions

### Where are reward functions defined in Microduck RL?

All reward functions are defined in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py), including leg and neck action rate penalties, upright rewards, and fallen state penalties. This centralization ensures consistent reward logic across all task configurations.

### How does Microduck RL handle joint indexing for rewards?

The MDP module provides helper functions `_servo_joint_ids` and `_servo_joint_pos` that filter passive joints and return indices for the canonical 14-servo layout. These helpers ensure reward calculations only consider active joints relevant to the learning objective.

### What utilities exist for resetting environments in Microduck RL?

The framework provides `reset_with_forward_velocity` for velocity initialization, `reset_action_history` for clearing temporal buffers, and `reset_rolling_entry` for wheel spin preparation. These are called from environment post-reset callbacks to standardize episode starts.

### How are MDP functions registered with the training pipeline?

The [`src/mjlab_microduck/tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/__init__.py) file registers task families and imports symbols from [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py), making them available to environment configuration files. Individual task configs then import specific functions from the `mdp` module to construct their reward managers.