# Custom MDP Functions in Microduck RL: A Complete Guide to Robot Learning Rewards and Terminations

> Explore custom MDP functions in Microduck RL. Learn how to define robot learning rewards and terminations using over 60 PyTorch functions in this comprehensive guide.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-02

---

**The Microduck RL repository defines all task-specific reward, termination, and utility logic in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py), with over 60 pure PyTorch functions that can be registered via `RewardManager` or `EventManager` in environment configurations.**

If you're building locomotion or manipulation policies for Pollen Robotics' Microduck platform, you'll spend most of your time configuring **custom MDP functions in Microduck RL**. These functions implement the dense reward shaping, safety penalties, and termination conditions that translate simulation-trained policies to hardware. This guide catalogues every major function category with exact source locations, signatures, and practical integration patterns.

---

## Environment-Wide Helper Functions

Before computing rewards, most MDP terms need filtered access to the robot's state. The helper functions in this category handle **servo-only joint indexing**—critical because the Microduck MJCF excludes `passive_*` joints from control.

| Function | Location | Purpose |
|----------|----------|---------|
| `_servo_joint_ids(env, asset)` | `mdp.py:26` | Returns indices of 14 controllable servo joints |
| `_servo_joint_pos(env, asset)` | `mdp.py:46` | Reads current joint positions (servo view) |
| `_servo_joint_vel(env, asset)` | `mdp.py:50` | Reads joint velocities (servo view) |
| `_servo_default_joint_pos(env, asset)` | `mdp.py:54` | Retrieves default/home joint angles |

These helpers enforce the **61-dimensional observation contract** that keeps sim-to-real transfer stable. Always use `_servo_joint_ids()` rather than raw MuJoCo joint indices to avoid indexing into passive roller or fixed joints.

```python
from mjlab_microduck.tasks.mdp import _servo_joint_pos, _servo_default_joint_pos

# Compute pose error for reward shaping

current_pos = _servo_joint_pos(env, env.asset)
default_pos = _servo_default_joint_pos(env, env.asset)
pose_error = torch.norm(current_pos - default_pos, dim=1)

```

---

## Joint-Space Penalties and Regularizers

Smooth, efficient motion requires penalizing high-frequency control signals and mechanical stress. The **custom MDP functions in Microduck RL** provide granular regularization across legs, neck, and individual joint groups.

### Action Rate and Acceleration Penalties

- **`leg_action_rate_l2(...)`** (`mdp.py:53`): L2 penalty on Δaction for leg joints (discourages jitter)
- **`neck_action_rate_l2(...)`** (`mdp.py:93`): L2 penalty on neck-joint action changes (head stability)
- **`leg_action_acceleration_l2(...)`** (`mdp.py:29`): Second-order L2 penalty on leg action acceleration
- **`neck_action_acceleration_l2(...)`** (`mdp.py:68`): Neck action acceleration penalty

### Velocity and Torque Penalties

- **`leg_joint_vel_l2(...)`** (`mdp.py:34`): Smoothness regularizer on leg joint velocities
- **`neck_joint_vel_l2(...)`** (`mdp.py:8`): Head stability via neck velocity penalty
- **`hip_pitch_knee_vel_l2(...)`** (`mdp.py:42`): Sagittal-plane gait regularizer
- **`joint_torques_l2(...)`** (`mdp.py:79`): Energy efficiency penalty on actuator torques
- **`joint_torque_rate_l2(...)`** (`mdp.py:16`): Gearbox protection via torque change penalty

### Position and Limit Penalties

- **`joint_deviation_l1(...)`** (`mdp.py:72`): L1 penalty on deviation from home pose
- **`joint_pos_limit_proximity(...)`** (`mdp.py:89`): Soft penalty as joints approach hard limits
- **`neck_joint_pos_l2(...)`** (`mdp.py:55`): Head pose regularization
- **`bilateral_symmetry_penalty(...)`** (`mdp.py:33`): L1 penalty on left-right leg asymmetry

```python

# Typical regularizer stack for velocity walking

from mjlab_microduck.tasks.mdp import (
    leg_action_rate_l2,
    leg_joint_vel_l2,
    joint_torques_l2,
    joint_deviation_l1,
)

reward_manager = _RewardManager(
    terms=[
        ("action_rate", leg_action_rate_l2, {"weight": -0.01}),
        ("joint_vel", leg_joint_vel_l2, {"weight": -0.001}),
        ("torque", joint_torques_l2, {"weight": -0.0001}),
        ("pose_dev", joint_deviation_l1, {"weight": -0.1}),
    ]
)

```

---

## Body Orientation Rewards and Uprightness Shaping

Maintaining vertical posture is fundamental to bipedal and roller locomotion. These functions provide **multiple mathematical formulations** for upright rewards, allowing task-specific gradient properties.

### Core Uprightness Functions

- **`_fallen_mask(env, asset, gate_z, gate_tilt)`** (`mdp.py:7`): Binary mask (1=fallen) gating other rewards
- **`body_upright_linear(...)`** (`mdp.py:89`): Linear `cos(tilt)` reward with gradient everywhere
- **`body_upright_gaussian(...)`** (`mdp.py:11`): Gaussian reward with sharp peak at vertical
- **`upright_gaussian_at_height(...)`** (`mdp.py:35`): Height-gated upright Gaussian (prevents crouch-cheating)

### Composite Standing Rewards

- **`standing_composite_score(...)`** (`mdp.py:10`): Multiplicative product of height, upright, and pose Gaussian scores
- **`standing_success_bonus(...)`** (`mdp.py:56`): Sparse binary bonus when all tolerances satisfied

### Potential-Based Shaping

- **`upright_progress(...)`** (`mdp.py:47`): Potential-based shaping on trunk tilt cosine
- **`height_progress(...)`** (`mdp.py:75`): Potential-based shaping on capped trunk height

### Recovery and Penalty Terms

- **`fallen_state_penalty(...)`** (`mdp.py:5`): Per-step tax while fallen
- **`recovery_success(...)`** (`mdp.py:46`): Sparse bounty on regaining uprightness
- **`body_ang_vel_at_height(...)`** (`mdp.py:64`): Angular velocity cost (height-gated)

```python

# Composite standing reward with height gating

from mjlab_microduck.tasks.mdp import (
    standing_composite_score,
    standing_success_bonus,
    fallen_state_penalty,
)

reward_manager = _RewardManager(
    terms=[
        ("standing", standing_composite_score, {
            "height_target": 0.32,
            "upright_sigma": 0.2,
            "pose_sigma": 0.3,
        }),
        ("stand_bonus", standing_success_bonus, {"tolerance": 0.05}),
        ("fallen_tax", fallen_state_penalty, {"gate_tilt_above_deg": 45.0}),
    ]
)

```

---

## Center-of-Mass Height and Velocity Rewards

Vertical control enables jumping, crouching, and terrain adaptation. The **custom MDP functions in Microduck RL** for CoM control support both fixed targets and **phase-varying trajectories**.

- **`com_upward_velocity(...)`** (`mdp.py:1`): Positive reward for upward CoM velocity below target height
- **`com_height_target(...)`** (`mdp.py:33`): Gaussian reward for CoM in desired height window
- **`crouch_height_target(...)`** (`mdp.py:78`): Trapezoidal phase-based height target for crouch-glide tasks
- **`crouch_glide_reward_from_values(...)`** (`mdp.py:13`): Gaussian reward matching `crouch_height_target`

```python

# Crouch-glide task: periodic height modulation

from mjlab_microduck.tasks.mdp import (
    crouch_height_target,
    crouch_glide_reward_from_values,
)

# Height target modulates over episode progress

height_target = crouch_height_target(
    env,
    low_height=0.15,
    high_height=0.30,
    crouch_fraction=0.4,
    glide_fraction=0.4,
)

# Reward matches the target

reward_manager = _RewardManager(
    terms=[
        ("crouch_glide", crouch_glide_reward_from_values, {
            "height_target": height_target,
            "sigma": 0.05,
        }),
    ]
)

```

---

## Wheel-Related Rewards for Roller Locomotion

The Microduck platform includes **passive roller wheels** for efficient forward motion. These functions shape policies to exploit wheel mechanics while preventing pathological slip.

- **`wheel_glide_reward(...)`** (`mdp.py:94`): Reward proportional to passive wheel spin (capped)
- **`wheel_speed_reward(...)`** (`mdp.py:74`): Aligned wheel spin with commanded forward speed
- **`descent_speed_reward(...)`** (`mdp.py:36`): Linear forward velocity reward for slope descent

```python

# Roller configuration: exploit wheel glide with speed alignment

from mjlab_microduck.tasks.mdp import wheel_glide_reward, wheel_speed_reward

reward_manager = _RewardManager(
    terms=[
        ("glide", wheel_glide_reward, {"cap_speed": 0.35, "weight": 1.0}),
        ("speed_match", wheel_speed_reward, {
            "command_key": "base_velocity",
            "bidirectional": False,
        }),
    ]
)

```

---

## Termination and Safety Events

Robust training requires **early termination of failed rollouts** and protection against numerical instability. These custom MDP functions integrate with Isaac Lab's `EventManager`.

| Function | Location | Trigger Condition |
|----------|----------|-------------------|
| `fallen_too_long(...)` | `mdp.py:37` | Fallen duration exceeds threshold |
| `robot_state_is_nan(...)` | `mdp.py:63` | Any state variable is NaN |
| `root_height_below(...)` | `mdp.py:17` | Trunk falls below world-z threshold |
| `body_impact_cost(...)` | `mdp.py:45` | Contact force exceeds protected-body threshold |
| `contact_frequency_penalty(...)` | `mdp.py:62` | Excessive foot contact changes |

### Contact and Grounding Rewards

- **`feet_grounded_reward(...)`** (`mdp.py:25`): Small positive reward per ground-contact foot
- **`feet_air_time_upright(...)`** (`mdp.py:27`): Standard air-time reward (zeroed while fallen)
- **`feet_flat_penalty(...)`** (`mdp.py:64`): Penalty for non-parallel foot-ground contact
- **`feet_tiptoe_alignment(...)`** (`mdp.py:15`): Reward for foot x-axis pointing downward

```python

# Safety-focused termination configuration

from mjlab_microduck.tasks.mdp import (
    fallen_too_long,
    robot_state_is_nan,
    root_height_below,
)

event_manager = _EventManager(
    terminations=[
        ("fallen_timeout", fallen_too_long, {
            "gate_z_below": 0.10,
            "max_duration_s": 4.0,
        }),
        ("nan_state", robot_state_is_nan, {}),
        ("fall_off_slope", root_height_below, {"min_height": -0.5}),
    ]
)

```

---

## Phase-Based and Ground-Pick Rewards

Complex behaviors like **sit-stand transitions and mouth-ground interaction** require temporal task decomposition. These functions provide phase-conditioned reward shaping.

### Mouth-Ground Interaction

- **`mouth_ground_proximity(...)`** (`mdp.py:38`): Gaussian reward for mouth tip approaching ground (phase-weighted)
- **`mouth_perpendicular_to_ground(...)`** (`mdp.py:66`): Rewards downward-pointing mouth x-axis during approach

### Sit-Stand Transitions

- **`sit_grounded(...)`** (`mdp.py:92`): Reward for trunk-ground contact while upright
- **`sit_stability(...)`** (`mdp.py:44`): Low angular velocity reward during sit phase
- **`phase_height_track(...)`** (`mdp.py:30`): Sinusoidal height target following

### Pose Interpolation Rewards

- **`pose_target_match(...)`** (`mdp.py:59`): Gaussian reward vs fixed target pose
- **`interpolated_pose_target_match(...)`** (`mdp.py:91`): Interpolated between two poses
- **`multistage_pose_target_match(...)`** (`mdp.py:25`): Arbitrary waypoint sequence (stand → fold → sit)
- **`interpolated_height_target(...)`** (`mdp.py:22`): Height interpolation across waypoints

```python

# Three-stage sit-stand-fold behavior

from mjlab_microduck.tasks.mdp import multistage_pose_target_match

pose_waypoints = [
    {"fraction": 0.0, "pose": "standing"},
    {"fraction": 0.4, "pose": "sitting"},
    {"fraction": 0.7, "pose": "folded"},
]

reward_manager = _RewardManager(
    terms=[
        ("multistage_pose", multistage_pose_target_match, {
            "waypoints": pose_waypoints,
            "sigma": 0.2,
        }),
    ]
)

```

---

## Reset Helpers for Curriculum and Warm-Start

Training stability often requires **non-uniform initialization**. These functions hook into environment reset callbacks.

- **`reset_with_forward_velocity(...)`** (`mdp.py:58`): Warm-start subset with random forward speed
- **`reset_action_history(...)`** (`mdp.py:29`): Clear cached action-rate and acceleration buffers
- **`reset_rolling_entry(...)`** (`mdp.py:55`): Initialize wheel velocities for non-slipping rollout

```python
from mjlab_microduck.tasks.mdp import reset_with_forward_velocity

def curriculum_reset(env, env_ids):
    # Progressive velocity curriculum: 50% at speed after step 5M

    step = env.common_step_counter
    fraction = 0.0 if step < 5_000_000 else 0.5
    
    reset_with_forward_velocity(
        env,
        env_ids,
        velocity_range=(0.3, 0.8),
        fraction_stages=[{"step": 0, "fraction": fraction}],
    )

```

---

## Summary

The **custom MDP functions in Microduck RL** provide a complete toolkit for robot learning:

- **Modular design**: All functions live in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) and register via `RewardManager`/`EventManager`
- **Servo-safe indexing**: Helper functions exclude passive joints to maintain the 61-dim observation contract
- **Multiple mathematical forms**: Linear, Gaussian, and L1/L2 penalties for gradient tuning
- **Phase-aware shaping**: Interpolation and waypoint support for complex behaviors
- **Hardware-aligned safety**: NaN detection, impact costs, and fall termination

By combining these primitives in task configurations, you can train policies that transfer reliably from Isaac Lab simulation to the physical Microduck robot.

---

## Frequently Asked Questions

### How do I add a custom reward to my Microduck RL environment?

Create or edit a task configuration file (e.g., [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py)), import your desired function from `mjlab_microduck.tasks.mdp`, and append a tuple to the `RewardManager` terms list: `("term_name", function, {params})`.

### What is the difference between `body_upright_linear` and `body_upright_gaussian`?

`body_upright_linear` returns `cos(tilt)` with non-zero gradient everywhere, providing continuous learning signal even when far from vertical. `body_upright_gaussian` uses `exp(-tilt²/σ²)` with sharp peak at vertical and near-zero gradient far from target—better for fine-tuning once approximately upright.

### Why do MDP functions use `_servo_joint_ids` instead of direct MuJoCo indexing?

The Microduck MJCF includes passive roller joints that must not be controlled or observed directly. `_servo_joint_ids` filters to the 14 actuated servo joints, ensuring policies respect the hardware's actual degrees of freedom and maintain sim-to-real consistency.

### How can I implement a curriculum that increases task difficulty over training?

Use `reset_with_forward_velocity` or time-varying reward weights in your configuration. Pass `fraction_stages` to progressively enable faster initial conditions, or modulate reward `weight` parameters via the `RewardManager`'s runtime interface based on `env.common_step_counter`.