Custom MDP Functions in Microduck RL: A Complete Guide to Robot Learning Rewards and Terminations

The Microduck RL repository defines all task-specific reward, termination, and utility logic in src/mjlab_microduck/tasks/mdp.py, with over 60 pure PyTorch functions that can be registered via RewardManager or EventManager in environment configurations.

If you're building locomotion or manipulation policies for Pollen Robotics' Microduck platform, you'll spend most of your time configuring custom MDP functions in Microduck RL. These functions implement the dense reward shaping, safety penalties, and termination conditions that translate simulation-trained policies to hardware. This guide catalogues every major function category with exact source locations, signatures, and practical integration patterns.


Environment-Wide Helper Functions

Before computing rewards, most MDP terms need filtered access to the robot's state. The helper functions in this category handle servo-only joint indexing—critical because the Microduck MJCF excludes passive_* joints from control.

Function Location Purpose
_servo_joint_ids(env, asset) mdp.py:26 Returns indices of 14 controllable servo joints
_servo_joint_pos(env, asset) mdp.py:46 Reads current joint positions (servo view)
_servo_joint_vel(env, asset) mdp.py:50 Reads joint velocities (servo view)
_servo_default_joint_pos(env, asset) mdp.py:54 Retrieves default/home joint angles

These helpers enforce the 61-dimensional observation contract that keeps sim-to-real transfer stable. Always use _servo_joint_ids() rather than raw MuJoCo joint indices to avoid indexing into passive roller or fixed joints.

from mjlab_microduck.tasks.mdp import _servo_joint_pos, _servo_default_joint_pos

# Compute pose error for reward shaping

current_pos = _servo_joint_pos(env, env.asset)
default_pos = _servo_default_joint_pos(env, env.asset)
pose_error = torch.norm(current_pos - default_pos, dim=1)

Joint-Space Penalties and Regularizers

Smooth, efficient motion requires penalizing high-frequency control signals and mechanical stress. The custom MDP functions in Microduck RL provide granular regularization across legs, neck, and individual joint groups.

Action Rate and Acceleration Penalties

  • leg_action_rate_l2(...) (mdp.py:53): L2 penalty on Δaction for leg joints (discourages jitter)
  • neck_action_rate_l2(...) (mdp.py:93): L2 penalty on neck-joint action changes (head stability)
  • leg_action_acceleration_l2(...) (mdp.py:29): Second-order L2 penalty on leg action acceleration
  • neck_action_acceleration_l2(...) (mdp.py:68): Neck action acceleration penalty

Velocity and Torque Penalties

  • leg_joint_vel_l2(...) (mdp.py:34): Smoothness regularizer on leg joint velocities
  • neck_joint_vel_l2(...) (mdp.py:8): Head stability via neck velocity penalty
  • hip_pitch_knee_vel_l2(...) (mdp.py:42): Sagittal-plane gait regularizer
  • joint_torques_l2(...) (mdp.py:79): Energy efficiency penalty on actuator torques
  • joint_torque_rate_l2(...) (mdp.py:16): Gearbox protection via torque change penalty

Position and Limit Penalties

  • joint_deviation_l1(...) (mdp.py:72): L1 penalty on deviation from home pose
  • joint_pos_limit_proximity(...) (mdp.py:89): Soft penalty as joints approach hard limits
  • neck_joint_pos_l2(...) (mdp.py:55): Head pose regularization
  • bilateral_symmetry_penalty(...) (mdp.py:33): L1 penalty on left-right leg asymmetry

# Typical regularizer stack for velocity walking

from mjlab_microduck.tasks.mdp import (
    leg_action_rate_l2,
    leg_joint_vel_l2,
    joint_torques_l2,
    joint_deviation_l1,
)

reward_manager = _RewardManager(
    terms=[
        ("action_rate", leg_action_rate_l2, {"weight": -0.01}),
        ("joint_vel", leg_joint_vel_l2, {"weight": -0.001}),
        ("torque", joint_torques_l2, {"weight": -0.0001}),
        ("pose_dev", joint_deviation_l1, {"weight": -0.1}),
    ]
)

Body Orientation Rewards and Uprightness Shaping

Maintaining vertical posture is fundamental to bipedal and roller locomotion. These functions provide multiple mathematical formulations for upright rewards, allowing task-specific gradient properties.

Core Uprightness Functions

  • _fallen_mask(env, asset, gate_z, gate_tilt) (mdp.py:7): Binary mask (1=fallen) gating other rewards
  • body_upright_linear(...) (mdp.py:89): Linear cos(tilt) reward with gradient everywhere
  • body_upright_gaussian(...) (mdp.py:11): Gaussian reward with sharp peak at vertical
  • upright_gaussian_at_height(...) (mdp.py:35): Height-gated upright Gaussian (prevents crouch-cheating)

Composite Standing Rewards

  • standing_composite_score(...) (mdp.py:10): Multiplicative product of height, upright, and pose Gaussian scores
  • standing_success_bonus(...) (mdp.py:56): Sparse binary bonus when all tolerances satisfied

Potential-Based Shaping

  • upright_progress(...) (mdp.py:47): Potential-based shaping on trunk tilt cosine
  • height_progress(...) (mdp.py:75): Potential-based shaping on capped trunk height

Recovery and Penalty Terms

  • fallen_state_penalty(...) (mdp.py:5): Per-step tax while fallen
  • recovery_success(...) (mdp.py:46): Sparse bounty on regaining uprightness
  • body_ang_vel_at_height(...) (mdp.py:64): Angular velocity cost (height-gated)

# Composite standing reward with height gating

from mjlab_microduck.tasks.mdp import (
    standing_composite_score,
    standing_success_bonus,
    fallen_state_penalty,
)

reward_manager = _RewardManager(
    terms=[
        ("standing", standing_composite_score, {
            "height_target": 0.32,
            "upright_sigma": 0.2,
            "pose_sigma": 0.3,
        }),
        ("stand_bonus", standing_success_bonus, {"tolerance": 0.05}),
        ("fallen_tax", fallen_state_penalty, {"gate_tilt_above_deg": 45.0}),
    ]
)

Center-of-Mass Height and Velocity Rewards

Vertical control enables jumping, crouching, and terrain adaptation. The custom MDP functions in Microduck RL for CoM control support both fixed targets and phase-varying trajectories.

  • com_upward_velocity(...) (mdp.py:1): Positive reward for upward CoM velocity below target height
  • com_height_target(...) (mdp.py:33): Gaussian reward for CoM in desired height window
  • crouch_height_target(...) (mdp.py:78): Trapezoidal phase-based height target for crouch-glide tasks
  • crouch_glide_reward_from_values(...) (mdp.py:13): Gaussian reward matching crouch_height_target

# Crouch-glide task: periodic height modulation

from mjlab_microduck.tasks.mdp import (
    crouch_height_target,
    crouch_glide_reward_from_values,
)

# Height target modulates over episode progress

height_target = crouch_height_target(
    env,
    low_height=0.15,
    high_height=0.30,
    crouch_fraction=0.4,
    glide_fraction=0.4,
)

# Reward matches the target

reward_manager = _RewardManager(
    terms=[
        ("crouch_glide", crouch_glide_reward_from_values, {
            "height_target": height_target,
            "sigma": 0.05,
        }),
    ]
)

The Microduck platform includes passive roller wheels for efficient forward motion. These functions shape policies to exploit wheel mechanics while preventing pathological slip.

  • wheel_glide_reward(...) (mdp.py:94): Reward proportional to passive wheel spin (capped)
  • wheel_speed_reward(...) (mdp.py:74): Aligned wheel spin with commanded forward speed
  • descent_speed_reward(...) (mdp.py:36): Linear forward velocity reward for slope descent

# Roller configuration: exploit wheel glide with speed alignment

from mjlab_microduck.tasks.mdp import wheel_glide_reward, wheel_speed_reward

reward_manager = _RewardManager(
    terms=[
        ("glide", wheel_glide_reward, {"cap_speed": 0.35, "weight": 1.0}),
        ("speed_match", wheel_speed_reward, {
            "command_key": "base_velocity",
            "bidirectional": False,
        }),
    ]
)

Termination and Safety Events

Robust training requires early termination of failed rollouts and protection against numerical instability. These custom MDP functions integrate with Isaac Lab's EventManager.

Function Location Trigger Condition
fallen_too_long(...) mdp.py:37 Fallen duration exceeds threshold
robot_state_is_nan(...) mdp.py:63 Any state variable is NaN
root_height_below(...) mdp.py:17 Trunk falls below world-z threshold
body_impact_cost(...) mdp.py:45 Contact force exceeds protected-body threshold
contact_frequency_penalty(...) mdp.py:62 Excessive foot contact changes

Contact and Grounding Rewards

  • feet_grounded_reward(...) (mdp.py:25): Small positive reward per ground-contact foot
  • feet_air_time_upright(...) (mdp.py:27): Standard air-time reward (zeroed while fallen)
  • feet_flat_penalty(...) (mdp.py:64): Penalty for non-parallel foot-ground contact
  • feet_tiptoe_alignment(...) (mdp.py:15): Reward for foot x-axis pointing downward

# Safety-focused termination configuration

from mjlab_microduck.tasks.mdp import (
    fallen_too_long,
    robot_state_is_nan,
    root_height_below,
)

event_manager = _EventManager(
    terminations=[
        ("fallen_timeout", fallen_too_long, {
            "gate_z_below": 0.10,
            "max_duration_s": 4.0,
        }),
        ("nan_state", robot_state_is_nan, {}),
        ("fall_off_slope", root_height_below, {"min_height": -0.5}),
    ]
)

Phase-Based and Ground-Pick Rewards

Complex behaviors like sit-stand transitions and mouth-ground interaction require temporal task decomposition. These functions provide phase-conditioned reward shaping.

Mouth-Ground Interaction

  • mouth_ground_proximity(...) (mdp.py:38): Gaussian reward for mouth tip approaching ground (phase-weighted)
  • mouth_perpendicular_to_ground(...) (mdp.py:66): Rewards downward-pointing mouth x-axis during approach

Sit-Stand Transitions

  • sit_grounded(...) (mdp.py:92): Reward for trunk-ground contact while upright
  • sit_stability(...) (mdp.py:44): Low angular velocity reward during sit phase
  • phase_height_track(...) (mdp.py:30): Sinusoidal height target following

Pose Interpolation Rewards

  • pose_target_match(...) (mdp.py:59): Gaussian reward vs fixed target pose
  • interpolated_pose_target_match(...) (mdp.py:91): Interpolated between two poses
  • multistage_pose_target_match(...) (mdp.py:25): Arbitrary waypoint sequence (stand → fold → sit)
  • interpolated_height_target(...) (mdp.py:22): Height interpolation across waypoints

# Three-stage sit-stand-fold behavior

from mjlab_microduck.tasks.mdp import multistage_pose_target_match

pose_waypoints = [
    {"fraction": 0.0, "pose": "standing"},
    {"fraction": 0.4, "pose": "sitting"},
    {"fraction": 0.7, "pose": "folded"},
]

reward_manager = _RewardManager(
    terms=[
        ("multistage_pose", multistage_pose_target_match, {
            "waypoints": pose_waypoints,
            "sigma": 0.2,
        }),
    ]
)

Reset Helpers for Curriculum and Warm-Start

Training stability often requires non-uniform initialization. These functions hook into environment reset callbacks.

  • reset_with_forward_velocity(...) (mdp.py:58): Warm-start subset with random forward speed
  • reset_action_history(...) (mdp.py:29): Clear cached action-rate and acceleration buffers
  • reset_rolling_entry(...) (mdp.py:55): Initialize wheel velocities for non-slipping rollout
from mjlab_microduck.tasks.mdp import reset_with_forward_velocity

def curriculum_reset(env, env_ids):
    # Progressive velocity curriculum: 50% at speed after step 5M

    step = env.common_step_counter
    fraction = 0.0 if step < 5_000_000 else 0.5
    
    reset_with_forward_velocity(
        env,
        env_ids,
        velocity_range=(0.3, 0.8),
        fraction_stages=[{"step": 0, "fraction": fraction}],
    )

Summary

The custom MDP functions in Microduck RL provide a complete toolkit for robot learning:

  • Modular design: All functions live in src/mjlab_microduck/tasks/mdp.py and register via RewardManager/EventManager
  • Servo-safe indexing: Helper functions exclude passive joints to maintain the 61-dim observation contract
  • Multiple mathematical forms: Linear, Gaussian, and L1/L2 penalties for gradient tuning
  • Phase-aware shaping: Interpolation and waypoint support for complex behaviors
  • Hardware-aligned safety: NaN detection, impact costs, and fall termination

By combining these primitives in task configurations, you can train policies that transfer reliably from Isaac Lab simulation to the physical Microduck robot.


Frequently Asked Questions

How do I add a custom reward to my Microduck RL environment?

Create or edit a task configuration file (e.g., microduck_velocity_env_cfg.py), import your desired function from mjlab_microduck.tasks.mdp, and append a tuple to the RewardManager terms list: ("term_name", function, {params}).

What is the difference between body_upright_linear and body_upright_gaussian?

body_upright_linear returns cos(tilt) with non-zero gradient everywhere, providing continuous learning signal even when far from vertical. body_upright_gaussian uses exp(-tilt²/σ²) with sharp peak at vertical and near-zero gradient far from target—better for fine-tuning once approximately upright.

Why do MDP functions use _servo_joint_ids instead of direct MuJoCo indexing?

The Microduck MJCF includes passive roller joints that must not be controlled or observed directly. _servo_joint_ids filters to the 14 actuated servo joints, ensuring policies respect the hardware's actual degrees of freedom and maintain sim-to-real consistency.

How can I implement a curriculum that increases task difficulty over training?

Use reset_with_forward_velocity or time-varying reward weights in your configuration. Pass fraction_stages to progressively enable faster initial conditions, or modulate reward weight parameters via the RewardManager's runtime interface based on env.common_step_counter.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →