Custom MDP Functions in Microduck RL: A Complete Guide to Robot Learning Rewards and Terminations
The Microduck RL repository defines all task-specific reward, termination, and utility logic in src/mjlab_microduck/tasks/mdp.py, with over 60 pure PyTorch functions that can be registered via RewardManager or EventManager in environment configurations.
If you're building locomotion or manipulation policies for Pollen Robotics' Microduck platform, you'll spend most of your time configuring custom MDP functions in Microduck RL. These functions implement the dense reward shaping, safety penalties, and termination conditions that translate simulation-trained policies to hardware. This guide catalogues every major function category with exact source locations, signatures, and practical integration patterns.
Environment-Wide Helper Functions
Before computing rewards, most MDP terms need filtered access to the robot's state. The helper functions in this category handle servo-only joint indexing—critical because the Microduck MJCF excludes passive_* joints from control.
| Function | Location | Purpose |
|---|---|---|
_servo_joint_ids(env, asset) |
mdp.py:26 |
Returns indices of 14 controllable servo joints |
_servo_joint_pos(env, asset) |
mdp.py:46 |
Reads current joint positions (servo view) |
_servo_joint_vel(env, asset) |
mdp.py:50 |
Reads joint velocities (servo view) |
_servo_default_joint_pos(env, asset) |
mdp.py:54 |
Retrieves default/home joint angles |
These helpers enforce the 61-dimensional observation contract that keeps sim-to-real transfer stable. Always use _servo_joint_ids() rather than raw MuJoCo joint indices to avoid indexing into passive roller or fixed joints.
from mjlab_microduck.tasks.mdp import _servo_joint_pos, _servo_default_joint_pos
# Compute pose error for reward shaping
current_pos = _servo_joint_pos(env, env.asset)
default_pos = _servo_default_joint_pos(env, env.asset)
pose_error = torch.norm(current_pos - default_pos, dim=1)
Joint-Space Penalties and Regularizers
Smooth, efficient motion requires penalizing high-frequency control signals and mechanical stress. The custom MDP functions in Microduck RL provide granular regularization across legs, neck, and individual joint groups.
Action Rate and Acceleration Penalties
leg_action_rate_l2(...)(mdp.py:53): L2 penalty on Δaction for leg joints (discourages jitter)neck_action_rate_l2(...)(mdp.py:93): L2 penalty on neck-joint action changes (head stability)leg_action_acceleration_l2(...)(mdp.py:29): Second-order L2 penalty on leg action accelerationneck_action_acceleration_l2(...)(mdp.py:68): Neck action acceleration penalty
Velocity and Torque Penalties
leg_joint_vel_l2(...)(mdp.py:34): Smoothness regularizer on leg joint velocitiesneck_joint_vel_l2(...)(mdp.py:8): Head stability via neck velocity penaltyhip_pitch_knee_vel_l2(...)(mdp.py:42): Sagittal-plane gait regularizerjoint_torques_l2(...)(mdp.py:79): Energy efficiency penalty on actuator torquesjoint_torque_rate_l2(...)(mdp.py:16): Gearbox protection via torque change penalty
Position and Limit Penalties
joint_deviation_l1(...)(mdp.py:72): L1 penalty on deviation from home posejoint_pos_limit_proximity(...)(mdp.py:89): Soft penalty as joints approach hard limitsneck_joint_pos_l2(...)(mdp.py:55): Head pose regularizationbilateral_symmetry_penalty(...)(mdp.py:33): L1 penalty on left-right leg asymmetry
# Typical regularizer stack for velocity walking
from mjlab_microduck.tasks.mdp import (
leg_action_rate_l2,
leg_joint_vel_l2,
joint_torques_l2,
joint_deviation_l1,
)
reward_manager = _RewardManager(
terms=[
("action_rate", leg_action_rate_l2, {"weight": -0.01}),
("joint_vel", leg_joint_vel_l2, {"weight": -0.001}),
("torque", joint_torques_l2, {"weight": -0.0001}),
("pose_dev", joint_deviation_l1, {"weight": -0.1}),
]
)
Body Orientation Rewards and Uprightness Shaping
Maintaining vertical posture is fundamental to bipedal and roller locomotion. These functions provide multiple mathematical formulations for upright rewards, allowing task-specific gradient properties.
Core Uprightness Functions
_fallen_mask(env, asset, gate_z, gate_tilt)(mdp.py:7): Binary mask (1=fallen) gating other rewardsbody_upright_linear(...)(mdp.py:89): Linearcos(tilt)reward with gradient everywherebody_upright_gaussian(...)(mdp.py:11): Gaussian reward with sharp peak at verticalupright_gaussian_at_height(...)(mdp.py:35): Height-gated upright Gaussian (prevents crouch-cheating)
Composite Standing Rewards
standing_composite_score(...)(mdp.py:10): Multiplicative product of height, upright, and pose Gaussian scoresstanding_success_bonus(...)(mdp.py:56): Sparse binary bonus when all tolerances satisfied
Potential-Based Shaping
upright_progress(...)(mdp.py:47): Potential-based shaping on trunk tilt cosineheight_progress(...)(mdp.py:75): Potential-based shaping on capped trunk height
Recovery and Penalty Terms
fallen_state_penalty(...)(mdp.py:5): Per-step tax while fallenrecovery_success(...)(mdp.py:46): Sparse bounty on regaining uprightnessbody_ang_vel_at_height(...)(mdp.py:64): Angular velocity cost (height-gated)
# Composite standing reward with height gating
from mjlab_microduck.tasks.mdp import (
standing_composite_score,
standing_success_bonus,
fallen_state_penalty,
)
reward_manager = _RewardManager(
terms=[
("standing", standing_composite_score, {
"height_target": 0.32,
"upright_sigma": 0.2,
"pose_sigma": 0.3,
}),
("stand_bonus", standing_success_bonus, {"tolerance": 0.05}),
("fallen_tax", fallen_state_penalty, {"gate_tilt_above_deg": 45.0}),
]
)
Center-of-Mass Height and Velocity Rewards
Vertical control enables jumping, crouching, and terrain adaptation. The custom MDP functions in Microduck RL for CoM control support both fixed targets and phase-varying trajectories.
com_upward_velocity(...)(mdp.py:1): Positive reward for upward CoM velocity below target heightcom_height_target(...)(mdp.py:33): Gaussian reward for CoM in desired height windowcrouch_height_target(...)(mdp.py:78): Trapezoidal phase-based height target for crouch-glide taskscrouch_glide_reward_from_values(...)(mdp.py:13): Gaussian reward matchingcrouch_height_target
# Crouch-glide task: periodic height modulation
from mjlab_microduck.tasks.mdp import (
crouch_height_target,
crouch_glide_reward_from_values,
)
# Height target modulates over episode progress
height_target = crouch_height_target(
env,
low_height=0.15,
high_height=0.30,
crouch_fraction=0.4,
glide_fraction=0.4,
)
# Reward matches the target
reward_manager = _RewardManager(
terms=[
("crouch_glide", crouch_glide_reward_from_values, {
"height_target": height_target,
"sigma": 0.05,
}),
]
)
Wheel-Related Rewards for Roller Locomotion
The Microduck platform includes passive roller wheels for efficient forward motion. These functions shape policies to exploit wheel mechanics while preventing pathological slip.
wheel_glide_reward(...)(mdp.py:94): Reward proportional to passive wheel spin (capped)wheel_speed_reward(...)(mdp.py:74): Aligned wheel spin with commanded forward speeddescent_speed_reward(...)(mdp.py:36): Linear forward velocity reward for slope descent
# Roller configuration: exploit wheel glide with speed alignment
from mjlab_microduck.tasks.mdp import wheel_glide_reward, wheel_speed_reward
reward_manager = _RewardManager(
terms=[
("glide", wheel_glide_reward, {"cap_speed": 0.35, "weight": 1.0}),
("speed_match", wheel_speed_reward, {
"command_key": "base_velocity",
"bidirectional": False,
}),
]
)
Termination and Safety Events
Robust training requires early termination of failed rollouts and protection against numerical instability. These custom MDP functions integrate with Isaac Lab's EventManager.
| Function | Location | Trigger Condition |
|---|---|---|
fallen_too_long(...) |
mdp.py:37 |
Fallen duration exceeds threshold |
robot_state_is_nan(...) |
mdp.py:63 |
Any state variable is NaN |
root_height_below(...) |
mdp.py:17 |
Trunk falls below world-z threshold |
body_impact_cost(...) |
mdp.py:45 |
Contact force exceeds protected-body threshold |
contact_frequency_penalty(...) |
mdp.py:62 |
Excessive foot contact changes |
Contact and Grounding Rewards
feet_grounded_reward(...)(mdp.py:25): Small positive reward per ground-contact footfeet_air_time_upright(...)(mdp.py:27): Standard air-time reward (zeroed while fallen)feet_flat_penalty(...)(mdp.py:64): Penalty for non-parallel foot-ground contactfeet_tiptoe_alignment(...)(mdp.py:15): Reward for foot x-axis pointing downward
# Safety-focused termination configuration
from mjlab_microduck.tasks.mdp import (
fallen_too_long,
robot_state_is_nan,
root_height_below,
)
event_manager = _EventManager(
terminations=[
("fallen_timeout", fallen_too_long, {
"gate_z_below": 0.10,
"max_duration_s": 4.0,
}),
("nan_state", robot_state_is_nan, {}),
("fall_off_slope", root_height_below, {"min_height": -0.5}),
]
)
Phase-Based and Ground-Pick Rewards
Complex behaviors like sit-stand transitions and mouth-ground interaction require temporal task decomposition. These functions provide phase-conditioned reward shaping.
Mouth-Ground Interaction
mouth_ground_proximity(...)(mdp.py:38): Gaussian reward for mouth tip approaching ground (phase-weighted)mouth_perpendicular_to_ground(...)(mdp.py:66): Rewards downward-pointing mouth x-axis during approach
Sit-Stand Transitions
sit_grounded(...)(mdp.py:92): Reward for trunk-ground contact while uprightsit_stability(...)(mdp.py:44): Low angular velocity reward during sit phasephase_height_track(...)(mdp.py:30): Sinusoidal height target following
Pose Interpolation Rewards
pose_target_match(...)(mdp.py:59): Gaussian reward vs fixed target poseinterpolated_pose_target_match(...)(mdp.py:91): Interpolated between two posesmultistage_pose_target_match(...)(mdp.py:25): Arbitrary waypoint sequence (stand → fold → sit)interpolated_height_target(...)(mdp.py:22): Height interpolation across waypoints
# Three-stage sit-stand-fold behavior
from mjlab_microduck.tasks.mdp import multistage_pose_target_match
pose_waypoints = [
{"fraction": 0.0, "pose": "standing"},
{"fraction": 0.4, "pose": "sitting"},
{"fraction": 0.7, "pose": "folded"},
]
reward_manager = _RewardManager(
terms=[
("multistage_pose", multistage_pose_target_match, {
"waypoints": pose_waypoints,
"sigma": 0.2,
}),
]
)
Reset Helpers for Curriculum and Warm-Start
Training stability often requires non-uniform initialization. These functions hook into environment reset callbacks.
reset_with_forward_velocity(...)(mdp.py:58): Warm-start subset with random forward speedreset_action_history(...)(mdp.py:29): Clear cached action-rate and acceleration buffersreset_rolling_entry(...)(mdp.py:55): Initialize wheel velocities for non-slipping rollout
from mjlab_microduck.tasks.mdp import reset_with_forward_velocity
def curriculum_reset(env, env_ids):
# Progressive velocity curriculum: 50% at speed after step 5M
step = env.common_step_counter
fraction = 0.0 if step < 5_000_000 else 0.5
reset_with_forward_velocity(
env,
env_ids,
velocity_range=(0.3, 0.8),
fraction_stages=[{"step": 0, "fraction": fraction}],
)
Summary
The custom MDP functions in Microduck RL provide a complete toolkit for robot learning:
- Modular design: All functions live in
src/mjlab_microduck/tasks/mdp.pyand register viaRewardManager/EventManager - Servo-safe indexing: Helper functions exclude passive joints to maintain the 61-dim observation contract
- Multiple mathematical forms: Linear, Gaussian, and L1/L2 penalties for gradient tuning
- Phase-aware shaping: Interpolation and waypoint support for complex behaviors
- Hardware-aligned safety: NaN detection, impact costs, and fall termination
By combining these primitives in task configurations, you can train policies that transfer reliably from Isaac Lab simulation to the physical Microduck robot.
Frequently Asked Questions
How do I add a custom reward to my Microduck RL environment?
Create or edit a task configuration file (e.g., microduck_velocity_env_cfg.py), import your desired function from mjlab_microduck.tasks.mdp, and append a tuple to the RewardManager terms list: ("term_name", function, {params}).
What is the difference between body_upright_linear and body_upright_gaussian?
body_upright_linear returns cos(tilt) with non-zero gradient everywhere, providing continuous learning signal even when far from vertical. body_upright_gaussian uses exp(-tilt²/σ²) with sharp peak at vertical and near-zero gradient far from target—better for fine-tuning once approximately upright.
Why do MDP functions use _servo_joint_ids instead of direct MuJoCo indexing?
The Microduck MJCF includes passive roller joints that must not be controlled or observed directly. _servo_joint_ids filters to the 14 actuated servo joints, ensuring policies respect the hardware's actual degrees of freedom and maintain sim-to-real consistency.
How can I implement a curriculum that increases task difficulty over training?
Use reset_with_forward_velocity or time-varying reward weights in your configuration. Pass fraction_stages to progressively enable faster initial conditions, or modulate reward weight parameters via the RewardManager's runtime interface based on env.common_step_counter.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →