Reward Function for the StandUp Recovery Task in Microduck RL: Architecture and Implementation

The StandUp Recovery reward function in Microduck RL combines Gaussian-shaped height and upright orientation bonuses with upward velocity incentives and lightly weighted penalties, implemented as a weighted sum of terms defined in mdp.py and configured in microduck_standup_env_cfg.py to enable robust standing from a fallen state.

The StandUp Recovery task in the Pollen Robotics Microduck RL repository trains a small quadruped robot to rise from a prone position and stabilize upright. The reward function that drives this behavior is intentionally minimal yet expressive, utilizing a carefully balanced set of terms defined in the core MDP module and assembled within the task configuration file.

Core Reward Components

The reward pipeline evaluates seven distinct terms every simulation step. Each term is implemented in src/mjlab_microduck/tasks/mdp.py and registered with specific weights inside src/mjlab_microduck/tasks/microduck_standup_env_cfg.py.

Height-Target Gaussian

The primary objective is encoded in height_target_gaussian(target_height, asset_cfg, std). This term creates a Gaussian-shaped reward centered on the target trunk height (approximately 0.115 m), providing a smooth gradient that encourages the center of mass to reach the standing posture. In the configuration, this term carries a positive weight of +5.0.

Upright Gaussian at Height

Once the target elevation is reached, the robot must remain vertical. upright_gaussian_at_height(std, height_low, height_high, asset_cfg) combines a height gate with a Gaussian penalty on torso tilt, rewarding the body for staying upright within the valid height range. This term contributes +2.0 to the total reward.

CoM Upward Velocity

To encourage active climbing rather than passive drifting, com_upward_velocity(asset_cfg, max_height) provides a linear bonus proportional to the vertical center-of-mass speed, capped at a maximum height threshold. This term is weighted +3.0.

Pose L1 Penalty

The pose_l1_penalty(target_overrides, asset_cfg, joint_indices) term applies an L1 loss on joint positions to keep configurations close to the target upright pose. As a penalty, it uses a negative weight of -6.0.

Regularization Terms

Three additional components ensure smooth control without blocking the explosive motions required for standing:

  • Joint Torque Rate L2: joint_torque_rate_l2() penalizes rapid torque changes with a weight of -0.002, damping oscillations while permitting the large torque variations necessary for lift-off.
  • Body Angular Velocity: A direct penalty on rotational velocity with weight -0.05 prevents uncontrollable spinning without inhibiting the flip-over motion required to transition from prone to standing.
  • Action Rate L2: action_rate_l2() smooths action deltas. The curriculum progressively ramps its weight from -0.4 to -0.8 and finally -1.0 during training phases.

Reward Assembly in the Configuration

Inside src/mjlab_microduck/tasks/microduck_standup_env_cfg.py, the make_microduck_standup_env_cfg factory combines these terms into the final scalar reward:


# src/mjlab_microduck/tasks/microduck_standup_env_cfg.py

cfg.rewards = [
    # Positive shaping terms

    RewardTerm(name="height_target_gaussian", weight=+5.0, 
               fn=height_target_gaussian(...)),
    RewardTerm(name="upright_gaussian",       weight=+2.0, 
               fn=upright_gaussian_at_height(...)),
    RewardTerm(name="com_upward_velocity",    weight=+3.0, 
               fn=com_upward_velocity(...)),
    
    # Penalties and regularization

    RewardTerm(name="pose_l1",                weight=-6.0, 
               fn=pose_l1_penalty(...)),
    RewardTerm(name="joint_torque_rate",      weight=-0.002, 
               fn=joint_torque_rate_l2()),
    RewardTerm(name="body_ang_vel",          weight=-0.05,  
               fn=body_ang_vel_penalty()),
    RewardTerm(name="action_rate",           weight=-1.0,   
               fn=action_rate_l2()),
]

The environment computes the step reward as a weighted sum across all active terms.

Design Philosophy and Curriculum Integration

According to the source code in microduck_standup_env_cfg.py, penalty weights are deliberately lightweight—particularly the body angular velocity term at -0.05—because aggressive penalization would freeze the robot during the necessary flip-over phase. This design choice reflects empirical findings documented in the configuration comments around line 463.

The Gaussian-shaped positive terms provide dense, smooth gradients even when the robot is far from the target height, preventing local optima where the agent might remain on the ground. Regularization terms survive the stand-up curriculum specifically because they do not block the large, rapid joint motions required to lift the body from a prone position.

Accessing and Computing Rewards Programmatically

You can inspect the reward structure at runtime using the configuration factory:

from mjlab_microduck.tasks.microduck_standup_env_cfg import make_microduck_standup_env_cfg

standup_cfg = make_microduck_standup_env_cfg()
for term in standup_cfg.rewards:
    print(f"{term.name:30s} weight={term.weight:6.2f}")

During rollout, the environment evaluates each term's function against the current observation and action. While the internal step logic handles this automatically, you can replicate the computation manually:


# Pseudo-code for manual reward calculation

obs = env.reset()
action = policy(obs)
next_obs, _, _, info = env.step(action)

total_reward = 0.0
for term in standup_cfg.rewards:
    term_value = term.fn(obs, action, next_obs)
    total_reward += term.weight * term_value
    
print(f"Step reward: {total_reward}")

Summary

  • The reward function combines three positive terms (height Gaussian, upright Gaussian, upward velocity) with four penalty terms (pose L1, torque rate, angular velocity, action rate).
  • All mathematical implementations reside in src/mjlab_microduck/tasks/mdp.py, while task-specific weights and curriculum scheduling are defined in src/mjlab_microduck/tasks/microduck_standup_env_cfg.py.
  • Penalty weights are intentionally lightweight (-0.05 for angular velocity, -0.002 for torque rate) to avoid blocking the flip-over recovery motion.
  • The curriculum progressively strengthens the action rate penalty from -0.4 to -1.0 during training, encouraging smooth control without constraining early exploration.

Frequently Asked Questions

Why is the body angular velocity penalty so small?

A weight of only -0.05 prevents the robot from spinning uncontrollably while still permitting the rapid rotational movements required to flip from a prone position to standing. According to comments in the configuration file at line 463, larger values would freeze the agent during the necessary recovery motion.

Where are the reward term implementations located?

The mathematical logic for each term is defined in src/mjlab_microduck/tasks/mdp.py, including functions like height_target_gaussian(), com_upward_velocity(), and joint_torque_rate_l2(). The task-specific weights, term selection, and curriculum parameters are configured in src/mjlab_microduck/tasks/microduck_standup_env_cfg.py.

How does the curriculum affect the reward function during training?

The curriculum primarily modulates the action_rate_l2 penalty weight, ramping it from -0.4 to -0.8 and finally -1.0 across training phases. This gradual tightening encourages smooth control policies without initially constraining the exploratory movements needed to discover the standing behavior.

Can I modify the target height for the stand-up task?

Yes. The height_target_gaussian term accepts a target_height parameter specified during instantiation in microduck_standup_env_cfg.py. Adjusting this value (default approximately 0.115 m) changes the vertical center-of-mass target that the policy attempts to reach.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →