How to Define Reward Functions for Microduck RL Tasks

Microduck RL reward functions are Python callables that receive a ManagerBasedRlEnv instance and return a 1-D torch.Tensor of per-environment reward values, registered via RewardTermCfg objects in task configuration files.

The pollen-robotics/microduck_rl repository provides a modular reward system for training bipedal locomotion policies. All reward implementations follow a consistent signature pattern, making it straightforward to compose existing terms or author custom shaping objectives.

Where Reward Functions Are Implemented

Every reward function lives in src/mjlab_microduck/tasks/mdp.py. These are ordinary Python functions that the RewardManager calls at each simulation step.

Key built-in rewards include:

  • Linear upright reward — body_upright_linear provides continuous gradient feedback across tilt angles (lines 68-78)
  • Gaussian upright reward — body_upright_gaussian creates a sharp attractor toward vertical orientation (lines 110-126)
  • Composite standing score — standing_composite_score multiplies height, uprightness, and pose Gaussians for stable standing (lines 208-226)
  • Recovery terms — penalties and bonuses for fall-recovery curricula, such as fallen_state_penalty (lines 106-144)

Each function shares the same signature structure:

def body_upright_linear(
    env: ManagerBasedRlEnv,
    asset_cfg: SceneEntityCfg = _DEFAULT_ASSET_CFG,
    gate_z_below: float | None = None,
    gate_tilt_above_deg: float = 40.0,
) -> torch.Tensor:
    ...

The env parameter provides access to the full simulation state. The asset_cfg specifies which robot asset to query. Additional keyword arguments allow per-task customization without code modification.

Registering Reward Functions in Environment Configurations

Rewards connect to environments through RewardTermCfg objects defined in task-specific configuration modules. The microduck_velocity_env_cfg.py file demonstrates this pattern around line 200.

from mjlab.managers import RewardTermCfg
from mjlab_microduck.tasks import mdp as microduck_mdp

reward_terms = [
    RewardTermCfg(
        name="upright_linear",
        weight=0.5,
        fn=microduck_mdp.body_upright_linear,
        args=dict(gate_z_below=None, gate_tilt_above_deg=40.0),
    ),
    RewardTermCfg(
        name="upright_gaussian",
        weight=0.2,
        fn=microduck_mdp.body_upright_gaussian,
        args=dict(std=0.1),
    ),
]

The RewardManager executes each registered function every step, multiplies the returned tensor by weight, and accumulates results into the episode reward. Positive weights encode bonuses; negative weights encode costs. The args dict passes custom parameters to the underlying function.

Creating Custom Reward Functions

Follow this four-step workflow to add new reward terms to Microduck RL tasks:

  1. Implement in mdp.py — Write a function with signature fn(env, ...) -> torch.Tensor
  2. Import in configuration — Expose via from mjlab_microduck.tasks import mdp as microduck_mdp
  3. Add RewardTermCfg — Specify unique name, scalar weight, and any additional arguments
  4. Validate with smoke test — Run uv run train <TASK_ID> --env.scene.num-envs 64 --agent.max_iterations 5 to catch integration errors

Complete Example: Joint Velocity Penalty

This custom reward penalizes high joint velocities for smoother motion policies.

Step 1: Define in src/mjlab_microduck/tasks/mdp.py

def joint_vel_penalty(
    env: ManagerBasedRlEnv,
    asset_cfg: SceneEntityCfg = _DEFAULT_ASSET_CFG,
    scale: float = 0.01,
) -> torch.Tensor:
    """Penalize high joint velocities (smoothness regularizer)."""
    asset = env.scene[asset_cfg.name]
    vel = _servo_joint_vel(env, asset)          # shape (N, 14)

    return -scale * torch.sum(vel ** 2, dim=1)

The implementation follows the same structure as leg_action_rate_l2 (lines 54-90), using internal helpers to extract joint states and returning a per-environment penalty vector.

Step 2: Register in task configuration

RewardTermCfg(
    name="joint_vel_penalty",
    weight=-0.01,                # negative weight applies as cost

    fn=microduck_mdp.joint_vel_penalty,
    args=dict(scale=0.01),
),

Core Source Files for Reward Development

File Purpose Location in Repository
src/mjlab_microduck/tasks/mdp.py Central reward and termination function library src/mjlab_microduck/tasks/mdp.py
src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py Example task configuration with reward term registration src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py
src/mjlab_microduck/tasks/symmetry.py Symmetry-aware training wrapper (uses standard reward interface) src/mjlab_microduck/tasks/symmetry.py

Summary

  • Location: All reward functions reside in src/mjlab_microduck/tasks/mdp.py as callables taking ManagerBasedRlEnv and returning torch.Tensor
  • Registration: Use RewardTermCfg objects with name, weight, fn, and args to wire rewards into task configs
  • Signatures: Follow the pattern def my_reward(env: ManagerBasedRlEnv, asset_cfg: SceneEntityCfg = ..., **kwargs) -> torch.Tensor
  • Weights: Positive for bonuses, negative for penalties; applied automatically by RewardManager
  • Validation: Always run a short training smoke test after adding custom terms

Frequently Asked Questions

What data does a reward function receive from the environment?

The env: ManagerBasedRlEnv parameter provides full access to simulation state through env.scene, including robot assets, joint positions/velocities, contact forces, and command targets. Query specific entities via env.scene[asset_cfg.name] where asset_cfg is a SceneEntityCfg object.

Can I use different reward weights for different training phases?

Yes. The configuration system supports curriculum updates through the RewardManager. Modify term weights programmatically during training by accessing the environment's reward manager instance, or define multiple environment variants with different static configurations.

How do I debug a reward function that returns NaN values?

Check for division by zero in normalization steps, ensure all tensor operations preserve batch dimensions, and verify that env.scene contains valid data for the requested asset. The smoke test (--agent.max_iterations 5) catches NaN propagation early. Add torch.nan_to_num() guards or explicit clamping where numerical instability occurs.

Are sparse rewards supported in Microduck RL?

Absolutely. Return zero tensors for most steps and non-zero values only at goal achievement. However, the default PPO configuration performs best with dense shaping rewards. For pure sparse formulations, consider increasing the number of parallel environments or adjusting the PPO entropy coefficient to maintain exploration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →