# How to Define Reward Functions for Microduck RL Tasks

> Learn how to define Microduck RL reward functions using Python callables and RewardTermCfg objects in task configurations. Master your reinforcement learning tasks.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: how-to-guide
- Published: 2026-09-02

---

**Microduck RL reward functions are Python callables that receive a `ManagerBasedRlEnv` instance and return a 1-D `torch.Tensor` of per-environment reward values, registered via `RewardTermCfg` objects in task configuration files.**

The pollen-robotics/microduck_rl repository provides a modular reward system for training bipedal locomotion policies. All reward implementations follow a consistent signature pattern, making it straightforward to compose existing terms or author custom shaping objectives.

## Where Reward Functions Are Implemented

Every reward function lives in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py). These are ordinary Python functions that the **`RewardManager`** calls at each simulation step.

Key built-in rewards include:

- **Linear upright reward** — `body_upright_linear` provides continuous gradient feedback across tilt angles (lines 68-78)
- **Gaussian upright reward** — `body_upright_gaussian` creates a sharp attractor toward vertical orientation (lines 110-126)
- **Composite standing score** — `standing_composite_score` multiplies height, uprightness, and pose Gaussians for stable standing (lines 208-226)
- **Recovery terms** — penalties and bonuses for fall-recovery curricula, such as `fallen_state_penalty` (lines 106-144)

Each function shares the same signature structure:

```python
def body_upright_linear(
    env: ManagerBasedRlEnv,
    asset_cfg: SceneEntityCfg = _DEFAULT_ASSET_CFG,
    gate_z_below: float | None = None,
    gate_tilt_above_deg: float = 40.0,
) -> torch.Tensor:
    ...

```

The `env` parameter provides access to the full simulation state. The `asset_cfg` specifies which robot asset to query. Additional keyword arguments allow per-task customization without code modification.

## Registering Reward Functions in Environment Configurations

Rewards connect to environments through **`RewardTermCfg`** objects defined in task-specific configuration modules. The [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py) file demonstrates this pattern around line 200.

```python
from mjlab.managers import RewardTermCfg
from mjlab_microduck.tasks import mdp as microduck_mdp

reward_terms = [
    RewardTermCfg(
        name="upright_linear",
        weight=0.5,
        fn=microduck_mdp.body_upright_linear,
        args=dict(gate_z_below=None, gate_tilt_above_deg=40.0),
    ),
    RewardTermCfg(
        name="upright_gaussian",
        weight=0.2,
        fn=microduck_mdp.body_upright_gaussian,
        args=dict(std=0.1),
    ),
]

```

The **`RewardManager`** executes each registered function every step, multiplies the returned tensor by `weight`, and accumulates results into the episode reward. Positive weights encode bonuses; negative weights encode costs. The `args` dict passes custom parameters to the underlying function.

## Creating Custom Reward Functions

Follow this four-step workflow to add new reward terms to Microduck RL tasks:

1. **Implement in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py)** — Write a function with signature `fn(env, ...) -> torch.Tensor`
2. **Import in configuration** — Expose via `from mjlab_microduck.tasks import mdp as microduck_mdp`
3. **Add `RewardTermCfg`** — Specify unique `name`, scalar `weight`, and any additional arguments
4. **Validate with smoke test** — Run `uv run train <TASK_ID> --env.scene.num-envs 64 --agent.max_iterations 5` to catch integration errors

### Complete Example: Joint Velocity Penalty

This custom reward penalizes high joint velocities for smoother motion policies.

**Step 1: Define in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py)**

```python
def joint_vel_penalty(
    env: ManagerBasedRlEnv,
    asset_cfg: SceneEntityCfg = _DEFAULT_ASSET_CFG,
    scale: float = 0.01,
) -> torch.Tensor:
    """Penalize high joint velocities (smoothness regularizer)."""
    asset = env.scene[asset_cfg.name]
    vel = _servo_joint_vel(env, asset)          # shape (N, 14)

    return -scale * torch.sum(vel ** 2, dim=1)

```

The implementation follows the same structure as `leg_action_rate_l2` (lines 54-90), using internal helpers to extract joint states and returning a per-environment penalty vector.

**Step 2: Register in task configuration**

```python
RewardTermCfg(
    name="joint_vel_penalty",
    weight=-0.01,                # negative weight applies as cost

    fn=microduck_mdp.joint_vel_penalty,
    args=dict(scale=0.01),
),

```

## Core Source Files for Reward Development

| File | Purpose | Location in Repository |
|------|---------|------------------------|
| [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) | Central reward and termination function library | [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) |
| [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) | Example task configuration with reward term registration | [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) |
| [`src/mjlab_microduck/tasks/symmetry.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/symmetry.py) | Symmetry-aware training wrapper (uses standard reward interface) | [`src/mjlab_microduck/tasks/symmetry.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/symmetry.py) |

## Summary

- **Location**: All reward functions reside in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) as callables taking `ManagerBasedRlEnv` and returning `torch.Tensor`
- **Registration**: Use `RewardTermCfg` objects with `name`, `weight`, `fn`, and `args` to wire rewards into task configs
- **Signatures**: Follow the pattern `def my_reward(env: ManagerBasedRlEnv, asset_cfg: SceneEntityCfg = ..., **kwargs) -> torch.Tensor`
- **Weights**: Positive for bonuses, negative for penalties; applied automatically by `RewardManager`
- **Validation**: Always run a short training smoke test after adding custom terms

## Frequently Asked Questions

### What data does a reward function receive from the environment?

The `env: ManagerBasedRlEnv` parameter provides full access to simulation state through `env.scene`, including robot assets, joint positions/velocities, contact forces, and command targets. Query specific entities via `env.scene[asset_cfg.name]` where `asset_cfg` is a `SceneEntityCfg` object.

### Can I use different reward weights for different training phases?

Yes. The configuration system supports curriculum updates through the `RewardManager`. Modify term weights programmatically during training by accessing the environment's reward manager instance, or define multiple environment variants with different static configurations.

### How do I debug a reward function that returns NaN values?

Check for division by zero in normalization steps, ensure all tensor operations preserve batch dimensions, and verify that `env.scene` contains valid data for the requested asset. The smoke test (`--agent.max_iterations 5`) catches NaN propagation early. Add `torch.nan_to_num()` guards or explicit clamping where numerical instability occurs.

### Are sparse rewards supported in Microduck RL?

Absolutely. Return zero tensors for most steps and non-zero values only at goal achievement. However, the default PPO configuration performs best with dense shaping rewards. For pure sparse formulations, consider increasing the number of parallel environments or adjusting the PPO entropy coefficient to maintain exploration.