# Episodic Trick Tasks in Microduck RL: Forward Roll, Ground Pick, and Ball Kick Explained

> Explore Microduck RL's episodic trick tasks like Forward Roll, Ground Pick, and Ball Kick. Learn about these fixed-duration maneuvers for robot learning.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: tutorial
- Published: 2026-09-08

---

**Episodic trick tasks in Microduck RL are short-duration behaviors with fixed-length episodes defined by `duration_s`, where the robot executes a maneuver and automatically terminates—distinct from perpetual standing or walking tasks.**

Microduck RL organizes its learning environments into **perpetual** tasks (ongoing walking, standing) and **episodic trick tasks** (finite-duration maneuvers). These episodic tasks are marked by a **constant-command** flag in the manifest and a fixed `duration_s` field that governs episode length. The repository currently implements three production trick tasks plus one test-case demonstration.

## What Defines an Episodic Trick Task

Episodic tasks share a common architectural pattern across the Microduck RL codebase. Each task enforces specific constraints that separate it from perpetual locomotion:

- **Fixed episode duration** – Controlled by `EPISODE_LENGTH_S` (typically 4–5 seconds) and enforced via `TerminationTermCfg` timer expiration
- **Constant-command policy** – The command slot remains frozen for the entire episode, ensuring the policy learns the maneuver without command-dependent adaptation
- **Domain randomization** – Full sim-to-real pipeline including mass, inertia, joint friction, and encoder bias randomization
- **Manifest validation** – The [`src/mjlab_microduck/publish/manifest.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/publish/manifest.py) validator rejects episodic policies missing `duration_s` or using non-constant command encodings

## Forward Roll (Roulade)

The **forward-roll** or **roulade** task teaches the Microduck robot to roll forward over its head and land upright—a gymnastic maneuver requiring precise rotational control.

### Mechanics and Reward Structure

The roll uses **support-gated rotation**: angular progress only accumulates while ground contact is detected. An overhead latch mechanism gates landing rewards to prevent premature credit assignment.

Key reward functions in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py):

| Function | Purpose |
|----------|---------|
| `microduck_mdp.roulade_progress` | Tracks rotation angle toward complete roll |
| `roulade_overspeed_penalty` | Penalizes excessive angular velocity |
| `roulade_head_pivot` | Rewards proper head positioning as rotation axis |
| `roulade_landing_composite` | Combined reward for successful touchdown |
| `roulade_upright_after_roll` | Final pose verification post-landing |

Configuration lives in [`src/mjlab_microduck/tasks/microduck_roulade_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_roulade_env_cfg.py), which sets `EPISODE_LENGTH_S = 5.0` for the complete maneuver duration.

### Training and Deployment

```bash

# Train the roulade task

uv run train microduck_roulade --env.scene.num-envs 64 --agent.max_iterations 1000

# Play trained policy (auto-terminates after ~5 seconds)

uv run play microduck_roulade --wandb-run-path your_org/your_project/run_id

# Export to ONNX with baked normalizer

uv run scripts/export.py microduck_roulade \
    --wandb-run-path your_org/your_project/run_id \
    --output-dir ./exported

# Publish to Hugging Face Hub

uv run publish \
    --task microduck_roulade \
    --wandb-run-path your_org/your_project/run_id \
    --checkpoint 10000 \
    --repo your_user/microduck-roulade \
    --kind episodic \
    --duration-s 5.0

```

## Ground Pick

The **ground-pick** task trains precise foot placement: the robot lifts a foot, positions it on a ground target, then returns to standing.

### Reward Components

Defined in [`src/mjlab_microduck/tasks/microduck_ground_pick_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_ground_pick_env_cfg.py):

- `microduck_mdp.ground_pick_progress` – Distance-to-target progress
- `ground_pick_lift_reward` – Foot clearance incentive
- `ground_pick_place_reward` – Accuracy of final foot touchdown

This task emphasizes **fine motor control** over explosive movement, with a shorter typical `duration_s` configured for the placement-and-recovery sequence.

## Ball Kick

The **ball-kick** task implements a single-leg striking maneuver against a ball placed in front of the robot.

### Kick-Specific Rewards

From [`src/mjlab_microduck/tasks/microduck_ball_kick_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_ball_kick_env_cfg.py) and [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py):

- `microduck_mdp.kick_progress` – Swing trajectory progress
- `kick_impact_penalty` – Penalizes premature or misaligned contact
- `kick_body_ang_vel` – Whole-body rotation stability

The ball-kick demonstrates how episodic tasks handle **external object interaction**—the ball physics are simulated but the policy focuses on leg trajectory rather than ball tracking.

## Polite Bow (Test Case)

The **polite-bow** appears only in [`tests/test_publish_manifest.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_publish_manifest.py) as a manifest validation example. It defines a 4-second pose sequence with:

```python
kind="episodic"
duration_s=4.0

```

This serves as a **reference implementation** for the minimum viable episodic task specification.

## How Episodic Tasks Work Under the Hood

### Manifest Validation

The [`src/mjlab_microduck/publish/manifest.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/publish/manifest.py) module enforces episodic task integrity:

```python

# Pseudocode illustrating validation logic

if kind == "episodic":
    assert duration_s > 0, "Episodic tasks require positive duration_s"
    if command_encoding == "constant":
        # Validate no command-dependent policy structure

        pass

```

### Episode Termination

All episodic tasks configure `TerminationTermCfg` with the episode timer condition, placed alongside other termination triggers (fall detection, joint limits) in the environment configuration.

### Curriculum Integration

Randomization schedules in episodic tasks mirror perpetual tasks—no reduction in sim-to-real robustness despite the shorter episode horizon.

## Summary

- **Episodic trick tasks** are finite-duration maneuvers with `duration_s` and constant-command policies, distinct from perpetual walking/standing
- **Three production tasks**: forward-roll (roulade), ground-pick, and ball-kick—each with specialized reward functions in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py)
- **Reward design** includes progress tracking, overspeed penalties, and composite landing rewards tailored to maneuver physics
- **Validation infrastructure** in [`manifest.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/manifest.py) ensures episodic policies declare duration and use appropriate command encodings
- **Deployment pipeline** supports training, ONNX export, and Hugging Face Hub publication with `--kind episodic` flag

## Frequently Asked Questions

### What makes a task "episodic" versus "perpetual" in Microduck RL?

An episodic task has a fixed `duration_s` field and automatically terminates when the timer expires, while perpetual tasks run indefinitely until an error condition (fall, joint limit) occurs. Episodic tasks also use **constant-command** policies that freeze the command input for the entire maneuver.

### How do I add a new episodic trick task to the repository?

Create a new `microduck_{task}_env_cfg.py` file in `src/mjlab_microduck/tasks/`, define `EPISODE_LENGTH_S`, implement reward functions in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py), and include `TerminationTermCfg` with a time-limit condition. Validate your manifest with `--kind episodic` and positive `duration_s` when publishing.

### Why does the forward-roll use support-gated rotation?

Support-gating prevents the robot from receiving progress rewards while airborne or after falling, ensuring credit assignment only occurs during ground-contact rotation. The overhead latch similarly delays landing rewards until the robot clears the peak rotation point, preventing premature reward collection from partial rolls.

### Can episodic policies run on the physical Microduck robot?

Yes—all episodic tasks use the same domain randomization pipeline as perpetual tasks, and the export pipeline bakes the normalizer into the ONNX model. The constant-command requirement simplifies deployment since no real-time command updates are needed during execution.