Episodic Trick Tasks in Microduck RL: Forward Roll, Ground Pick, and Ball Kick Explained

Episodic trick tasks in Microduck RL are short-duration behaviors with fixed-length episodes defined by duration_s, where the robot executes a maneuver and automatically terminates—distinct from perpetual standing or walking tasks.

Microduck RL organizes its learning environments into perpetual tasks (ongoing walking, standing) and episodic trick tasks (finite-duration maneuvers). These episodic tasks are marked by a constant-command flag in the manifest and a fixed duration_s field that governs episode length. The repository currently implements three production trick tasks plus one test-case demonstration.

What Defines an Episodic Trick Task

Episodic tasks share a common architectural pattern across the Microduck RL codebase. Each task enforces specific constraints that separate it from perpetual locomotion:

  • Fixed episode duration – Controlled by EPISODE_LENGTH_S (typically 4–5 seconds) and enforced via TerminationTermCfg timer expiration
  • Constant-command policy – The command slot remains frozen for the entire episode, ensuring the policy learns the maneuver without command-dependent adaptation
  • Domain randomization – Full sim-to-real pipeline including mass, inertia, joint friction, and encoder bias randomization
  • Manifest validation – The src/mjlab_microduck/publish/manifest.py validator rejects episodic policies missing duration_s or using non-constant command encodings

Forward Roll (Roulade)

The forward-roll or roulade task teaches the Microduck robot to roll forward over its head and land upright—a gymnastic maneuver requiring precise rotational control.

Mechanics and Reward Structure

The roll uses support-gated rotation: angular progress only accumulates while ground contact is detected. An overhead latch mechanism gates landing rewards to prevent premature credit assignment.

Key reward functions in src/mjlab_microduck/tasks/mdp.py:

Function Purpose
microduck_mdp.roulade_progress Tracks rotation angle toward complete roll
roulade_overspeed_penalty Penalizes excessive angular velocity
roulade_head_pivot Rewards proper head positioning as rotation axis
roulade_landing_composite Combined reward for successful touchdown
roulade_upright_after_roll Final pose verification post-landing

Configuration lives in src/mjlab_microduck/tasks/microduck_roulade_env_cfg.py, which sets EPISODE_LENGTH_S = 5.0 for the complete maneuver duration.

Training and Deployment


# Train the roulade task

uv run train microduck_roulade --env.scene.num-envs 64 --agent.max_iterations 1000

# Play trained policy (auto-terminates after ~5 seconds)

uv run play microduck_roulade --wandb-run-path your_org/your_project/run_id

# Export to ONNX with baked normalizer

uv run scripts/export.py microduck_roulade \
    --wandb-run-path your_org/your_project/run_id \
    --output-dir ./exported

# Publish to Hugging Face Hub

uv run publish \
    --task microduck_roulade \
    --wandb-run-path your_org/your_project/run_id \
    --checkpoint 10000 \
    --repo your_user/microduck-roulade \
    --kind episodic \
    --duration-s 5.0

Ground Pick

The ground-pick task trains precise foot placement: the robot lifts a foot, positions it on a ground target, then returns to standing.

Reward Components

Defined in src/mjlab_microduck/tasks/microduck_ground_pick_env_cfg.py:

  • microduck_mdp.ground_pick_progress – Distance-to-target progress
  • ground_pick_lift_reward – Foot clearance incentive
  • ground_pick_place_reward – Accuracy of final foot touchdown

This task emphasizes fine motor control over explosive movement, with a shorter typical duration_s configured for the placement-and-recovery sequence.

Ball Kick

The ball-kick task implements a single-leg striking maneuver against a ball placed in front of the robot.

Kick-Specific Rewards

From src/mjlab_microduck/tasks/microduck_ball_kick_env_cfg.py and mdp.py:

  • microduck_mdp.kick_progress – Swing trajectory progress
  • kick_impact_penalty – Penalizes premature or misaligned contact
  • kick_body_ang_vel – Whole-body rotation stability

The ball-kick demonstrates how episodic tasks handle external object interaction—the ball physics are simulated but the policy focuses on leg trajectory rather than ball tracking.

Polite Bow (Test Case)

The polite-bow appears only in tests/test_publish_manifest.py as a manifest validation example. It defines a 4-second pose sequence with:

kind="episodic"
duration_s=4.0

This serves as a reference implementation for the minimum viable episodic task specification.

How Episodic Tasks Work Under the Hood

Manifest Validation

The src/mjlab_microduck/publish/manifest.py module enforces episodic task integrity:


# Pseudocode illustrating validation logic

if kind == "episodic":
    assert duration_s > 0, "Episodic tasks require positive duration_s"
    if command_encoding == "constant":
        # Validate no command-dependent policy structure

        pass

Episode Termination

All episodic tasks configure TerminationTermCfg with the episode timer condition, placed alongside other termination triggers (fall detection, joint limits) in the environment configuration.

Curriculum Integration

Randomization schedules in episodic tasks mirror perpetual tasks—no reduction in sim-to-real robustness despite the shorter episode horizon.

Summary

  • Episodic trick tasks are finite-duration maneuvers with duration_s and constant-command policies, distinct from perpetual walking/standing
  • Three production tasks: forward-roll (roulade), ground-pick, and ball-kick—each with specialized reward functions in mdp.py
  • Reward design includes progress tracking, overspeed penalties, and composite landing rewards tailored to maneuver physics
  • Validation infrastructure in manifest.py ensures episodic policies declare duration and use appropriate command encodings
  • Deployment pipeline supports training, ONNX export, and Hugging Face Hub publication with --kind episodic flag

Frequently Asked Questions

What makes a task "episodic" versus "perpetual" in Microduck RL?

An episodic task has a fixed duration_s field and automatically terminates when the timer expires, while perpetual tasks run indefinitely until an error condition (fall, joint limit) occurs. Episodic tasks also use constant-command policies that freeze the command input for the entire maneuver.

How do I add a new episodic trick task to the repository?

Create a new microduck_{task}_env_cfg.py file in src/mjlab_microduck/tasks/, define EPISODE_LENGTH_S, implement reward functions in mdp.py, and include TerminationTermCfg with a time-limit condition. Validate your manifest with --kind episodic and positive duration_s when publishing.

Why does the forward-roll use support-gated rotation?

Support-gating prevents the robot from receiving progress rewards while airborne or after falling, ensuring credit assignment only occurs during ground-contact rotation. The overhead latch similarly delays landing rewards until the robot clears the peak rotation point, preventing premature reward collection from partial rolls.

Can episodic policies run on the physical Microduck robot?

Yes—all episodic tasks use the same domain randomization pipeline as perpetual tasks, and the export pipeline bakes the normalizer into the ONNX model. The constant-command requirement simplifies deployment since no real-time command updates are needed during execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →