Episodic Trick Tasks in Microduck RL: Forward Roll, Ground Pick, and Ball Kick Explained
Episodic trick tasks in Microduck RL are short-duration behaviors with fixed-length episodes defined by duration_s, where the robot executes a maneuver and automatically terminates—distinct from perpetual standing or walking tasks.
Microduck RL organizes its learning environments into perpetual tasks (ongoing walking, standing) and episodic trick tasks (finite-duration maneuvers). These episodic tasks are marked by a constant-command flag in the manifest and a fixed duration_s field that governs episode length. The repository currently implements three production trick tasks plus one test-case demonstration.
What Defines an Episodic Trick Task
Episodic tasks share a common architectural pattern across the Microduck RL codebase. Each task enforces specific constraints that separate it from perpetual locomotion:
- Fixed episode duration – Controlled by
EPISODE_LENGTH_S(typically 4–5 seconds) and enforced viaTerminationTermCfgtimer expiration - Constant-command policy – The command slot remains frozen for the entire episode, ensuring the policy learns the maneuver without command-dependent adaptation
- Domain randomization – Full sim-to-real pipeline including mass, inertia, joint friction, and encoder bias randomization
- Manifest validation – The
src/mjlab_microduck/publish/manifest.pyvalidator rejects episodic policies missingduration_sor using non-constant command encodings
Forward Roll (Roulade)
The forward-roll or roulade task teaches the Microduck robot to roll forward over its head and land upright—a gymnastic maneuver requiring precise rotational control.
Mechanics and Reward Structure
The roll uses support-gated rotation: angular progress only accumulates while ground contact is detected. An overhead latch mechanism gates landing rewards to prevent premature credit assignment.
Key reward functions in src/mjlab_microduck/tasks/mdp.py:
| Function | Purpose |
|---|---|
microduck_mdp.roulade_progress |
Tracks rotation angle toward complete roll |
roulade_overspeed_penalty |
Penalizes excessive angular velocity |
roulade_head_pivot |
Rewards proper head positioning as rotation axis |
roulade_landing_composite |
Combined reward for successful touchdown |
roulade_upright_after_roll |
Final pose verification post-landing |
Configuration lives in src/mjlab_microduck/tasks/microduck_roulade_env_cfg.py, which sets EPISODE_LENGTH_S = 5.0 for the complete maneuver duration.
Training and Deployment
# Train the roulade task
uv run train microduck_roulade --env.scene.num-envs 64 --agent.max_iterations 1000
# Play trained policy (auto-terminates after ~5 seconds)
uv run play microduck_roulade --wandb-run-path your_org/your_project/run_id
# Export to ONNX with baked normalizer
uv run scripts/export.py microduck_roulade \
--wandb-run-path your_org/your_project/run_id \
--output-dir ./exported
# Publish to Hugging Face Hub
uv run publish \
--task microduck_roulade \
--wandb-run-path your_org/your_project/run_id \
--checkpoint 10000 \
--repo your_user/microduck-roulade \
--kind episodic \
--duration-s 5.0
Ground Pick
The ground-pick task trains precise foot placement: the robot lifts a foot, positions it on a ground target, then returns to standing.
Reward Components
Defined in src/mjlab_microduck/tasks/microduck_ground_pick_env_cfg.py:
microduck_mdp.ground_pick_progress– Distance-to-target progressground_pick_lift_reward– Foot clearance incentiveground_pick_place_reward– Accuracy of final foot touchdown
This task emphasizes fine motor control over explosive movement, with a shorter typical duration_s configured for the placement-and-recovery sequence.
Ball Kick
The ball-kick task implements a single-leg striking maneuver against a ball placed in front of the robot.
Kick-Specific Rewards
From src/mjlab_microduck/tasks/microduck_ball_kick_env_cfg.py and mdp.py:
microduck_mdp.kick_progress– Swing trajectory progresskick_impact_penalty– Penalizes premature or misaligned contactkick_body_ang_vel– Whole-body rotation stability
The ball-kick demonstrates how episodic tasks handle external object interaction—the ball physics are simulated but the policy focuses on leg trajectory rather than ball tracking.
Polite Bow (Test Case)
The polite-bow appears only in tests/test_publish_manifest.py as a manifest validation example. It defines a 4-second pose sequence with:
kind="episodic"
duration_s=4.0
This serves as a reference implementation for the minimum viable episodic task specification.
How Episodic Tasks Work Under the Hood
Manifest Validation
The src/mjlab_microduck/publish/manifest.py module enforces episodic task integrity:
# Pseudocode illustrating validation logic
if kind == "episodic":
assert duration_s > 0, "Episodic tasks require positive duration_s"
if command_encoding == "constant":
# Validate no command-dependent policy structure
pass
Episode Termination
All episodic tasks configure TerminationTermCfg with the episode timer condition, placed alongside other termination triggers (fall detection, joint limits) in the environment configuration.
Curriculum Integration
Randomization schedules in episodic tasks mirror perpetual tasks—no reduction in sim-to-real robustness despite the shorter episode horizon.
Summary
- Episodic trick tasks are finite-duration maneuvers with
duration_sand constant-command policies, distinct from perpetual walking/standing - Three production tasks: forward-roll (roulade), ground-pick, and ball-kick—each with specialized reward functions in
mdp.py - Reward design includes progress tracking, overspeed penalties, and composite landing rewards tailored to maneuver physics
- Validation infrastructure in
manifest.pyensures episodic policies declare duration and use appropriate command encodings - Deployment pipeline supports training, ONNX export, and Hugging Face Hub publication with
--kind episodicflag
Frequently Asked Questions
What makes a task "episodic" versus "perpetual" in Microduck RL?
An episodic task has a fixed duration_s field and automatically terminates when the timer expires, while perpetual tasks run indefinitely until an error condition (fall, joint limit) occurs. Episodic tasks also use constant-command policies that freeze the command input for the entire maneuver.
How do I add a new episodic trick task to the repository?
Create a new microduck_{task}_env_cfg.py file in src/mjlab_microduck/tasks/, define EPISODE_LENGTH_S, implement reward functions in mdp.py, and include TerminationTermCfg with a time-limit condition. Validate your manifest with --kind episodic and positive duration_s when publishing.
Why does the forward-roll use support-gated rotation?
Support-gating prevents the robot from receiving progress rewards while airborne or after falling, ensuring credit assignment only occurs during ground-contact rotation. The overhead latch similarly delays landing rewards until the robot clears the peak rotation point, preventing premature reward collection from partial rolls.
Can episodic policies run on the physical Microduck robot?
Yes—all episodic tasks use the same domain randomization pipeline as perpetual tasks, and the export pipeline bakes the normalizer into the ONNX model. The constant-command requirement simplifies deployment since no real-time command updates are needed during execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →