VelStand Task in Microduck RL: Combining Walking and Recovery Policies
The VelStand task is a unified reinforcement learning environment that merges the Microduck robot's velocity-based walking policy with gated recovery rewards, enabling a single neural network to both locomote and autonomously stand up after falls.
The VelStand task in the pollen-robotics/microduck_rl repository solves a critical challenge in bipedal robotics: training one policy that handles both stable locomotion and emergency recovery. Unlike separate walking and stand-up controllers, this environment reuses the proven velocity recipe while injecting conditional rewards that activate only when the robot has fallen.
Architecture of the VelStand Task
The VelStand environment operates as a single-policy system that seamlessly transitions between walking and recovery behaviors. According to the source code in src/mjlab_microduck/tasks/microduck_velstand_env_cfg.py, the architecture combines three core components: the walking layer, a collision-enabled robot model, and gated recovery rewards.
The Walking Layer (Velocity Recipe)
At its foundation, VelStand reuses the entire velocity environment defined in make_microduck_velocity_env_cfg. This inheritance preserves all tracking weights, air-time rewards, turn-in-place buckets, domain randomization, and observation noise from the standard walking task. The configuration maintains the identical 61-dimensional observation space used across other Microduck policies, ensuring compatibility with existing training pipelines and hardware deployment.
The Recovery Reward Layer
The innovation lies in gated recovery rewards that fire exclusively when the robot detects a fallen state—defined as a trunk tilt exceeding 40° or a trunk Z-height below 0.08 meters. These rewards, implemented starting at line 98 of the configuration file, include:
- upright_progress: Potential-based reward for returning to vertical orientation
- height_progress: Incentive for raising the center of mass
- com_upward_velocity: Reward for upward momentum during recovery
- fallen_tax: Penalty for remaining in a fallen state
- recovery_success: Bounty awarded upon achieving upright standing posture
During stable walking, these rewards remain zero, ensuring the policy focuses on velocity tracking until a fall actually occurs.
Impact Penalties and Safety
To discourage violent recovery maneuvers, the environment includes ungated impact penalties that penalize hard trunk and head landings regardless of state. The joint_torque_rate_l2 reward (lines 39-42) and collision penalties from the docstring (lines 23-24) encourage smooth, controlled movements during both walking and recovery phases.
Robot Configuration for Recovery
Unlike standard walking environments that use simplified collision models, VelStand utilizes the all-collision stand-up MJCF specified in MICRODUCK_STANDUP_ROBOT_CFG (src/mjlab_microduck/robot/microduck_constants.py, lines 71-73). This robot configuration enables the simulated Microduck to lie flat on the ground and generate sufficient friction to push off during recovery attempts.
Training Curriculum and Phases
The VelStand task implements a three-phase curriculum defined in lines 62-70 and the PRONE_RAMP_STAGES list (lines 53-59) to balance walking and recovery data:
- Phase 1: Clean walking with
fell_overtermination enabled to establish baseline locomotion - Phase 2: Fall-over disabled, converting falls into recovery training opportunities
- Phase 3: Prone-init ramp that gradually introduces prone and crouch start positions
This structured progression guarantees the policy receives adequate walking demonstrations while accumulating dense recovery experience.
Termination Logic
The environment employs specific termination conditions to prevent infinite loops. The fallen_too_long termination triggers after 8 seconds in a fallen state (line 305), ending episodes where the robot fails to recover. During play mode, the fell_over termination is removed to allow continuous evaluation of recovery capabilities.
Implementation Example
To instantiate the VelStand environment for training or inference:
from mjlab_microduck.tasks.microduck_velstand_env_cfg import (
make_microduck_velstand_env_cfg,
)
# Training configuration with curriculum active
train_cfg = make_microduck_velstand_env_cfg(play=False, rough=False)
# Play configuration with curriculum disabled
play_cfg = make_microduck_velstand_env_cfg(play=True, rough=False)
Launch a quick smoke test using the CLI:
uv run train Mjlab-VelStand-Flat-MicroDuck \
--env.scene.num-envs 64 \
--agent.max_iterations 5
Export the trained policy to ONNX format for real robot deployment:
uv run scripts/export.py Mjlab-VelStand-Flat-MicroDuck \
--wandb-run-path <entity/project/run_id>
Summary
- The VelStand task unifies walking and recovery into a single policy by extending the velocity environment with conditional rewards
- Recovery rewards activate only when tilt exceeds 40° or height drops below 0.08m, remaining dormant during normal walking
- The MICRODUCK_STANDUP_ROBOT_CFG collision model allows the robot to interact with the ground plane for push-off maneuvers
- A three-phase curriculum ensures balanced training data between locomotion and recovery scenarios
- The 61-dimensional observation space and BAM actuator model maintain compatibility with existing deployment pipelines
Frequently Asked Questions
What makes VelStand different from standard walking tasks?
Standard walking tasks terminate episodes immediately upon falling, while VelStand treats falls as learning opportunities. By disabling terminations in later curriculum phases and adding gated recovery rewards, the environment trains policies that can autonomously transition from prone positions back to walking without human intervention or controller switching.
How does the robot detect when to trigger recovery rewards?
The environment monitors trunk orientation and height through the upright_progress and height_progress functions defined in src/mjlab_microduck/tasks/mdp.py. When the trunk tilt exceeds 40 degrees or the Z-position drops below 0.08 meters, the system activates recovery rewards including com_upward_velocity and recovery_success while applying the fallen_tax penalty until upright posture is restored.
Can the trained VelStand policy be deployed to real hardware?
Yes. Because VelStand maintains the same 61-D observation contract and BAM actuator model as other Microduck environments, trained policies export directly to ONNX format without additional wiring. The observation normalization statistics bake into the exported model, allowing seamless transfer from simulation to the physical Microduck robot.
What files define the VelStand environment configuration?
The primary configuration resides in src/mjlab_microduck/tasks/microduck_velstand_env_cfg.py, which registers the task IDs Mjlab-VelStand-Flat-MicroDuck and Mjlab-VelStand-Rough-MicroDuck in src/mjlab_microduck/tasks/__init__.py. Reward function implementations are located in src/mjlab_microduck/tasks/mdp.py, while the stand-up robot model is defined in src/mjlab_microduck/robot/microduck_constants.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →