# VelStand Task in Microduck RL: Combining Walking and Recovery Policies

> Discover the VelStand task in Microduck RL. Learn how this unified RL environment combines walking and recovery policies for autonomous locomotion and standing after falls.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-01

---

**The VelStand task is a unified reinforcement learning environment that merges the Microduck robot's velocity-based walking policy with gated recovery rewards, enabling a single neural network to both locomote and autonomously stand up after falls.**

The **VelStand task** in the `pollen-robotics/microduck_rl` repository solves a critical challenge in bipedal robotics: training one policy that handles both stable locomotion and emergency recovery. Unlike separate walking and stand-up controllers, this environment reuses the proven velocity recipe while injecting conditional rewards that activate only when the robot has fallen.

## Architecture of the VelStand Task

The VelStand environment operates as a **single-policy** system that seamlessly transitions between walking and recovery behaviors. According to the source code in [`src/mjlab_microduck/tasks/microduck_velstand_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velstand_env_cfg.py), the architecture combines three core components: the walking layer, a collision-enabled robot model, and gated recovery rewards.

### The Walking Layer (Velocity Recipe)

At its foundation, VelStand reuses the entire **velocity environment** defined in `make_microduck_velocity_env_cfg`. This inheritance preserves all tracking weights, air-time rewards, turn-in-place buckets, domain randomization, and observation noise from the standard walking task. The configuration maintains the identical **61-dimensional observation space** used across other Microduck policies, ensuring compatibility with existing training pipelines and hardware deployment.

### The Recovery Reward Layer

The innovation lies in **gated recovery rewards** that fire exclusively when the robot detects a fallen state—defined as a trunk tilt exceeding 40° or a trunk Z-height below 0.08 meters. These rewards, implemented starting at line 98 of the configuration file, include:

- **upright_progress**: Potential-based reward for returning to vertical orientation
- **height_progress**: Incentive for raising the center of mass
- **com_upward_velocity**: Reward for upward momentum during recovery
- **fallen_tax**: Penalty for remaining in a fallen state
- **recovery_success**: Bounty awarded upon achieving upright standing posture

During stable walking, these rewards remain zero, ensuring the policy focuses on velocity tracking until a fall actually occurs.

### Impact Penalties and Safety

To discourage violent recovery maneuvers, the environment includes **ungated impact penalties** that penalize hard trunk and head landings regardless of state. The `joint_torque_rate_l2` reward (lines 39-42) and collision penalties from the docstring (lines 23-24) encourage smooth, controlled movements during both walking and recovery phases.

## Robot Configuration for Recovery

Unlike standard walking environments that use simplified collision models, VelStand utilizes the **all-collision stand-up MJCF** specified in `MICRODUCK_STANDUP_ROBOT_CFG` ([`src/mjlab_microduck/robot/microduck_constants.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/robot/microduck_constants.py), lines 71-73). This robot configuration enables the simulated Microduck to lie flat on the ground and generate sufficient friction to push off during recovery attempts.

## Training Curriculum and Phases

The VelStand task implements a **three-phase curriculum** defined in lines 62-70 and the `PRONE_RAMP_STAGES` list (lines 53-59) to balance walking and recovery data:

1. **Phase 1**: Clean walking with `fell_over` termination enabled to establish baseline locomotion
2. **Phase 2**: Fall-over disabled, converting falls into recovery training opportunities
3. **Phase 3**: Prone-init ramp that gradually introduces prone and crouch start positions

This structured progression guarantees the policy receives adequate walking demonstrations while accumulating dense recovery experience.

## Termination Logic

The environment employs specific termination conditions to prevent infinite loops. The `fallen_too_long` termination triggers after **8 seconds** in a fallen state (line 305), ending episodes where the robot fails to recover. During play mode, the `fell_over` termination is removed to allow continuous evaluation of recovery capabilities.

## Implementation Example

To instantiate the VelStand environment for training or inference:

```python
from mjlab_microduck.tasks.microduck_velstand_env_cfg import (
    make_microduck_velstand_env_cfg,
)

# Training configuration with curriculum active

train_cfg = make_microduck_velstand_env_cfg(play=False, rough=False)

# Play configuration with curriculum disabled

play_cfg = make_microduck_velstand_env_cfg(play=True, rough=False)

```

Launch a quick smoke test using the CLI:

```bash
uv run train Mjlab-VelStand-Flat-MicroDuck \
    --env.scene.num-envs 64 \
    --agent.max_iterations 5

```

Export the trained policy to ONNX format for real robot deployment:

```bash
uv run scripts/export.py Mjlab-VelStand-Flat-MicroDuck \
    --wandb-run-path <entity/project/run_id>

```

## Summary

- The **VelStand task** unifies walking and recovery into a single policy by extending the velocity environment with conditional rewards
- **Recovery rewards** activate only when tilt exceeds 40° or height drops below 0.08m, remaining dormant during normal walking
- The **MICRODUCK_STANDUP_ROBOT_CFG** collision model allows the robot to interact with the ground plane for push-off maneuvers
- A **three-phase curriculum** ensures balanced training data between locomotion and recovery scenarios
- The **61-dimensional observation space** and **BAM actuator model** maintain compatibility with existing deployment pipelines

## Frequently Asked Questions

### What makes VelStand different from standard walking tasks?

Standard walking tasks terminate episodes immediately upon falling, while VelStand treats falls as learning opportunities. By disabling terminations in later curriculum phases and adding gated recovery rewards, the environment trains policies that can autonomously transition from prone positions back to walking without human intervention or controller switching.

### How does the robot detect when to trigger recovery rewards?

The environment monitors trunk orientation and height through the `upright_progress` and `height_progress` functions defined in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py). When the trunk tilt exceeds 40 degrees or the Z-position drops below 0.08 meters, the system activates recovery rewards including `com_upward_velocity` and `recovery_success` while applying the `fallen_tax` penalty until upright posture is restored.

### Can the trained VelStand policy be deployed to real hardware?

Yes. Because VelStand maintains the same **61-D observation contract** and **BAM actuator model** as other Microduck environments, trained policies export directly to ONNX format without additional wiring. The observation normalization statistics bake into the exported model, allowing seamless transfer from simulation to the physical Microduck robot.

### What files define the VelStand environment configuration?

The primary configuration resides in [`src/mjlab_microduck/tasks/microduck_velstand_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velstand_env_cfg.py), which registers the task IDs `Mjlab-VelStand-Flat-MicroDuck` and `Mjlab-VelStand-Rough-MicroDuck` in [`src/mjlab_microduck/tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/__init__.py). Reward function implementations are located in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py), while the stand-up robot model is defined in [`src/mjlab_microduck/robot/microduck_constants.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/robot/microduck_constants.py).