# Microduck RL Locomotion Task: Velocity Tracking Environment and Design Decisions

> Explore the Microduck RL velocity tracking locomotion task. Learn about the bipedal robot's design decisions for following velocity and head-pose commands.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-01

---

**The central locomotion task in Microduck RL is the velocity-tracking (walking) environment defined in [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py), which trains the 800g bipedal robot to follow linear and angular velocity commands while simultaneously tracking head-pose targets.**

The Microduck RL repository provides reinforcement learning environments specifically tailored for a lightweight bipedal hardware platform. The primary locomotion task focuses on velocity command tracking, engineered with precise constraints to ensure robust Sim-to-Real transfer. This environment encapsulates the physical robot’s limitations while providing a challenging learning domain for end-to-end gait policies.

## The Velocity-Tracking Locomotion Task

The main locomotion task resides in [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py), where the `make_microduck_velocity_env_cfg` function constructs the environment. According to the module docstring (lines 3–15), the task is explicitly defined as **velocity-command tracking plus head-pose commands**【/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L3-L15】.

The 800g bipedal robot must track velocity commands comprising `lin_vel_x`, `lin_vel_y`, and `ang_vel_z` while maintaining specified head orientations. This dual-objective design reflects the physical robot’s operational requirements, where stable walking must coexist with directed gaze control.

## Core Design Decisions for Locomotion

### Command Structure and Curriculum Strategy

**Turn-in-place sampling** is set to **15%** of episodes (`rel_turn_in_place_envs=0.15`) with linear velocity set to zero and angular velocity sampled uniformly from [0.4, 1.0] rad/s. Without this explicit curriculum bucket, the policy rarely learns to spin on the spot, as natural training data contains only approximately 2% of such instances【L11-L13】.

**Fixed command ranges** limit linear velocity to ±0.4 m/s and angular velocity to ±1.0 rad/s. Wider curricula were found to outpace the robot’s physical capabilities; these restricted ranges keep turning learnable while maintaining task stability【L9-L10, L48-L51】.

### Reward Function Engineering

**Foot slip penalty** is deliberately weakened to **-0.1** rather than the default -1.0. A strong slip penalty proved too restrictive for the robot’s pivot-heavy turning mechanics, so the softened penalty allows necessary foot sliding during rotation while still discouraging excessive slipping【L7-L9】.

**Head-pose tracking** carries a primary reward weight of **2.0** with per-joint Gaussian standard deviation of 0.5. An EMA-based **head-pose-bias** penalty is added later in training to correct systematic droop, providing a strong initial signal for head orientation while allowing subsequent refinement【L13-L15, L11-L14】.

**Body-pose slot** remains active with weight **0** as a placeholder. This maintains the observation layout consistency across the task family, though actual body-pose tracking is enabled exclusively in the stand-up environment【L15-L16】.

### Domain Randomization for Sim-to-Real

The environment implements extensive randomization toggles to bridge the simulation-to-reality gap. **Center-of-mass randomization** (`ENABLE_COM_RANDOMIZATION = True` at line 31), **head-COM perturbations**, **joint friction variation**, **armature randomization**, **velocity pushes**, **IMU orientation noise**, and other physics parameters are explicitly listed and enabled【L30-L42】.

**Encoder bias domain randomization** applies a per-environment constant joint-encoder offset exclusively to the actor observations, keeping the critic privileged. This aligns simulation with real hardware where encoders exhibit static biases【L26-L34】.

**Observation noise** uses reduced levels for gravity, joint positions, velocities, and angular velocity (e.g., `n_min=-0.01, n_max=0.01` for gravity). These values mirror realistic sensor noise without destabilizing the learned policy【L86-L90】.

### Observation Space Layout

The actor receives a **61-dimensional observation vector**: 48 dimensions of proprioception combined with a 13-dimensional command block. The command block strictly follows the structure: `twist` (3D), `head_pose` (4D), and `body_pose` (6D). This invariant layout is enforced across the entire policy family to maintain stable ONNX export contracts, as documented in [`AGENTS.md`](https://github.com/pollen-robotics/microduck_rl/blob/main/AGENTS.md).

### Terrain and Contact Handling

Optional **rough terrain** generation uses a Microduck-specific generator that produces harsher terrain than default implementations. To maintain contact solver stability despite this difficulty, the environment employs `_soften_terrain_contacts` and increases solver iterations (lines 55–71).

## Implementation and Training Workflow

Instantiate the velocity-tracking environment with configuration flags for play mode and terrain roughness:

```python
from mjlab_microduck.tasks.microduck_velocity_env_cfg import make_microduck_velocity_env_cfg

env_cfg = make_microduck_velocity_env_cfg(play=False, rough=False)

```

Run a smoke test to validate configuration before full training:

```bash
uv run train velocity --env.scene.num-envs 64 --agent.max_iterations 5

```

Inspect critical training parameters programmatically:

```python
print("Linear velocity range:", env_cfg.commands["twist"].ranges.lin_vel_x)
print("Turn-in-place fraction:", env_cfg.commands["twist"].rel_turn_in_place_envs)
print("Action-rate weight schedule:", env_cfg.curriculum["action_rate_weight"].params["weight_stages"])

```

Export trained policies to ONNX format including the observation normalizer:

```bash
uv run scripts/export.py velocity --wandb-run-path <entity/project/run_id>

```

## Summary

- The primary locomotion task is **velocity tracking with head-pose control**, implemented in [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py).
- **15% turn-in-place sampling** and constrained command ranges (±0.4 m/s linear, ±1.0 rad/s angular) ensure learnable pivoting behavior.
- **Softened foot-slip penalties (-0.1)** accommodate necessary pivoting while head-pose tracking uses a high reward weight (2.0) with Gaussian shaping.
- **61-dimensional observations** maintain strict contracts across the policy family, with 48 proprioception and 13 command dimensions.
- **Comprehensive domain randomization** (COM, friction, encoder bias, observation noise) enables robust Sim-to-Real transfer for the 800g hardware.

## Frequently Asked Questions

### What defines the main locomotion task in Microduck RL?

The main locomotion task is the **velocity-tracking environment** that commands the 800g bipedal robot to follow specified linear and angular velocities while tracking head-orientation targets. It is defined in [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) and represents the primary walking skill used for Sim-to-Real deployment.

### Why is the foot slip penalty set to -0.1 instead of -1.0?

The **weak foot-slip penalty (-0.1)** prevents over-restriction during the robot’s pivot-heavy turning maneuvers. A penalty of -1.0 constrained the policy too severely, preventing effective in-place rotation learned via the 15% turn-in-place curriculum fraction.

### How does the observation space support Sim-to-Real deployment?

The **fixed 61-dimensional observation layout** (48 proprioception + 13 commands) creates a stable interface for ONNX export. Domain randomization techniques including encoder bias injection and reduced sensor noise levels (e.g., ±0.01 gravity noise) prepare policies to handle real-hardware sensor characteristics without destabilizing training.

### What curriculum stages are used during training?

The environment employs staged curriculum ramps for **action-rate weight**, **standing-environment fraction**, **head-pose range**, **CoM randomization**, and **head-pose-bias weight** (beginning at line 76). These stages ensure smooth skill acquisition before introducing aggressive regularization or randomization that could destabilize early learning.