Why Turn-in-Place Buckets Are Used in Microduck RL Velocity Walking

Turn-in-place buckets explicitly reserve a fixed fraction of training environments for zero-linear-velocity commands, ensuring the policy learns to rotate the robot’s pivot-heavy chassis without relying on statistically rare random samples.

The Velocity Walking task in pollen-robotics/microduck_rl trains legged locomotion policies to execute both forward translation and on-the-spot rotation. Because the robot’s kinematic design features pivot-heavy geometry, unstructured command sampling rarely produces the specific condition of zero linear velocity paired with high angular velocity required to learn in-place turning. The turn-in-place bucket mechanism solves this data imbalance by deterministically injecting pivot-training episodes into the curriculum.

The Problem with Random Command Sampling

In standard velocity-tracking tasks, commands are sampled randomly from continuous distributions. For pivot-heavy robot geometries, this approach creates a critical learning gap.

  • Statistical rarity: The probability of randomly sampling a command with exactly zero linear velocity and significant angular velocity (≈ ±1 rad s⁻¹) is vanishingly small.
  • Policy short-circuiting: Without explicit exposure to pure rotation, the policy learns to minimize loss by ignoring the angular command component, effectively treating the heading as noise.
  • Real-world failure: When deployed, the robot cannot execute turn-in-place maneuvers, manifesting as either immobility or unstable spinning that was never encountered during training.

How Turn-in-Place Buckets Work

The bucket mechanism functions as a deterministic curriculum injection system that overrides random sampling for a subset of environments.

Command Distribution Mechanics

At each episode reset, the environment allocates a fixed percentage of parallel simulation instances to the turn-in-place bucket:

  • Fraction allocation: Controlled by TURN_IN_PLACE_FRACTION (default 0.15), meaning 15% of environments receive specialized commands.
  • Linear velocity constraint: Forces the linear velocity component to exactly 0.
  • Angular velocity sampling: Samples from a high-magnitude range, specifically |angular velocity| ∈ [0.4, 1.0] rad s⁻¹, ensuring the policy experiences meaningful pivot torques.

Architectural Implementation

The injection occurs inside make_velocity_env_cfg() in src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py. Lines 9–13 define the bucket constant and bind it to the command generator:


# src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py

TURN_IN_PLACE_FRACTION = 0.15  # 15% of envs get lin=0, |ang|∈[0.4, 1.0]

def make_velocity_env_cfg(...):
    cfg = MIT_VelocityRoughEnvCfg()
    cfg.commands["twist"].rel_turn_in_place_envs = TURN_IN_PLACE_FRACTION
    return cfg

This configuration guarantees that, across thousands of parallel environments, a steady 15% of episodes explicitly train the robot to pivot without translation.

Configuring Turn-in-Place Buckets

The fraction is fully configurable for experimental tuning or ablation studies.

Default Configuration (15% Bucket)

Standard training runs automatically use the default distribution:

uv run train velocity-walk --env.scene.num-envs 4096

This invokes the configuration where TURN_IN_PLACE_FRACTION = 0.15 remains active, providing sufficient pivot examples without over-saturating the dataset.

Increasing the Turn Fraction

For robots with particularly difficult turning dynamics or when pivot agility is prioritized, increase the allocation to 30%:


# src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py

TURN_IN_PLACE_FRACTION = 0.30
cfg.commands["twist"].rel_turn_in_place_envs = TURN_IN_PLACE_FRACTION

This doubles the frequency of pure rotation episodes, accelerating convergence on turning behavior at the potential cost of forward-walking sample efficiency.

For ablation studies comparing curriculum versus non-curriculum learning, set the fraction to zero:

TURN_IN_PLACE_FRACTION = 0.0   # No dedicated turn-in-place episodes

cfg.commands["twist"].rel_turn_in_place_envs = TURN_IN_PLACE_FRACTION

Warning: Disabling the bucket typically produces policies that cannot turn on-spot, as documented in the repository’s troubleshooting guide and AGENTS.md design notes.

Source Code References

The implementation spans two primary locations:

  • src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py — Defines the TURN_IN_PLACE_FRACTION constant and injects the value into cfg.commands["twist"].rel_turn_in_place_envs within the make_velocity_env_cfg() factory function.
  • AGENTS.md — Documents the design invariant in the "Velocity Walking" section, specifying that fixed command ranges with angular velocities of ±1.0 rad s⁻¹ make turning learnable through explicit bucketing.

Summary

  • Turn-in-place buckets solve the data imbalance problem for pivot-heavy robots by deterministically allocating 15% of training episodes to zero-linear-velocity, high-angular-velocity commands.
  • Implementation occurs in microduck_velocity_env_cfg.py via the TURN_IN_PLACE_FRACTION constant mapped to rel_turn_in_place_envs.
  • Consequences of removal include policy failure to learn rotation and poor real-world turning performance.
  • Tuning allows increasing the fraction for difficult turning geometries or disabling entirely for ablation studies, though the latter is not recommended for production policies.

Frequently Asked Questions

What happens if I disable turn-in-place buckets?

Without the bucket, the policy rarely encounters pure rotation commands during training. Consequently, it learns to ignore angular velocity targets or produces unstable, untrained pivot motions. As reported in the repository’s troubleshooting documentation, this results in robots that can walk forward but fail to turn on-spot in real-world deployment.

Why is the default fraction set to 15%?

The 15% value represents a balance between providing sufficient pivot-training samples and maintaining adequate coverage of forward-walking behaviors. According to the design notes in AGENTS.md, this fraction ensures the policy receives enough turning experience to learn pivot dynamics without over-fitting to rotation at the expense of translation skills.

Can I use turn-in-place buckets with other locomotion tasks?

The bucket mechanism is specific to the Velocity Walking task implemented in microduck_velocity_env_cfg.py. Other tasks, such as standing or trotting configurations, may use different command distributions. However, the architectural pattern—using rel_turn_in_place_envs in the command configuration—can be adapted to any task requiring explicit rotation training by setting the corresponding parameter in the environment configuration.

Where is the angular velocity range for turning defined?

The specific range of [0.4, 1.0] rad s⁻¹ for angular velocity magnitude during turn-in-place episodes is defined within the command sampling logic referenced by cfg.commands["twist"]. While the configuration file sets the bucket fraction, the underlying sampling distribution parameters are implemented in the task’s MDP configuration files, ensuring that turn-in-place episodes provide sufficiently challenging rotational commands.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →