Why Microduck RL Zero-Pads Command Slots in Its Observation Space

Microduck RL zero-pads unused command slots to maintain a fixed 61-dimensional observation vector across all training tasks, ensuring that policies trained on one environment can be deployed on another without buffer reshaping or ONNX model recompilation.

The pollen-robotics/microduck_rl repository employs a unified observation space architecture where every environment emits the same 61-dimensional tensor, regardless of which motion commands the specific task requires. This design relies on the zero_command_padding mechanism to fill unused command slots with constant zeros, preserving neural network input semantics and runtime compatibility.

The Fixed 61-Dimensional Observation Layout

Microduck RL policies share a single observation layout consisting of three fixed command blocks. The repository standardizes these dimensions to ensure tensor consistency across all environments:

Block Size Description
twist 3 Linear and angular velocity command
head_command 4 Desired neck and head pose
body_command 6 Desired torso pose

Not every environment utilizes all three blocks. For instance, the "sit-stand" and "ground-pick" tasks only process twist commands while lacking active head or body control. Without zero-padding, these environments would produce 47-dimensional observations, breaking the invariant that all policies receive the same buffer layout.

The zero_command_padding Implementation

The helper function zero_command_padding in src/mjlab_microduck/tasks/mdp.py (lines 5115-5225) generates constant-zero tensors to fill missing command dimensions:

def zero_command_padding(env: ManagerBasedRlEnv, dim: int) -> torch.Tensor:
    """Constant-zero obs term of width `dim`.

    Used by envs that don't actively track head/body commands (e.g. sit-stand,
    ground-pick) but still need the unified 61D obs shape so the runtime can
    feed all policies with the same buffer layout.
    """
    return torch.zeros(env.num_envs, dim, device=env.device)

Environments that lack real commands for specific slots invoke this function during configuration to insert the padding terms.

Configuring Zero-Padded Commands in Environment Files

In src/mjlab_microduck/tasks/microduck_velocity_rollers_env_cfg.py (lines 30-38), the roller task demonstrates how to configure zero-padding for unused head and body command slots:


# Command obs parity with the 61D family layout: head/body slots zero-padded

# (the roller task drives heading through the twist slot instead).

for group in ("actor", "critic"):
    cfg.observations[group].terms["head_command"] = ObservationTermCfg(
        func=microduck_mdp.zero_command_padding, params={"dim": 4},
    )
    cfg.observations[group].terms["body_command"] = ObservationTermCfg(
        func=microduck_mdp.zero_command_padding, params={"dim": 6},
    )

Similarly, src/mjlab_microduck/tasks/microduck_standup_env_cfg.py (lines 655-658) implements identical padding for stand-up configurations that do not require head or body pose tracking.

Purpose and Benefits of Zero-Padding

Zero-padding command slots serves four critical functions in the Microduck RL architecture:

  1. Ensures ONNX Model Compatibility – The runtime loads a single ONNX model that expects exactly 61 input dimensions. Zero-padding prevents shape mismatches that would invalidate the compiled model.
  2. Enables Policy Hot-Swapping – By preserving the ordering of twist, head_command, and body_command blocks, a policy trained on one task can immediately execute on another without buffer reshaping or weight matrix realignment.
  3. Stabilizes Neural Network Weights – The first linear layer of each policy contains a fixed weight matrix sized for 61 inputs. Changing the observation dimension would invalidate pretrained weights, requiring retraining from scratch.
  4. Simplifies Buffer Management – Curriculum learning and replay buffers store uniform tensor shapes, eliminating special-case handling for environments that omit specific command types.

Accessing Padded Observations in Training

During rollout, the zero-padded commands appear as contiguous tensor segments. You can extract them from the 61-dimensional observation buffer as follows:

obs = env.step(action)           # shape: (num_envs, 61)

head_cmd = obs[:, 3:7]           # always present, may be all zeros

body_cmd = obs[:, 7:13]          # always present, may be all zeros

By treating "no command" as an explicit neutral input (all zeros), the system allows policies to learn idle behavior while maintaining the capacity to process full command vectors when available.

Summary

  • Microduck RL maintains a fixed 61-dimensional observation space across all environments by zero-padding unused command slots.
  • The zero_command_padding function in src/mjlab_microduck/tasks/mdp.py generates constant-zero tensors for missing head_command (4D) and body_command (6D) blocks.
  • Environments like the velocity roller task configure padding in their observation terms to ensure compatibility with the unified policy architecture.
  • Zero-padding enables ONNX model reuse, policy interchangeability, and simplified buffer management without requiring architectural changes between tasks.

Frequently Asked Questions

Why can't environments simply omit unused command observations?

Omitting unused commands would produce variable tensor shapes (e.g., 47 dimensions instead of 61), breaking the ONNX runtime compatibility and preventing policy transfer between tasks. As implemented in pollen-robotics/microduck_rl, the neural network expects a fixed input size that matches its first layer weight matrix dimensions.

Does zero-padding affect policy learning performance?

No. The zero-padding represents an explicit "no command" signal that the policy can learn to recognize as idle behavior. According to the repository's AGENTS.md documentation, zero-command behavior must be explicitly trained, allowing the policy to interpret the neutral input correctly without destabilizing learning.

Which environments use zero-padded command slots?

Several tasks in the microduck_rl repository utilize zero-padding, including the velocity roller configuration (microduck_velocity_rollers_env_cfg.py) and the stand-up task (microduck_standup_env_cfg.py). Any environment that controls only base velocity while ignoring head or body pose requires this padding mechanism.

How do I add zero-padding to a custom environment?

Import microduck_mdp from mjlab_microduck.tasks and configure an ObservationTermCfg using zero_command_padding with the appropriate dim parameter for the missing command slot (4 for head, 6 for body). Apply this configuration to both actor and critic observation groups to maintain consistency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →