How Microduck RL Handles Unused Command Slots in the Observation Space
Microduck RL environments always expose a fixed-size 61‑dimensional observation vector where unused command slots are zero‑padded by the zero_command_padding helper to maintain consistent layouts for policy portability.
The pollen-robotics/microduck_rl repository implements a canonical observation space that remains constant across every task family. When a specific environment does not utilize certain command modalities—such as head pose or body pose controls—the framework substitutes missing values with zero tensors rather than collapsing the vector dimensions.
The Fixed 61‑Dimensional Observation Architecture
Microduck RL constructs observations from three distinct command blocks that collectively occupy 13 dimensions within the full 61‑D vector:
| Block | Size | Meaning |
|---|---|---|
| Twist command | 3 | Linear‑X, Linear‑Y, Angular‑Z velocities |
| Head‑pose command | 4 | Δneck‑pitch, Δhead‑pitch, Δhead‑yaw, Δhead‑roll |
| Body‑pose command | 6 | Δx, Δy, Δz, Δroll, Δpitch, Δyaw |
The remaining dimensions contain proprioceptive state data such as joint positions, velocities, and base orientation. Because the policy network expects identical input shapes for every task—including during ONNX export—the observation layout must never change, even when specific command slots are irrelevant to the current task.
Zero‑Padding Unused Command Slots with zero_command_padding
When an environment does not require a particular command slot, the framework invokes zero_command_padding defined in src/mjlab_microduck/tasks/mdp.py (around line 5115). This helper returns a tensor of zeros with the requested shape, ensuring the slot contributes a constant‑zero vector to the observation.
# src/mjlab_microduck/tasks/mdp.py
def zero_command_padding(dim: int) -> torch.Tensor:
"""Return a tensor of zeros with the requested dimension."""
# `env` is injected by the observation manager; we ignore the actual command.
return torch.zeros(env.num_envs, dim, device=env.device)
The function ignores any actual command data from the environment and instead emits a fixed zero tensor sized to match the expected slot dimension—either 4 for head commands or 6 for body commands.
Configuring Padded Slots in Environment Definitions
Each task configuration explicitly declares every command slot and assigns the zero‑padding function to slots the task does not actively command. For example, in the "rollers" velocity task defined in src/mjlab_microduck/tasks/microduck_velocity_rollers_env_cfg.py, the head and body command slots are padded while the twist command contains real values:
# Example from the "rollers" task configuration
cfg.observations[group].terms["head_command"] = ObservationTermCfg(
func=microduck_mdp.zero_command_padding, params={"dim": 4},
)
cfg.observations[group].terms["body_command"] = ObservationTermCfg(
func=microduck_mdp.zero_command_padding, params={"dim": 6},
)
Other configurations like microduck_spin_env_cfg.py and microduck_standup_env_cfg.py follow the same pattern, padding body commands while spinning or standing. Conversely, microduck_velocity_swizzle_env_cfg.py uses the real head command while keeping the body slot zero‑padded, illustrating mixed usage patterns.
Why Observation Layout Invariants Matter
According to the obs‑layout invariants documented in AGENTS.md, the 61‑D observation shape must remain constant across all policies. Changing the number of command dimensions would:
- Break policy compatibility: Pre-trained weights expect fixed input dimensions.
- Prevent ONNX export: Runtime crashes occur when exporting policies if input shapes fluctuate between tasks.
- Force costly re‑training: Every task variant would require separate network architectures without zero‑padding.
Zero‑padding preserves the vector semantics so that unused commands contribute null information without altering the tensor geometry.
Runtime Behavior and Accessing Padded Observations
During simulation, the command manager may contain no entry for a padded slot, yet the observation manager still reads it via the zero‑padding term. The returned observation maintains exactly 61 dimensions regardless of which commands are active.
# Accessing the padded command in an episode (will be all zeros)
obs = env.observe() # shape (num_envs, 61)
head_cmd = obs[:, 51:55] # zeros if the task does not use head pose
twist_cmd = obs[:, 48:51] # actual velocity commands
The slice indices above illustrate how specific command blocks are extracted from the fixed observation buffer; padded sections consistently return tensors of zeros.
Summary
- Microduck RL observation space is fixed at 61 dimensions across all tasks.
- Unused command slots are handled by
zero_command_paddinginsrc/mjlab_microduck/tasks/mdp.py, which returns zero tensors of size 4 or 6. - Configuration files such as
microduck_velocity_rollers_env_cfg.pyexplicitly assign zero‑padding to inactive slots usingObservationTermCfg. - Policy portability relies on this zero‑padding strategy to maintain layout invariants required for ONNX export and cross‑task weight sharing.
Frequently Asked Questions
What is the exact size of the Microduck RL observation vector?
The observation vector is always 61 dimensions, composed of 13 command dimensions (3 twist + 4 head + 6 body) plus 48 dimensions of proprioceptive state data. This size is invariant across all task configurations in the repository.
Which command slots can be zero-padded in Microduck RL?
The head‑pose command (4 dimensions) and body‑pose command (6 dimensions) slots can be zero‑padded. The twist command (3 dimensions) is typically active in most tasks, though the framework supports padding any slot via the zero_command_padding helper.
Where is the zero‑padding logic implemented?
The zero‑padding logic is implemented in src/mjlab_microduck/tasks/mdp.py within the zero_command_padding function (approximately line 5115). This function is referenced by environment configuration files to fill unused observation slots with tensors of zeros.
Why can't the observation space size change between different Microduck RL tasks?
Changing the observation space size would violate the obs‑layout invariants specified in AGENTS.md, breaking policy network compatibility and preventing successful ONNX export. Fixed dimensions ensure that a single policy architecture can operate across diverse tasks without re‑initialization or architecture modifications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →