# Why Microduck RL Zero-Pads Command Slots in Its Observation Space

> Microduck RL zero-pads command slots to ensure fixed observation vectors, enabling seamless policy deployment across diverse tasks without recompilation. Learn how.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: internals
- Published: 2026-09-01

---

**Microduck RL zero-pads unused command slots to maintain a fixed 61-dimensional observation vector across all training tasks, ensuring that policies trained on one environment can be deployed on another without buffer reshaping or ONNX model recompilation.**

The `pollen-robotics/microduck_rl` repository employs a unified observation space architecture where every environment emits the same 61-dimensional tensor, regardless of which motion commands the specific task requires. This design relies on the `zero_command_padding` mechanism to fill unused command slots with constant zeros, preserving neural network input semantics and runtime compatibility.

## The Fixed 61-Dimensional Observation Layout

Microduck RL policies share a single observation layout consisting of three fixed command blocks. The repository standardizes these dimensions to ensure tensor consistency across all environments:

| Block | Size | Description |
|-------|------|-------------|
| `twist` | 3 | Linear and angular velocity command |
| `head_command` | 4 | Desired neck and head pose |
| `body_command` | 6 | Desired torso pose |

Not every environment utilizes all three blocks. For instance, the "sit-stand" and "ground-pick" tasks only process twist commands while lacking active head or body control. Without zero-padding, these environments would produce 47-dimensional observations, breaking the invariant that **all policies receive the same buffer layout**.

## The zero_command_padding Implementation

The helper function `zero_command_padding` in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) (lines 5115-5225) generates constant-zero tensors to fill missing command dimensions:

```python
def zero_command_padding(env: ManagerBasedRlEnv, dim: int) -> torch.Tensor:
    """Constant-zero obs term of width `dim`.

    Used by envs that don't actively track head/body commands (e.g. sit-stand,
    ground-pick) but still need the unified 61D obs shape so the runtime can
    feed all policies with the same buffer layout.
    """
    return torch.zeros(env.num_envs, dim, device=env.device)

```

Environments that lack real commands for specific slots invoke this function during configuration to insert the padding terms.

## Configuring Zero-Padded Commands in Environment Files

In [`src/mjlab_microduck/tasks/microduck_velocity_rollers_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_rollers_env_cfg.py) (lines 30-38), the roller task demonstrates how to configure zero-padding for unused head and body command slots:

```python

# Command obs parity with the 61D family layout: head/body slots zero-padded

# (the roller task drives heading through the twist slot instead).

for group in ("actor", "critic"):
    cfg.observations[group].terms["head_command"] = ObservationTermCfg(
        func=microduck_mdp.zero_command_padding, params={"dim": 4},
    )
    cfg.observations[group].terms["body_command"] = ObservationTermCfg(
        func=microduck_mdp.zero_command_padding, params={"dim": 6},
    )

```

Similarly, [`src/mjlab_microduck/tasks/microduck_standup_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_standup_env_cfg.py) (lines 655-658) implements identical padding for stand-up configurations that do not require head or body pose tracking.

## Purpose and Benefits of Zero-Padding

Zero-padding command slots serves four critical functions in the Microduck RL architecture:

1. **Ensures ONNX Model Compatibility** – The runtime loads a single ONNX model that expects exactly 61 input dimensions. Zero-padding prevents shape mismatches that would invalidate the compiled model.
2. **Enables Policy Hot-Swapping** – By preserving the ordering of `twist`, `head_command`, and `body_command` blocks, a policy trained on one task can immediately execute on another without buffer reshaping or weight matrix realignment.
3. **Stabilizes Neural Network Weights** – The first linear layer of each policy contains a fixed weight matrix sized for 61 inputs. Changing the observation dimension would invalidate pretrained weights, requiring retraining from scratch.
4. **Simplifies Buffer Management** – Curriculum learning and replay buffers store uniform tensor shapes, eliminating special-case handling for environments that omit specific command types.

## Accessing Padded Observations in Training

During rollout, the zero-padded commands appear as contiguous tensor segments. You can extract them from the 61-dimensional observation buffer as follows:

```python
obs = env.step(action)           # shape: (num_envs, 61)

head_cmd = obs[:, 3:7]           # always present, may be all zeros

body_cmd = obs[:, 7:13]          # always present, may be all zeros

```

By treating "no command" as an explicit neutral input (all zeros), the system allows policies to learn idle behavior while maintaining the capacity to process full command vectors when available.

## Summary

- Microduck RL maintains a **fixed 61-dimensional observation space** across all environments by zero-padding unused command slots.
- The `zero_command_padding` function in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) generates constant-zero tensors for missing `head_command` (4D) and `body_command` (6D) blocks.
- Environments like the velocity roller task configure padding in their observation terms to ensure compatibility with the unified policy architecture.
- Zero-padding enables **ONNX model reuse**, **policy interchangeability**, and **simplified buffer management** without requiring architectural changes between tasks.

## Frequently Asked Questions

### Why can't environments simply omit unused command observations?

Omitting unused commands would produce variable tensor shapes (e.g., 47 dimensions instead of 61), breaking the ONNX runtime compatibility and preventing policy transfer between tasks. As implemented in `pollen-robotics/microduck_rl`, the neural network expects a fixed input size that matches its first layer weight matrix dimensions.

### Does zero-padding affect policy learning performance?

No. The zero-padding represents an explicit "no command" signal that the policy can learn to recognize as idle behavior. According to the repository's [`AGENTS.md`](https://github.com/pollen-robotics/microduck_rl/blob/main/AGENTS.md) documentation, zero-command behavior must be explicitly trained, allowing the policy to interpret the neutral input correctly without destabilizing learning.

### Which environments use zero-padded command slots?

Several tasks in the `microduck_rl` repository utilize zero-padding, including the velocity roller configuration ([`microduck_velocity_rollers_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_rollers_env_cfg.py)) and the stand-up task ([`microduck_standup_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_standup_env_cfg.py)). Any environment that controls only base velocity while ignoring head or body pose requires this padding mechanism.

### How do I add zero-padding to a custom environment?

Import `microduck_mdp` from `mjlab_microduck.tasks` and configure an `ObservationTermCfg` using `zero_command_padding` with the appropriate `dim` parameter for the missing command slot (4 for head, 6 for body). Apply this configuration to both actor and critic observation groups to maintain consistency.