# How Command Slots Are Handled in the Microduck RL Observation Space

> Understand how Microduck RL manages command slots within its observation space. Learn about the 13-element command block and stable observation contracts for seamless task execution.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: how-to-guide
- Published: 2026-09-02

---

**Command slots in Microduck RL are managed through a fixed 61‑dimensional observation vector where the last 13 elements form a command block, populated either with live command values via `mdp.generated_commands` or zero‑padded via `mdp.zero_command_padding` to maintain a stable observation contract across all tasks.**

The `pollen-robotics/microduck_rl` repository implements a modular reinforcement learning framework for humanoid robot control. A critical design decision is how **high-level commands**—such as velocity targets, head poses, and body poses—are exposed to the policy network. Rather than dynamically sized observations, Microduck RL enforces a rigid observation layout that allows model swapping and curriculum learning without reshaping inputs.

## Understanding the Command Block Structure

The observation space is explicitly documented in [`src/mjlab_microduck/tasks/symmetry.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/symmetry.py) and task configuration files. The command block occupies indices 48–60 of the 61‑dimensional vector:

| Index Range | Command Slot | Dimensions | Description |
|-------------|--------------|------------|-------------|
| 48–50 | `twist` | 3 | Linear‑x, linear‑y, angular‑z velocity commands |
| 51–54 | `head_pose` | 4 | Neck pitch, head pitch, head yaw, head roll |
| 55–60 | `body_pose` | 6 | Δx, Δy, Δz position change plus Δroll, Δpitch, Δyaw orientation change |

This fixed layout ensures that every policy, regardless of which commands are active, receives identically shaped inputs.

## How Commands Are Registered and Managed

### Step 1: Define Commands in Task Configuration

Commands are declared in each task's configuration through the `cfg.commands` dictionary. For example, in [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py):

```python
from mjlab_microduck.tasks import mdp as microduck_mdp

cfg.commands["twist"] = microduck_mdp.UniformVelocityCommandCfg(
    resampling_time_range=(0.0, 0.0),
    ranges=((-0.5, 0.5), (-0.5, 0.5), (-0.5, 0.5)),  # x, y, yaw velocities

)

cfg.commands["head_pose"] = microduck_mdp.UniformPoseCommandCfg(
    resampling_time_range=(2.0, 5.0),
    ranges=((-0.3, 0.3), (-0.4, 0.4), (-0.6, 0.6), (-0.3, 0.3)),  # neck_pitch, head_pitch, head_yaw, head_roll

)

cfg.commands["body_pose"] = microduck_mdp.UniformPoseCommandCfg(
    resampling_time_range=(2.0, 5.0),
    ranges=((-0.1, 0.1), (-0.1, 0.1), (-0.1, 0.1),  # Δx, Δy, Δz

            (-0.05, 0.05), (-0.05, 0.05), (-0.05, 0.05)),  # Δroll, Δpitch, Δyaw

)

```

The `CommandManager` instantiates these command generators and maintains a tensor `env.command_manager.active_terms` containing current values for all environments.

### Step 2: Expose Commands Through Observation Terms

The observation manager adds command slots using `ObservationTermCfg` entries in both "actor" and "critic" observation groups. Two patterns handle command slot integration:

**Pattern A: Live Command Injection**

When a command is active, use `mdp.generated_commands`:

```python
from omni.isaac.lab.managers import ObservationTermCfg
from mjlab_microduck.tasks import mdp

for group in ("actor", "critic"):
    cfg.observations[group].terms["twist_command"] = ObservationTermCfg(
        func=mdp.generated_commands,
        params={"command_name": "twist"},
    )
    
    cfg.observations[group].terms["head_command"] = ObservationTermCfg(
        func=mdp.generated_commands,
        params={"command_name": "head_pose"},
    )
    
    cfg.observations[group].terms["body_command"] = ObservationTermCfg(
        func=mdp.generated_commands,
        params={"command_name": "body_pose"},
    )

```

The `generated_commands` function retrieves the specified command from the `CommandManager` and returns it as an observation term.

**Pattern B: Zero Padding for Unused Commands**

When a command slot must be preserved but the command is inactive, use `mdp.zero_command_padding`:

```python
cfg.observations["actor"].terms["head_command"] = ObservationTermCfg(
    func=mdp.zero_command_padding,
    params={"dim": 4},  # Matches head_pose dimensionality

)

```

This returns a zero‑filled vector, maintaining the 61‑D observation shape while providing no actionable signal for that slot.

## Real-World Example: Swapping Command Types

The task configuration [`microduck_velocity_swizzle_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_swizzle_env_cfg.py) demonstrates conditional command slot handling. Some task variants inject live head commands while others zero‑pad:

```python

# Variant WITH head pose control

cfg.commands["head_pose"] = microduck_mdp.UniformPoseCommandCfg(...)
for group in ("actor", "critic"):
    cfg.observations[group].terms["head_command"] = ObservationTermCfg(
        func=mdp.generated_commands,
        params={"command_name": "head_pose"},
    )

# Variant WITHOUT head pose control (head fixed/ignored)

# No "head_pose" entry in cfg.commands

cfg.observations["actor"].terms["head_command"] = ObservationTermCfg(
    func=mdp.zero_command_padding, params={"dim": 4}
)
cfg.observations["critic"].terms["head_command"] = ObservationTermCfg(
    func=mdp.zero_command_padding, params={"dim": 4}
)

```

Both variants produce identical 61‑D observations—the policy cannot distinguish between "zero command" and "no command" without additional context, which simplifies model deployment.

## Implementation Details in Source Files

| File Path | Role in Command Slot Handling |
|-----------|-------------------------------|
| [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) | Defines `generated_commands` and `zero_command_padding` utility functions |
| [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) | Standard task with all three command types fully active |
| [`src/mjlab_microduck/tasks/microduck_velocity_swizzle_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_swizzle_env_cfg.py) | Shows selective command activation and zero‑padding patterns |
| [`src/mjlab_microduck/tasks/symmetry.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/symmetry.py) | Documents the complete 61‑D observation layout including command indices |
| [`src/mjlab_microduck/tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/__init__.py) | Registers task variants and ensures observation compatibility |

## Why Fixed-Size Command Slots Matter

The Microduck RL design prioritizes three operational requirements:

- **Model Portability**: Trained policies can be transferred between tasks without input dimension mismatches
- **Curriculum Compatibility**: Command ranges can be widened gradually while the observation structure remains constant
- **Hardware Deployment**: ONNX exports have predictable input shapes, enabling safe real-time inference on robot controllers

The `zero_command_padding` mechanism is essential for this design—it provides a "don't care" signal that preserves network connectivity without introducing noise from uninitialized or random values.

## Summary

- **Command slots occupy fixed indices** (48–60) in the 61‑D observation vector
- **`mdp.generated_commands`** injects live command values from the `CommandManager`
- **`mdp.zero_command_padding`** supplies zero vectors for inactive command slots
- **Both actor and critic observation groups** receive identical command block treatment
- **Task configurations** in [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py) and variants demonstrate the full pattern

## Frequently Asked Questions

### What happens if a command is defined in `cfg.commands` but never added to observations?

The command will be sampled and updated by the `CommandManager` but remain invisible to the policy. This wastes computation but causes no errors. Best practice is to either expose commands via `generated_commands` or explicitly zero‑pad to document the intentional exclusion.

### Can command slot dimensions change between training and deployment?

No—the 61‑D observation contract is rigid. Changing dimensions requires retraining. The fixed layout in [`symmetry.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/symmetry.py) serves as the cross‑task specification that all policies must respect.

### How does `generated_commands` handle batched environment instances?

The function operates on the full command tensor from `CommandManager`, which has shape `(num_envs, command_dim)`. It returns identically shaped outputs, maintaining batch consistency for vectorized simulation.

### Why zero‑pad instead of removing the slot entirely?

Zero‑padding preserves neural network weight matrices. If a slot were removed, pre‑trained layers expecting 61 inputs would fail. Zero values also provide a consistent "no signal" state that policies can learn to ignore, whereas variable dimensions would require architectural modifications.