Microduck RL Observation Contracts: Understanding the 61-Dimensional Policy Interface

Microduck RL policies require a fixed 61-dimensional observation vector composed of 48 proprioceptive sensor values and 13 command parameters, shared identically between actor and critic networks.

The observation contract in the pollen-robotics/microduck_rl repository defines a strict, immutable interface that every policy must follow. This design ensures compatibility across training, ONNX export, and real-robot deployment without runtime surprises.

The Fixed 61-Dimensional Observation Structure

The observation contract splits the 61-dimensional vector into two distinct blocks:

  • Proprioceptive sensors (48 values) — Raw robot feedback including joint positions, joint velocities, IMU readings, and foot contact states. This block remains identical for every policy in the family.
  • Command block (13 values) — Structured control inputs concatenated in strict order:
    • twist (3 values): linear x/y/z and angular x/y/z velocity commands
    • head_pose (4 values): head-pitch, head-yaw, head-roll, plus a bias term
    • body_pose (6 values): body-pitch, body-yaw, body-roll, plus three additional pose parameters

Unused command slots are zero-padded to preserve the fixed 61-dimensional layout. This padding guarantees that policy networks always receive the same tensor shape regardless of which command modes are active.

Observation Groups for Actor and Critic

The codebase defines two observation groups that both consume the identical 61-dimensional structure:

  • actor — Feeds the policy's action-selection network
  • critic — Feeds the value-estimation network

Both groups set nan_policy = "sanitize" to replace transient NaN values (e.g., from contact detection glitches) with safe defaults rather than crashing training. This defensive configuration is verified by unit tests and applied consistently across all environment configurations.

Where the Contract Lives in the Source Code

The observation schema is defined and enforced across several key files:

Core Schema Definition

In src/mjlab_microduck/tasks/mdp.py, the observation schema assembles the 48-element proprioception tensor with the 13-element command block. This file serves as the single source of truth for observation composition.

Symmetry and Augmentation Guarantees

src/mjlab_microduck/tasks/symmetry.py contains two critical tensors that preserve the contract during data augmentation:

  • _OBS_PERM — Permutation indices ensuring mirrored observations maintain field ordering
  • _OBS_SIGN — Sign flips for vector quantities that reverse direction under mirroring

These tensors hardcode the exact 61-field structure, making any layout change immediately detectable.

Environment Configurations

All task environments inherit this layout. For example, src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py registers the observation contract for the primary walking task, while src/mjlab_microduck/tasks/microduck_spin_env_cfg.py reuses the same structure for spinning behaviors.

Working with Observations in Training

The observation dictionary returned by environments follows this access pattern:

import torch
from mjlab_microduck.tasks.symmetry import augment_obs

def step(env, policy):
    # env.obs returns a dict with 'actor' and 'critic' tensors

    obs = env.obs  # TensorDict { "actor": [B, 61], "critic": [B, 61] }

    # Optional data-augmentation (mirroring) – keeps the 61-dim contract

    aug_obs, _ = augment_obs(obs)

    # Forward pass through the actor network

    actions = policy.actor(aug_obs["actor"])

    # Apply actions in the environment

    env.step(actions)

The augment_obs function applies the permutation and sign tensors from symmetry.py, ensuring that mirrored observations remain valid 61-dimensional inputs that the policy can process.

Export and Deployment Implications

The observation contract becomes immutable once baked into exported models. The scripts/export.py workflow embeds the 61-dimensional expectation directly into the ONNX graph:


# scripts/export.py (excerpt)

obs_cfg = env.cfg.observations["actor"]
if obs_cfg.obs_normalization:
    # The normalizer is baked; the exported ONNX expects the same 61-dim vector

    policy.export_onnx("policy.onnx", obs_dim=61)

Any deviation between training-time and deployment-time observations causes immediate runtime failure on the physical robot. This constraint motivates the extensive test coverage that locks the contract in place.

Test Coverage for Contract Enforcement

Several unit tests verify that the observation contract remains intact:

These tests run in CI to catch any regression that would break policy compatibility.

Inference-Time Observation Handling

The scripts/infer_policy.py script demonstrates how deployed policies consume the 61-dimensional vector:


# scripts/infer_policy.py demonstrates ONNX runtime usage

# The exported model expects exactly 61 float32 values in the defined order

Real-world deployment requires sensor stack implementations that produce the identical 48+13 value ordering used during training.

Summary

Frequently Asked Questions

What happens if my sensors provide fewer than 48 proprioceptive values?

You must zero-pad or compute missing values to reach exactly 48 dimensions before concatenating with the 13 command values. The policy network was trained on the full 61-dimensional layout and will fail silently or produce garbage outputs with incorrect input shapes.

Can I change the command block ordering for my specific task?

No. The _OBS_PERM and _OBS_SIGN tensors in src/mjlab_microduck/tasks/symmetry.py hardcode field positions. Reordering commands without updating these tensors breaks data augmentation and would require retraining all policies from scratch.

How does the NaN sanitization policy work in practice?

Both observation groups set nan_policy = "sanitize", which replaces any NaN sensor reading with a predefined safe value (typically zero or the running mean) before the tensor reaches the neural network. This prevents gradient corruption and training crashes from transient hardware glitches.

Why do actor and critic use the same 61-dimensional observation instead of different views?

Shared observations simplify the export pipeline and guarantee that value estimates remain coherent with the actor's state representation. The design choice trades potential representational efficiency for deployment reliability—critical when the same ONNX model must run on diverse hardware targets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →