Microduck RL Observation Contract: 61-Dimensional Vector Specification
Microduck RL policies enforce a fixed 61-dimensional observation vector split into 48-dimensional proprioception and 13-dimensional command blocks, standardized across actor and critic networks for hot-swappable deployment.
The pollen-robotics/microduck_rl repository defines a strict observation contract that ensures reinforcement learning policies can be trained and deployed interchangeably without tensor reshaping. This specification guarantees that every environment variant produces identical observation structures, enabling policy hot-swapping at runtime.
61-Dimensional Observation Vector Structure
The observation contract mandates a fixed-size 61-dimensional vector used identically by both the actor and critic networks. According to the repository's AGENTS.md documentation and the implementation in src/mjlab_microduck/tasks/symmetry.py, the vector divides into two logical blocks.
Base Proprioception Block (48 Dimensions)
The first 48 elements contain low-level robot state information:
- 14 joint positions (relative)
- 14 joint velocities (relative)
- 14 last-action values (previous policy output)
- Base angular velocity (3 dimensions: roll, pitch, yaw)
- Projected gravity (3 dimensions: gx, gy, gz)
Command Block (13 Dimensions)
The remaining 13 elements encode high-level motion commands:
- Twist command (3 dimensions: linear-x, linear-y, angular-z)
- Head command (4 dimensions: neck pitch, head pitch, head yaw, head roll)
- Body command (6 dimensions: x, y, z, roll, pitch, yaw)
Exact Index Layout and Tensor Ordering
When flattened, the observation tensor follows this strict index ordering as documented in src/mjlab_microduck/tasks/symmetry.py:
| Index Range | Component | Dimensions |
|---|---|---|
[0:3] |
Base angular velocity | 3 |
[3:6] |
Projected gravity | 3 |
[6:20] |
Joint positions (relative) | 14 |
[20:34] |
Joint velocities (relative) | 14 |
[34:48] |
Last action | 14 |
[48:51] |
Twist command | 3 |
[51:55] |
Head command | 4 |
[55:61] |
Body command | 6 |
This layout ensures that zero-padding occurs uniformly when a specific task does not utilize particular command slots. Unused command dimensions must still exist in the observation vector and contain zeros rather than being omitted.
Actor and Critic Observation Groups
Both networks receive the identical 61-dimensional observation, though they access it through different keys in a TensorDict structure. The src/mjlab_microduck/tasks/mdp.py module constructs these observation groups during environment initialization.
# Example: Accessing observations inside a training loop
import torch
from tensordict import TensorDict
def step(env):
# Environment returns a TensorDict with keys "actor" and "critic"
obs: TensorDict = env.reset()
actor_obs: torch.Tensor = obs["actor"] # shape [B, 61]
critic_obs: torch.Tensor = obs["critic"] # shape [B, 61]
# Extract base angular velocity (indices 0-2)
base_ang_vel = actor_obs[:, 0:3]
# Extract joint positions (indices 6-19)
joint_pos = actor_obs[:, 6:20]
# Extract twist command (indices 48-50)
twist_cmd = actor_obs[:, 48:51]
When exported to ONNX format via scripts/export.py, the observation is passed as a plain NumPy array of shape (batch_size, 61) with the input node named "obs".
Contract Enforcement and Validation
The repository enforces the observation contract through compile-time constants, automated testing, and configuration validation.
Manifest Constants and OBS_LEN
The src/mjlab_microduck/publish/manifest.py module defines the canonical constants:
OBS_LEN = 61
ACTION_LEN = 14
These values are validated against the daemon's expectations in tests/test_publish_manifest.py, ensuring that published policies advertise the correct tensor dimensions. Any policy deviating from these constants fails the manifest validation suite.
Zero-Padding and Normalization Requirements
The observation contract specifies three critical behavioral requirements:
- Fixed dimensions: The observation vector must always contain exactly 61 elements, regardless of task configuration.
- Zero-padding: Unused command slots (e.g., head commands in locomotion-only tasks) must be explicitly zero-filled rather than removed.
- Normalization: Observations are normalized using statistics baked into the exported ONNX model. The raw environment outputs are not normalized; the policy model handles normalization internally.
Additionally, the configuration specifies nan_policy = "sanitize" to handle rare sensor glitches, ensuring training stability when physical sensors produce NaN values.
# Example: Loading an exported ONNX policy with the 61-dim contract
import onnxruntime as ort
import numpy as np
session = ort.InferenceSession("my_policy.onnx")
# Create dummy observation respecting the contract
obs = np.zeros((1, 61), dtype=np.float32)
# ONNX model expects input named "obs"
actions = session.run(["actions"], {"obs": obs})[0]
print("Action shape:", actions.shape) # (1, 14)
# Example: Validating a custom environment respects the contract
def test_observation_shape(cfg):
assert cfg.observations["actor"].obs_len == 61
assert cfg.observations["critic"].obs_len == 61
# Verify all required observation terms exist
for grp in cfg.observations:
terms = cfg.observations[grp].terms
assert set(terms.keys()) == {
"base_ang_vel", "gravity", "joint_pos", "joint_vel",
"last_action", "twist_cmd", "head_cmd", "body_cmd"
}
Summary
- Microduck RL policies require a 61-dimensional observation vector defined by the
OBS_LENconstant insrc/mjlab_microduck/publish/manifest.py. - The vector comprises 48 dimensions of proprioception (joint states, base velocity, gravity) and 13 dimensions of commands (twist, head, body).
- Both actor and critic networks receive identical observations through a
TensorDictwith keys"actor"and"critic". - The contract mandates zero-padding for unused command slots and strict adherence to the index ordering specified in
src/mjlab_microduck/tasks/symmetry.py. - Exported ONNX models embed normalization statistics and expect input tensors of shape
[batch, 61]with the node name"obs".
Frequently Asked Questions
What happens if my observation vector has a different length than 61?
The policy will fail to load or execute. The manifest validation in tests/test_publish_manifest.py explicitly checks that OBS_LEN == 61, and the ONNX model's input layer expects exactly 61 features. Deviating from this dimensionality breaks the hot-swapping capability and causes runtime shape mismatches.
How are unused command slots handled in the observation vector?
Unused slots must be zero-padded. The observation contract requires all 61 dimensions to be present even if a specific task does not utilize head or body commands. The src/mjlab_microduck/tasks/mdp.py module ensures these slots are explicitly filled with zeros rather than omitted, maintaining consistent tensor shapes across all environment variants.
Where is the observation layout documented in the source code?
The canonical documentation resides in AGENTS.md (lines 49-55), with the implementation details explicitly commented in src/mjlab_microduck/tasks/symmetry.py (lines 9-18). The src/mjlab_microduck/tasks/mdp.py module constructs the observation groups according to this specification, while tests/test_spin_cfg.py validates that all environment variants maintain identical observation term sets.
Does the observation contract differ between training and deployment?
No. The contract remains identical between training and deployment phases. During training, observations are delivered as TensorDict objects with separate "actor" and "critic" keys, each containing the 61-dimensional vector. During deployment, the exported ONNX model receives the same 61-dimensional array as a single input tensor named "obs", with normalization handled internally by the model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →