# How Runtime Policy Hot‑Swapping Is Enabled in Microduck RL

> Discover how Microduck RL enables runtime policy hot-swapping using an immutable observation contract and a JSON manifest for automated robotctl commands. Learn more.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: internals
- Published: 2026-09-08

---

**Microduck RL enables runtime policy hot‑swapping by enforcing a single, immutable 61‑dimensional observation contract across all training environments and publishing a JSON manifest that directs the robot daemon to load ONNX policies via automated `robotctl` commands.**

Runtime policy hot‑swapping allows physical robots to switch between reinforcement learning policies without restarting services or redeploying firmware. In the `pollen-robotics/microduck_rl` repository, this capability is built on a strict observation schema, baked‑in normalization, and a declarative manifest system that the runtime daemon consumes to validate and activate new policies instantly. The architecture is documented in [`AGENTS.md`](https://github.com/pollen-robotics/microduck_rl/blob/main/AGENTS.md), which describes the constant‑command family and the importance of a unified observation contract.

## The Foundation: A Unified Observation Contract

All policies in Microduck RL must conform to an identical observation layout. This immutability guarantee ensures that any trained network can be substituted for another at runtime without shape mismatches or preprocessing errors.

### The 61‑Dimensional Observation Layout

Every training environment emits a fixed **61‑dimensional observation vector** comprising **48 proprioception values** plus **13 command dimensions**. This layout is hard‑coded in task configuration modules such as [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) and enforced by the export pipeline.

Because the observation length never varies, the robot daemon can allocate fixed buffers and validation logic once, then accept any policy trained with this schema. The corresponding action space is always **14 dimensions**, matching the robot's actuator count.

## ONNX Export with Baked Normalization

The normalization layer is eliminated from runtime overhead by baking it directly into the exported neural network.

In [`src/mjlab_microduck/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/export.py), the `run_export` function serializes `actor(normalizer(obs))` as a single ONNX graph. This means the robot feeds raw MJCF observations directly into the network without additional preprocessing, reducing latency and eliminating version skew between training statistics and runtime normalization.

## The Policy Manifest System

Microduck RL uses a declarative manifest to bridge the training framework and the robot daemon. The CLI wrapper in [`src/mjlab_microduck/publish/cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/publish/cli.py) orchestrates the full publish flow, calling the manifest builder and export utilities.

### Schema 2 and Slot Hints

[`src/mjlab_microduck/publish/manifest.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/publish/manifest.py) generates a JSON manifest (schema version 2) containing:

- **`obs_len`**: Fixed at `61`
- **`action_len`**: Fixed at `14`
- **Slot hints**: Categorical labels such as `walk`, `stand`, or `sitstand` that map the policy to a runtime slot in the daemon

These slot hints determine which behavioral mode the policy replaces when hot‑swapped.

### Auto‑Generated robotctl Commands

The manifest builder automatically generates the appropriate daemon commands via `install_commands`. For **perpetual** policies that run indefinitely, it produces:

```bash
sudo robotctl policy load <slot> <repo_id>

```

For **episodic** policies that run for fixed durations, it generates:

```bash
sudo robotctl policy add <slot> <repo_id>

```

The helper logic inspects the `kind` field in the manifest to select the correct command variant, ensuring the daemon receives the proper instruction without manual scripting.

## Runtime Daemon Integration

While the daemon itself lives in the separate `pollen-robotics/microduck` repository, Microduck RL produces all artifacts required for its hot‑swap protocol. When a new policy is published:

1. The daemon reads the manifest and validates that the ONNX input shape matches `obs_len` (61)
2. It registers the policy under the indicated slot (e.g., `walk`)
3. It executes the generated `robotctl` command to unload the current policy and activate the new one

Because all policies share the same observation shape and normalization is baked into the ONNX file, the transition occurs without restarting the robot or recompiling code.

## Practical Implementation

The following workflow demonstrates how to export a policy, generate its manifest, and prepare the hot‑swap commands.

Export a trained checkpoint to ONNX with the normalizer baked in:

```python
from mjlab_microduck.export import run_export, ExportConfig

cfg = ExportConfig(
    onnx_file="policy.onnx",
    agent="trained",
    wandb_run_path="user/project/run_id",
    checkpoint=3000,
)
result = run_export("microduck-velocity", cfg)
print(f"Exported ONNX: {result.onnx_path}")

```

Build a manifest for a perpetual walking policy with the appropriate slot hint:

```python
from mjlab_microduck.publish.manifest import build_manifest, dump_manifest

manifest = build_manifest(
    name="walk",
    kind="perpetual",
    description="Standard walking gait",
    slot="walk",
    action_scale=1.0,
    training={"task_id": "microduck-velocity", "checkpoint": 3000},
)
print(dump_manifest(manifest))

```

Generate the daemon commands that enable hot‑swapping:

```python
from mjlab_microduck.publish.manifest import install_commands

repo_id = "user/microduck-walk"
print(install_commands(manifest, repo_id))

# Output: sudo robotctl policy load walk user/microduck-walk

```

## Summary

- **Fixed observation contract**: All policies use a 61‑D vector (48 proprioception + 13 command), ensuring shape compatibility at runtime.
- **Baked normalization**: [`src/mjlab_microduck/export.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/export.py) embeds the normalizer into the ONNX graph, allowing raw observation feeds.
- **Manifest‑driven deployment**: [`src/mjlab_microduck/publish/manifest.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/publish/manifest.py) generates schema‑2 JSON with slot hints and auto‑generated `robotctl` commands.
- **Zero‑downtime swapping**: The robot daemon validates ONNX shapes against the manifest and loads policies into named slots without service restarts.

## Frequently Asked Questions

### What happens if a policy uses a different observation size than 61 dimensions?

The robot daemon will reject the policy during manifest validation. Because Microduck RL enforces the 61‑D contract at the training configuration level ([`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py)) and the export level, shape mismatches are caught before they reach the robot.

### Why is the normalizer baked into the ONNX file instead of running on the robot?

Baking the normalizer into the exported graph eliminates runtime preprocessing dependencies and reduces latency. The robot can feed raw MJCF observations directly to the network, ensuring that training‑time and runtime statistics never diverge.

### What is the difference between perpetual and episodic policy slots?

Perpetual policies run indefinitely until explicitly unloaded, using the `robotctl policy load` command. Episodic policies execute for a fixed duration or until completion, using `robotctl policy add`. The `install_commands` helper in [`src/mjlab_microduck/publish/manifest.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/publish/manifest.py) selects the appropriate command based on the manifest's `kind` field.

### Where is the hot‑swap daemon implemented?

The daemon itself is implemented in the separate `pollen-robotics/microduck` repository. Microduck RL provides the training‑side infrastructure—observation contracts, ONNX export, and manifest generation—that makes the daemon's hot‑swap validation and loading possible.