How Runtime Policy Hot‑Swapping Is Enabled in Microduck RL
Microduck RL enables runtime policy hot‑swapping by enforcing a single, immutable 61‑dimensional observation contract across all training environments and publishing a JSON manifest that directs the robot daemon to load ONNX policies via automated robotctl commands.
Runtime policy hot‑swapping allows physical robots to switch between reinforcement learning policies without restarting services or redeploying firmware. In the pollen-robotics/microduck_rl repository, this capability is built on a strict observation schema, baked‑in normalization, and a declarative manifest system that the runtime daemon consumes to validate and activate new policies instantly. The architecture is documented in AGENTS.md, which describes the constant‑command family and the importance of a unified observation contract.
The Foundation: A Unified Observation Contract
All policies in Microduck RL must conform to an identical observation layout. This immutability guarantee ensures that any trained network can be substituted for another at runtime without shape mismatches or preprocessing errors.
The 61‑Dimensional Observation Layout
Every training environment emits a fixed 61‑dimensional observation vector comprising 48 proprioception values plus 13 command dimensions. This layout is hard‑coded in task configuration modules such as src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py and enforced by the export pipeline.
Because the observation length never varies, the robot daemon can allocate fixed buffers and validation logic once, then accept any policy trained with this schema. The corresponding action space is always 14 dimensions, matching the robot's actuator count.
ONNX Export with Baked Normalization
The normalization layer is eliminated from runtime overhead by baking it directly into the exported neural network.
In src/mjlab_microduck/export.py, the run_export function serializes actor(normalizer(obs)) as a single ONNX graph. This means the robot feeds raw MJCF observations directly into the network without additional preprocessing, reducing latency and eliminating version skew between training statistics and runtime normalization.
The Policy Manifest System
Microduck RL uses a declarative manifest to bridge the training framework and the robot daemon. The CLI wrapper in src/mjlab_microduck/publish/cli.py orchestrates the full publish flow, calling the manifest builder and export utilities.
Schema 2 and Slot Hints
src/mjlab_microduck/publish/manifest.py generates a JSON manifest (schema version 2) containing:
obs_len: Fixed at61action_len: Fixed at14- Slot hints: Categorical labels such as
walk,stand, orsitstandthat map the policy to a runtime slot in the daemon
These slot hints determine which behavioral mode the policy replaces when hot‑swapped.
Auto‑Generated robotctl Commands
The manifest builder automatically generates the appropriate daemon commands via install_commands. For perpetual policies that run indefinitely, it produces:
sudo robotctl policy load <slot> <repo_id>
For episodic policies that run for fixed durations, it generates:
sudo robotctl policy add <slot> <repo_id>
The helper logic inspects the kind field in the manifest to select the correct command variant, ensuring the daemon receives the proper instruction without manual scripting.
Runtime Daemon Integration
While the daemon itself lives in the separate pollen-robotics/microduck repository, Microduck RL produces all artifacts required for its hot‑swap protocol. When a new policy is published:
- The daemon reads the manifest and validates that the ONNX input shape matches
obs_len(61) - It registers the policy under the indicated slot (e.g.,
walk) - It executes the generated
robotctlcommand to unload the current policy and activate the new one
Because all policies share the same observation shape and normalization is baked into the ONNX file, the transition occurs without restarting the robot or recompiling code.
Practical Implementation
The following workflow demonstrates how to export a policy, generate its manifest, and prepare the hot‑swap commands.
Export a trained checkpoint to ONNX with the normalizer baked in:
from mjlab_microduck.export import run_export, ExportConfig
cfg = ExportConfig(
onnx_file="policy.onnx",
agent="trained",
wandb_run_path="user/project/run_id",
checkpoint=3000,
)
result = run_export("microduck-velocity", cfg)
print(f"Exported ONNX: {result.onnx_path}")
Build a manifest for a perpetual walking policy with the appropriate slot hint:
from mjlab_microduck.publish.manifest import build_manifest, dump_manifest
manifest = build_manifest(
name="walk",
kind="perpetual",
description="Standard walking gait",
slot="walk",
action_scale=1.0,
training={"task_id": "microduck-velocity", "checkpoint": 3000},
)
print(dump_manifest(manifest))
Generate the daemon commands that enable hot‑swapping:
from mjlab_microduck.publish.manifest import install_commands
repo_id = "user/microduck-walk"
print(install_commands(manifest, repo_id))
# Output: sudo robotctl policy load walk user/microduck-walk
Summary
- Fixed observation contract: All policies use a 61‑D vector (48 proprioception + 13 command), ensuring shape compatibility at runtime.
- Baked normalization:
src/mjlab_microduck/export.pyembeds the normalizer into the ONNX graph, allowing raw observation feeds. - Manifest‑driven deployment:
src/mjlab_microduck/publish/manifest.pygenerates schema‑2 JSON with slot hints and auto‑generatedrobotctlcommands. - Zero‑downtime swapping: The robot daemon validates ONNX shapes against the manifest and loads policies into named slots without service restarts.
Frequently Asked Questions
What happens if a policy uses a different observation size than 61 dimensions?
The robot daemon will reject the policy during manifest validation. Because Microduck RL enforces the 61‑D contract at the training configuration level (src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) and the export level, shape mismatches are caught before they reach the robot.
Why is the normalizer baked into the ONNX file instead of running on the robot?
Baking the normalizer into the exported graph eliminates runtime preprocessing dependencies and reduces latency. The robot can feed raw MJCF observations directly to the network, ensuring that training‑time and runtime statistics never diverge.
What is the difference between perpetual and episodic policy slots?
Perpetual policies run indefinitely until explicitly unloaded, using the robotctl policy load command. Episodic policies execute for a fixed duration or until completion, using robotctl policy add. The install_commands helper in src/mjlab_microduck/publish/manifest.py selects the appropriate command based on the manifest's kind field.
Where is the hot‑swap daemon implemented?
The daemon itself is implemented in the separate pollen-robotics/microduck repository. Microduck RL provides the training‑side infrastructure—observation contracts, ONNX export, and manifest generation—that makes the daemon's hot‑swap validation and loading possible.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →