How to Deploy Policies on the Real Microduck Robot: A Complete Pipeline Guide
Policies are deployed to the real Microduck robot through a three-stage pipeline that exports the trained network to ONNX, publishes a manifest to the Hugging Face Hub, and executes generated robotctl commands on the target hardware.
Deploying reinforcement learning policies from simulation to physical hardware requires rigorous tooling. In the pollen-robotics/microduck_rl repository, the path from trained model to real robot execution follows a structured workflow involving ONNX serialization, JSON manifests, and the robotctl daemon. This guide explains exactly how policies are deployed on the real Microduck robot using the source implementation.
Understanding the Policy Deployment Pipeline
The deployment system consists of three coordinated stages that bridge machine learning training and embedded execution. Each stage is implemented in specific source files that handle the transition from Python training code to on-device inference.
- ONNX export and manifest generation handled in
src/mjlab_microduck/export.py - Publishing to remote repositories managed by
src/mjlab_microduck/publish/cli.py - Robot-side loading and execution controlled through
robotctlcommands generated bysrc/mjlab_microduck/publish/manifest.py
Step 1: ONNX Export and Manifest Generation
Before deploying policies on the real Microduck robot, the trained network must be converted into a hardware-agnostic format. The src/mjlab_microduck/export.py module performs ONNX export while baking the observation normalizer directly into the network graph to eliminate runtime preprocessing overhead.
The scripts/export.py CLI wrapper orchestrates this process, creating two critical artifacts in your output directory:
policy.onnx: The serialized neural network with fused normalizersmanifest.json: Runtime metadata describing policy shape, type, and execution parameters
The manifest is constructed by build_manifest and validated by validate_manifest, both defined in src/mjlab_microduck/publish/manifest.py. Key configuration fields include:
kind: The execution type (episodic,perpetual, orheld)slot: The target actuator slot for perpetual gaits (e.g.,walk,stand)duration_s: Runtime limit in seconds for episodic skills
from src.mjlab_microduck.publish.manifest import build_manifest, render_readme
manifest = build_manifest(
name="walk",
kind="perpetual",
description="A walking gait for Microduck",
slot="walk",
duration_s=None, # perpetual policies omit duration
)
repo_id = "user/microduck-walk"
print(render_readme(manifest, repo_id))
The render_readme function automatically generates deployment documentation that stays synchronized with the manifest, ensuring users always have the correct robotctl syntax.
Step 2: Publishing to the Hugging Face Hub
Once exported, the policy package is pushed to the Hugging Face Hub using the uv run publish command. This CLI tool is implemented in src/mjlab_microduck/publish/cli.py and handles authentication, repository creation, and artifact upload.
The command pushes both the ONNX file and the JSON manifest to a remote repository (e.g., user/microduck-walk). The robotctl daemon running on the physical Microduck robot can subsequently pull these artifacts directly from the hub using standard Git LFS protocols.
Step 3: Loading and Executing Policies on the Robot
The robotctl daemon handles the final stage of deploying policies on the real Microduck robot. Based on the manifest's kind field, the install_commands function in src/mjlab_microduck/publish/manifest.py generates specific shell commands for each execution mode.
Episodic Skills
For one-shot behaviors that terminate automatically after a fixed duration, use the episodic kind. The generated commands install the policy temporarily and execute it:
sudo robotctl policy add <name> <repo>
robotctl robot do <name>
The skill runs for exactly duration_s seconds as specified in the manifest, then automatically unwinds and returns control to the default controller.
Perpetual Gaits
Continuous behaviors like walking use the perpetual kind and load into fixed actuator slots. The install_commands function generates:
sudo robotctl policy load <slot> <repo>
Valid slots include walk, stand, and other predefined gait categories defined in the robot's configuration. The gait persists until explicitly replaced by another policy or the robot powers down.
Held Poses
For static configurations that must remain active indefinitely until explicitly released, use the held kind. The generated command includes the --hold flag:
sudo robotctl policy add <name> <repo> --hold <seconds>
This loads a pose that stays locked until the control system issues a release command or the specified timeout expires.
Robot-Side Validation and Execution
When robotctl receives a deployment command to load a policy on the real Microduck robot, it performs several validation steps before issuing motor commands:
- Pulls the policy repository from the Hugging Face Hub to local storage
- Validates the ONNX input/output tensor shapes (typically
61 → 14) using thecheck_onnxutility - Confirms the observation normalizer is properly baked into the network graph
- Registers the policy under the requested slot or name in the daemon's registry
- Executes the policy according to the manifest's timing rules (
duration_s,unwind_s, or perpetual)
Summary
- ONNX export converts trained policies into hardware-agnostic format using
src/mjlab_microduck/export.py, embedding observation normalizers directly into the network - Manifest generation defines runtime behavior through
build_manifestinsrc/mjlab_microduck/publish/manifest.py, specifying execution kind and timing parameters - Remote publishing pushes artifacts via
uv run publishimplemented insrc/mjlab_microduck/publish/cli.pyfor version-controlled distribution - Robot loading uses auto-generated
robotctlcommands frominstall_commandsto execute episodic, perpetual, or held behaviors on the physical hardware - Runtime validation ensures ONNX shape compatibility (
61 → 14) and normalizer presence before real robot execution
Frequently Asked Questions
What file format does Microduck use for policy deployment?
Microduck deployments use ONNX for the neural network weights and a JSON manifest for runtime metadata. This combination ensures cross-platform compatibility and self-describing parameters that the robotctl daemon can interpret without additional configuration files.
How does the robot validate downloaded policies before execution?
The robotctl daemon validates policies by checking ONNX tensor shapes against expected dimensions (typically 61 inputs to 14 outputs) and verifying that observation normalizers are baked into the network graph. This validation occurs in check_onnx before any motor commands are issued, preventing runtime shape mismatches.
What distinguishes a perpetual gait from an episodic skill?
Perpetual gaits load into persistent slots like walk or stand using robotctl policy load and run indefinitely until replaced. Episodic skills use robotctl policy add followed by robotctl robot do, executing for a specific duration_s before automatically terminating and returning to the default controller.
Where is the deployment command generation tested?
Unit tests for the install_commands function and manifest schema validation reside in tests/test_publish_manifest.py. These tests confirm that generated robotctl syntax matches expected patterns for all three policy kinds (episodic, perpetual, and held) and that the render_readme output contains valid bash commands.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →