How to Deploy Policies on the Real Microduck Robot: A Complete Pipeline Guide

Policies are deployed to the real Microduck robot through a three-stage pipeline that exports the trained network to ONNX, publishes a manifest to the Hugging Face Hub, and executes generated robotctl commands on the target hardware.

Deploying reinforcement learning policies from simulation to physical hardware requires rigorous tooling. In the pollen-robotics/microduck_rl repository, the path from trained model to real robot execution follows a structured workflow involving ONNX serialization, JSON manifests, and the robotctl daemon. This guide explains exactly how policies are deployed on the real Microduck robot using the source implementation.

Understanding the Policy Deployment Pipeline

The deployment system consists of three coordinated stages that bridge machine learning training and embedded execution. Each stage is implemented in specific source files that handle the transition from Python training code to on-device inference.

Step 1: ONNX Export and Manifest Generation

Before deploying policies on the real Microduck robot, the trained network must be converted into a hardware-agnostic format. The src/mjlab_microduck/export.py module performs ONNX export while baking the observation normalizer directly into the network graph to eliminate runtime preprocessing overhead.

The scripts/export.py CLI wrapper orchestrates this process, creating two critical artifacts in your output directory:

  • policy.onnx: The serialized neural network with fused normalizers
  • manifest.json: Runtime metadata describing policy shape, type, and execution parameters

The manifest is constructed by build_manifest and validated by validate_manifest, both defined in src/mjlab_microduck/publish/manifest.py. Key configuration fields include:

  • kind: The execution type (episodic, perpetual, or held)
  • slot: The target actuator slot for perpetual gaits (e.g., walk, stand)
  • duration_s: Runtime limit in seconds for episodic skills
from src.mjlab_microduck.publish.manifest import build_manifest, render_readme

manifest = build_manifest(
    name="walk",
    kind="perpetual",
    description="A walking gait for Microduck",
    slot="walk",
    duration_s=None,  # perpetual policies omit duration

)
repo_id = "user/microduck-walk"
print(render_readme(manifest, repo_id))

The render_readme function automatically generates deployment documentation that stays synchronized with the manifest, ensuring users always have the correct robotctl syntax.

Step 2: Publishing to the Hugging Face Hub

Once exported, the policy package is pushed to the Hugging Face Hub using the uv run publish command. This CLI tool is implemented in src/mjlab_microduck/publish/cli.py and handles authentication, repository creation, and artifact upload.

The command pushes both the ONNX file and the JSON manifest to a remote repository (e.g., user/microduck-walk). The robotctl daemon running on the physical Microduck robot can subsequently pull these artifacts directly from the hub using standard Git LFS protocols.

Step 3: Loading and Executing Policies on the Robot

The robotctl daemon handles the final stage of deploying policies on the real Microduck robot. Based on the manifest's kind field, the install_commands function in src/mjlab_microduck/publish/manifest.py generates specific shell commands for each execution mode.

Episodic Skills

For one-shot behaviors that terminate automatically after a fixed duration, use the episodic kind. The generated commands install the policy temporarily and execute it:

sudo robotctl policy add <name> <repo>
robotctl robot do <name>

The skill runs for exactly duration_s seconds as specified in the manifest, then automatically unwinds and returns control to the default controller.

Perpetual Gaits

Continuous behaviors like walking use the perpetual kind and load into fixed actuator slots. The install_commands function generates:

sudo robotctl policy load <slot> <repo>

Valid slots include walk, stand, and other predefined gait categories defined in the robot's configuration. The gait persists until explicitly replaced by another policy or the robot powers down.

Held Poses

For static configurations that must remain active indefinitely until explicitly released, use the held kind. The generated command includes the --hold flag:

sudo robotctl policy add <name> <repo> --hold <seconds>

This loads a pose that stays locked until the control system issues a release command or the specified timeout expires.

Robot-Side Validation and Execution

When robotctl receives a deployment command to load a policy on the real Microduck robot, it performs several validation steps before issuing motor commands:

  1. Pulls the policy repository from the Hugging Face Hub to local storage
  2. Validates the ONNX input/output tensor shapes (typically 61 → 14) using the check_onnx utility
  3. Confirms the observation normalizer is properly baked into the network graph
  4. Registers the policy under the requested slot or name in the daemon's registry
  5. Executes the policy according to the manifest's timing rules (duration_s, unwind_s, or perpetual)

Summary

  • ONNX export converts trained policies into hardware-agnostic format using src/mjlab_microduck/export.py, embedding observation normalizers directly into the network
  • Manifest generation defines runtime behavior through build_manifest in src/mjlab_microduck/publish/manifest.py, specifying execution kind and timing parameters
  • Remote publishing pushes artifacts via uv run publish implemented in src/mjlab_microduck/publish/cli.py for version-controlled distribution
  • Robot loading uses auto-generated robotctl commands from install_commands to execute episodic, perpetual, or held behaviors on the physical hardware
  • Runtime validation ensures ONNX shape compatibility (61 → 14) and normalizer presence before real robot execution

Frequently Asked Questions

What file format does Microduck use for policy deployment?

Microduck deployments use ONNX for the neural network weights and a JSON manifest for runtime metadata. This combination ensures cross-platform compatibility and self-describing parameters that the robotctl daemon can interpret without additional configuration files.

How does the robot validate downloaded policies before execution?

The robotctl daemon validates policies by checking ONNX tensor shapes against expected dimensions (typically 61 inputs to 14 outputs) and verifying that observation normalizers are baked into the network graph. This validation occurs in check_onnx before any motor commands are issued, preventing runtime shape mismatches.

What distinguishes a perpetual gait from an episodic skill?

Perpetual gaits load into persistent slots like walk or stand using robotctl policy load and run indefinitely until replaced. Episodic skills use robotctl policy add followed by robotctl robot do, executing for a specific duration_s before automatically terminating and returning to the default controller.

Where is the deployment command generation tested?

Unit tests for the install_commands function and manifest schema validation reside in tests/test_publish_manifest.py. These tests confirm that generated robotctl syntax matches expected patterns for all three policy kinds (episodic, perpetual, and held) and that the render_readme output contains valid bash commands.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →