How the Observation Normalizer Is Baked Into the ONNX Export for Microduck RL
The observation normalizer is automatically embedded into the ONNX graph during export through the EmpiricalNormalization layer, which is a sub-module of the policy's MLP when obs_normalization=True is configured.
Microduck RL is a reinforcement learning framework developed by Pollen Robotics for training locomotion policies. A critical challenge in deploying RL models is ensuring that observations received at runtime are processed identically to those seen during training. This article explains how Microduck RL solves this by baking the observation normalizer directly into the ONNX export, eliminating the need for separate preprocessing code in deployment.
Where the Export Logic Lives
The central export implementation is found in src/mjlab_microduck/export.py. This file orchestrates the conversion of trained checkpoints into deployment-ready ONNX models with normalization included.
The export flow begins by loading the environment configuration, which enables normalization through the obs_normalization=True flag in the task's RslRlModelCfg. You can see this configuration in any task configuration file, such as src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py (lines 916-926):
# RslRlModelCfg with observation normalization enabled
class RslRlModelCfg:
obs_normalization = True # Embeds EmpiricalNormalization into the policy
# ... other model parameters
How the Normalizer Becomes Part of the Graph
When obs_normalization=True, the EmpiricalNormalization layer is instantiated as a sub-module of the policy's MLPModel. This architectural choice is what enables automatic embedding during export.
The key export call is runner.export_policy_to_onnx(path, filename) at line 64 of export.py. The comment at lines 52-57 clarifies the mechanism:
# Lines 52-57 of export.py
# mjlab 1.3.0: ONNX export + metadata moved to mjlab.rl.exporter_utils and
# the runner's built-in export_policy_to_onnx. Observation normalization is
# baked into the exported graph automatically — EmpiricalNormalization is a
# submodule of the policy's MLPModel (obs_normalization=True in RslRlModelCfg),
# so export_policy_to_onnx emits actor(normalizer(obs)). No manual normalizer
# handling needed (the old export_velocity_policy_as_onnx path is gone).
The resulting ONNX graph structure is exactly actor(normalizer(obs)). The exporter traverses the policy's module hierarchy and includes the EmpiricalNormalization computation automatically—no manual insertion of normalization nodes is required.
Complete Export Workflow
Command-Line Export
Export a trained policy using the provided CLI script:
uv run scripts/export.py microduck_velocity --onnx_file=policy.onnx \
--wandb-run-path=user/project/run_id --checkpoint=3000
This invokes run_export in src/mjlab_microduck/export.py, which:
- Loads the checkpoint and configuration
- Calls
runner.export_policy_to_onnx()to generate the ONNX file with embedded normalizer - Attaches metadata using
attach_metadata_to_onnx(lines 66-67)
Runtime Usage
Because the normalizer is baked into the graph, deployment requires zero preprocessing:
import onnxruntime as ort
import numpy as np
sess = ort.InferenceSession("policy.onnx")
obs = np.load("sample_obs.npy") # raw observations from the robot
action = sess.run(["output"], {"obs": obs})[0]
The obs array contains raw sensor values from the robot. The ONNX model internally applies the learned mean and standard deviation from training before feeding normalized values to the policy network.
Key Files and Their Roles
| File | Purpose |
|---|---|
src/mjlab_microduck/export.py |
Core export logic; calls runner.export_policy_to_onnx which embeds the observation normalizer |
src/mjlab_microduck/tasks/*_env_cfg.py (e.g., microduck_velocity_env_cfg.py) |
Enables normalization via obs_normalization=True in RslRlModelCfg |
src/mjlab_microduck/scripts/export.py |
CLI entry point that forwards arguments to the export function |
mjlab.rl.exporter_utils (external dependency) |
Provides get_base_metadata and attach_metadata_to_onnx for post-export metadata attachment |
Design Benefits of This Approach
Training-deployment parity is guaranteed because the same EmpiricalNormalization code path executes in both contexts. There's no risk of mismatched scaling between Python training and C++ or Python runtime.
Simplified deployment eliminates the need to maintain separate normalization statistics or preprocessing implementations on the robot. The ONNX file is self-contained.
Version compatibility is handled transparently. As noted in the source comments, the old export_velocity_policy_as_onnx path has been removed in favor of this unified, automatic approach introduced in mjlab 1.3.0.
Summary
- The observation normalizer is embedded automatically when
obs_normalization=Trueis set inRslRlModelCfg - The
EmpiricalNormalizationlayer exists as a sub-module of the policy's MLP, not as external preprocessing runner.export_policy_to_onnx()traverses this module hierarchy to produce a graph with structureactor(normalizer(obs))- No manual normalizer handling is required during export or deployment
- Metadata attachment via
attach_metadata_to_onnxcompletes the deployment artifact
Frequently Asked Questions
What happens if I set obs_normalization=False in the configuration?
The EmpiricalNormalization layer is not added to the policy's MLPModel, and the exported ONNX graph contains only actor(obs) with no normalization step. You would need to implement equivalent preprocessing externally if your observations require scaling.
Can I inspect the normalizer parameters in the exported ONNX file?
Yes. The EmpiricalNormalization layer's learned running mean and variance are stored as constants in the ONNX graph. You can visualize the model with Netron or extract these values using ONNX manipulation tools if needed for debugging.
Does this approach work with quantized or optimized ONNX models?
The normalization operations are standard element-wise computations that survive typical ONNX optimization passes. If you apply post-export quantization or graph optimization, the normalizer remains part of the computation graph.
How do I update the normalizer statistics if I collect more training data?
You must resume training from your checkpoint to update the EmpiricalNormalization running statistics, then re-export to ONNX. The normalizer parameters are frozen at export time and cannot be updated in the deployed model without regeneration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →