What Is mjlab in the Microduck RL Framework? Core Engine Explained
mjlab is the core simulation-and-training engine that powers the Microduck RL repository, providing the MuJoCo Warp physics backend, parallel environment management, PPO implementation, and export utilities that all Microduck-specific robot learning code depends on.
The pollen-robotics/microduck_rl repository is built entirely on top of mjlab. This framework handles everything from high-speed physics simulation to the reinforcement learning algorithms that train policies for the Microduck quadruped robot. Understanding mjlab's role is essential for anyone extending the Microduck RL codebase or deploying trained policies to real hardware.
Physics Backend: MuJoCo Warp Simulation
mjlab wraps MuJoCo Warp—the high-performance GPU-accelerated evolution of the MuJoCo physics engine—through its mjlab.envs module.
The Microduck repository stores robot MJCF files under src/mjlab_microduck/robot/. mjlab loads these models and runs dynamics for all environments at 50 Hz, matching the control frequency required for real-world deployment.
Key implementation
In src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py, task configurations import SceneEntityCfg from mjlab to define robot entities:
from mjlab.envs import SceneEntityCfg
# Robot entity configuration using mjlab's scene system
robot_entity = SceneEntityCfg(
"robot",
joint_names=".*",
body_names="base_link",
)
Without mjlab's physics abstraction, the repository would need custom MuJoCo bindings and manual GPU batching.
Parallel Environment Management
mjlab provides ManagerBasedRlEnv—a unified base class for vectorized RL environments. This enables training thousands of parallel robot instances on a single GPU.
Microduck RL task modules inherit from this base. For example, microduck_velocity_env_cfg.py builds environments that mjlab manages:
from mjlab_microduck.tasks.microduck_velocity_env_cfg import make_microduck_velocity_env_cfg
# Build a parallel environment with 4096 Microduck instances
cfg = make_microduck_velocity_env_cfg(play=False, rough=False)
env = cfg.build() # Returns ManagerBasedRlEnv instance
obs = env.reset() # Resets all 4096 environments simultaneously
The ManagerBasedRlEnv interface standardizes:
- Resetting environments (
env.reset()) - Stepping physics and controllers (
env.step(actions)) - Observing states (
env.obs_buf)
PPO Training Infrastructure
mjlab integrates the rsl_rl PPO implementation through mjlab.rl utilities. The training script at src/mjlab_microduck/train_cli.py delegates optimization to mjlab's algorithms.
Training invocation
uv run train Mjlab-Velocity-Flat-MicroDuck \
--env.scene.num-envs 4096 \
--agent.max_iterations 1000 \
--agent.learning_rate 3e-4
This CLI forwards to mjlab's trainer while injecting Microduck-specific hooks. The underlying rsl_rl.algorithms.ppo.PPO class handles advantage estimation, policy gradients, and value function updates.
NaN-safe training patches
The repository extends mjlab's training in src/mjlab_microduck/tasks/mdp.py (lines 11–58). Here, RewardManager.compute and PPO.compute_returns are wrapped to guard against numerical instability:
# From mdp.py: patched reward computation for stability
from mjlab.managers import RewardManager
def nan_safe_compute(self, *args, **kwargs):
rewards = self._original_compute(*args, **kwargs)
return torch.nan_to_num(rewards, nan=0.0, posinf=1e6, neginf=-1e6)
Curriculum and Domain Randomization
mjlab supplies non-accumulating randomizers and event management through mjlab.managers.event_manager and dr.* operations. Microduck RL builds custom domain randomization on these foundations.
Custom DR implementations
The repository implements voltage sag, friction scaling, and motor parameter variations. These use mjlab's requires_model_fields decorator and event system:
from mjlab.utils import requires_model_fields
@requires_model_fields("motor_model")
def randomize_voltage_sag(env, env_ids, cfg):
"""Applied via mjlab's event manager during reset."""
env.motor_model.voltage_scale.uniform_(*cfg.voltage_range)
Task Registration and Discovery
mjlab exposes a plug-in entry point at mjlab.tasks for automatic task discovery. The Microduck repository registers its task families in src/mjlab_microduck/tasks/__init__.py:
# Task registration for mjlab discovery
import mjlab
from .microduck_velocity_env_cfg import MicroduckVelocityEnvCfg
mjlab.tasks.register("Mjlab-Velocity-Flat-MicroDuck", MicroduckVelocityEnvCfg)
After registration, uv run list-envs displays available tasks:
Mjlab-Velocity-Flat-MicroDuckMjlab-Velocity-Rough-MicroDuck
This eliminates manual environment string parsing and enables type-safe configuration loading.
ONNX Export for Real Robot Deployment
mjlab provides mjlab.rl.exporter_utils for policy extraction. The Microduck repository patches this system in mdp.py (lines 68–84) to handle joint metadata correctly.
Export command
uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck \
--wandb-run-path pollen-robotics/microduck-rl/run_id
The patch filters passive joints before baking the observation normalizer. This ensures the exported ONNX graph accepts the exact 61-dimensional observation vector used on the real Microduck robot's embedded controller.
Test Utilities and Stability Guarantees
mjlab includes NaN-guarding, reward statistics tracking, and environment sanity checks. The Microduck test suite (tests/*.py) leverages these to validate training stability across:
- NVIDIA and AMD GPUs
- Different CUDA versions
- Batch size variations
Summary
- mjlab is the foundational engine that hosts all Microduck RL functionality—not a peripheral dependency.
- It provides MuJoCo Warp physics, parallel environment management via
ManagerBasedRlEnv, and PPO training through rsl_rl integration. - Custom Microduck code extends mjlab through task registration, reward patches, domain randomization, and ONNX export modifications.
- Key integration files:
mdp.py(training patches),microduck_velocity_env_cfg.py(task definition),train_cli.py(CLI forwarding), andtasks/__init__.py(entry-point registration).
Frequently Asked Questions
What does mjlab stand for in the Microduck RL repository?
mjlab denotes the internal simulation and learning framework developed for Pollen Robotics' MuJoCo-based RL workflows. It is built on MuJoCo Warp (hence "mj") and provides the laboratory ("lab") infrastructure for robot learning experiments. The README explicitly states the repo is "built on mjlab (MuJoCo Warp) with PPO."
Can Microduck RL run without mjlab installed?
No. Every component—from environment stepping to policy optimization—imports from mjlab. The requirements.txt or pyproject.toml lists mjlab as a mandatory dependency. Attempting to import mjlab_microduck modules without mjlab raises ModuleNotFoundError on mjlab.envs, mjlab.managers, or mjlab.rl.
How does mjlab differ from Isaac Gym or other simulators?
mjlab specifically targets MuJoCo Warp for physics, whereas Isaac Gym uses PhysX. This choice enables exact model compatibility with MuJoCo MJCF files and deterministic simulation. The 50 Hz control frequency and motor model implementations in Microduck RL are tuned for mjlab's MuJoCo backend.
Where is the PPO implementation located—in mjlab or Microduck RL?
The PPO algorithm resides in mjlab's rsl_rl integration. Microduck RL only patches specific methods (reward computation, returns calculation) in mdp.py for numerical safety. The core policy network, value function, and PPO loss computation are imported from mjlab.rl or rsl_rl.algorithms.ppo.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →