What Is mjlab in the Microduck RL Framework? Core Engine Explained

mjlab is the core simulation-and-training engine that powers the Microduck RL repository, providing the MuJoCo Warp physics backend, parallel environment management, PPO implementation, and export utilities that all Microduck-specific robot learning code depends on.

The pollen-robotics/microduck_rl repository is built entirely on top of mjlab. This framework handles everything from high-speed physics simulation to the reinforcement learning algorithms that train policies for the Microduck quadruped robot. Understanding mjlab's role is essential for anyone extending the Microduck RL codebase or deploying trained policies to real hardware.

Physics Backend: MuJoCo Warp Simulation

mjlab wraps MuJoCo Warp—the high-performance GPU-accelerated evolution of the MuJoCo physics engine—through its mjlab.envs module.

The Microduck repository stores robot MJCF files under src/mjlab_microduck/robot/. mjlab loads these models and runs dynamics for all environments at 50 Hz, matching the control frequency required for real-world deployment.

Key implementation

In src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py, task configurations import SceneEntityCfg from mjlab to define robot entities:

from mjlab.envs import SceneEntityCfg

# Robot entity configuration using mjlab's scene system

robot_entity = SceneEntityCfg(
    "robot",
    joint_names=".*",
    body_names="base_link",
)

Without mjlab's physics abstraction, the repository would need custom MuJoCo bindings and manual GPU batching.

Parallel Environment Management

mjlab provides ManagerBasedRlEnv—a unified base class for vectorized RL environments. This enables training thousands of parallel robot instances on a single GPU.

Microduck RL task modules inherit from this base. For example, microduck_velocity_env_cfg.py builds environments that mjlab manages:

from mjlab_microduck.tasks.microduck_velocity_env_cfg import make_microduck_velocity_env_cfg

# Build a parallel environment with 4096 Microduck instances

cfg = make_microduck_velocity_env_cfg(play=False, rough=False)
env = cfg.build()  # Returns ManagerBasedRlEnv instance

obs = env.reset()  # Resets all 4096 environments simultaneously

The ManagerBasedRlEnv interface standardizes:

  • Resetting environments (env.reset())
  • Stepping physics and controllers (env.step(actions))
  • Observing states (env.obs_buf)

PPO Training Infrastructure

mjlab integrates the rsl_rl PPO implementation through mjlab.rl utilities. The training script at src/mjlab_microduck/train_cli.py delegates optimization to mjlab's algorithms.

Training invocation

uv run train Mjlab-Velocity-Flat-MicroDuck \
    --env.scene.num-envs 4096 \
    --agent.max_iterations 1000 \
    --agent.learning_rate 3e-4

This CLI forwards to mjlab's trainer while injecting Microduck-specific hooks. The underlying rsl_rl.algorithms.ppo.PPO class handles advantage estimation, policy gradients, and value function updates.

NaN-safe training patches

The repository extends mjlab's training in src/mjlab_microduck/tasks/mdp.py (lines 11–58). Here, RewardManager.compute and PPO.compute_returns are wrapped to guard against numerical instability:


# From mdp.py: patched reward computation for stability

from mjlab.managers import RewardManager

def nan_safe_compute(self, *args, **kwargs):
    rewards = self._original_compute(*args, **kwargs)
    return torch.nan_to_num(rewards, nan=0.0, posinf=1e6, neginf=-1e6)

Curriculum and Domain Randomization

mjlab supplies non-accumulating randomizers and event management through mjlab.managers.event_manager and dr.* operations. Microduck RL builds custom domain randomization on these foundations.

Custom DR implementations

The repository implements voltage sag, friction scaling, and motor parameter variations. These use mjlab's requires_model_fields decorator and event system:

from mjlab.utils import requires_model_fields

@requires_model_fields("motor_model")
def randomize_voltage_sag(env, env_ids, cfg):
    """Applied via mjlab's event manager during reset."""
    env.motor_model.voltage_scale.uniform_(*cfg.voltage_range)

Task Registration and Discovery

mjlab exposes a plug-in entry point at mjlab.tasks for automatic task discovery. The Microduck repository registers its task families in src/mjlab_microduck/tasks/__init__.py:


# Task registration for mjlab discovery

import mjlab
from .microduck_velocity_env_cfg import MicroduckVelocityEnvCfg

mjlab.tasks.register("Mjlab-Velocity-Flat-MicroDuck", MicroduckVelocityEnvCfg)

After registration, uv run list-envs displays available tasks:

  • Mjlab-Velocity-Flat-MicroDuck
  • Mjlab-Velocity-Rough-MicroDuck

This eliminates manual environment string parsing and enables type-safe configuration loading.

ONNX Export for Real Robot Deployment

mjlab provides mjlab.rl.exporter_utils for policy extraction. The Microduck repository patches this system in mdp.py (lines 68–84) to handle joint metadata correctly.

Export command

uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck \
    --wandb-run-path pollen-robotics/microduck-rl/run_id

The patch filters passive joints before baking the observation normalizer. This ensures the exported ONNX graph accepts the exact 61-dimensional observation vector used on the real Microduck robot's embedded controller.

Test Utilities and Stability Guarantees

mjlab includes NaN-guarding, reward statistics tracking, and environment sanity checks. The Microduck test suite (tests/*.py) leverages these to validate training stability across:

  • NVIDIA and AMD GPUs
  • Different CUDA versions
  • Batch size variations

Summary

  • mjlab is the foundational engine that hosts all Microduck RL functionality—not a peripheral dependency.
  • It provides MuJoCo Warp physics, parallel environment management via ManagerBasedRlEnv, and PPO training through rsl_rl integration.
  • Custom Microduck code extends mjlab through task registration, reward patches, domain randomization, and ONNX export modifications.
  • Key integration files: mdp.py (training patches), microduck_velocity_env_cfg.py (task definition), train_cli.py (CLI forwarding), and tasks/__init__.py (entry-point registration).

Frequently Asked Questions

What does mjlab stand for in the Microduck RL repository?

mjlab denotes the internal simulation and learning framework developed for Pollen Robotics' MuJoCo-based RL workflows. It is built on MuJoCo Warp (hence "mj") and provides the laboratory ("lab") infrastructure for robot learning experiments. The README explicitly states the repo is "built on mjlab (MuJoCo Warp) with PPO."

Can Microduck RL run without mjlab installed?

No. Every component—from environment stepping to policy optimization—imports from mjlab. The requirements.txt or pyproject.toml lists mjlab as a mandatory dependency. Attempting to import mjlab_microduck modules without mjlab raises ModuleNotFoundError on mjlab.envs, mjlab.managers, or mjlab.rl.

How does mjlab differ from Isaac Gym or other simulators?

mjlab specifically targets MuJoCo Warp for physics, whereas Isaac Gym uses PhysX. This choice enables exact model compatibility with MuJoCo MJCF files and deterministic simulation. The 50 Hz control frequency and motor model implementations in Microduck RL are tuned for mjlab's MuJoCo backend.

Where is the PPO implementation located—in mjlab or Microduck RL?

The PPO algorithm resides in mjlab's rsl_rl integration. Microduck RL only patches specific methods (reward computation, returns calculation) in mdp.py for numerical safety. The core policy network, value function, and PPO loss computation are imported from mjlab.rl or rsl_rl.algorithms.ppo.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →