# What Is mjlab in the Microduck RL Framework? Core Engine Explained

> Discover mjlab, the core engine of Microduck RL. Learn how it powers robot learning with its MuJoCo backend, parallel environments, and PPO implementation for efficient training.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-02

---

**mjlab is the core simulation-and-training engine that powers the Microduck RL repository, providing the MuJoCo Warp physics backend, parallel environment management, PPO implementation, and export utilities that all Microduck-specific robot learning code depends on.**

The `pollen-robotics/microduck_rl` repository is built entirely on top of mjlab. This framework handles everything from high-speed physics simulation to the reinforcement learning algorithms that train policies for the Microduck quadruped robot. Understanding mjlab's role is essential for anyone extending the Microduck RL codebase or deploying trained policies to real hardware.

## Physics Backend: MuJoCo Warp Simulation

mjlab wraps **MuJoCo Warp**—the high-performance GPU-accelerated evolution of the MuJoCo physics engine—through its `mjlab.envs` module.

The Microduck repository stores robot MJCF files under `src/mjlab_microduck/robot/`. mjlab loads these models and runs dynamics for all environments at 50 Hz, matching the control frequency required for real-world deployment.

### Key implementation

In [`src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py), task configurations import `SceneEntityCfg` from mjlab to define robot entities:

```python
from mjlab.envs import SceneEntityCfg

# Robot entity configuration using mjlab's scene system

robot_entity = SceneEntityCfg(
    "robot",
    joint_names=".*",
    body_names="base_link",
)

```

Without mjlab's physics abstraction, the repository would need custom MuJoCo bindings and manual GPU batching.

## Parallel Environment Management

mjlab provides `ManagerBasedRlEnv`—a unified base class for vectorized RL environments. This enables training thousands of parallel robot instances on a single GPU.

Microduck RL task modules inherit from this base. For example, [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py) builds environments that mjlab manages:

```python
from mjlab_microduck.tasks.microduck_velocity_env_cfg import make_microduck_velocity_env_cfg

# Build a parallel environment with 4096 Microduck instances

cfg = make_microduck_velocity_env_cfg(play=False, rough=False)
env = cfg.build()  # Returns ManagerBasedRlEnv instance

obs = env.reset()  # Resets all 4096 environments simultaneously

```

The `ManagerBasedRlEnv` interface standardizes:
- **Resetting** environments (`env.reset()`)
- **Stepping** physics and controllers (`env.step(actions)`)
- **Observing** states (`env.obs_buf`)

## PPO Training Infrastructure

mjlab integrates the **rsl_rl** PPO implementation through `mjlab.rl` utilities. The training script at [`src/mjlab_microduck/train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/train_cli.py) delegates optimization to mjlab's algorithms.

### Training invocation

```bash
uv run train Mjlab-Velocity-Flat-MicroDuck \
    --env.scene.num-envs 4096 \
    --agent.max_iterations 1000 \
    --agent.learning_rate 3e-4

```

This CLI forwards to mjlab's trainer while injecting Microduck-specific hooks. The underlying `rsl_rl.algorithms.ppo.PPO` class handles advantage estimation, policy gradients, and value function updates.

### NaN-safe training patches

The repository extends mjlab's training in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) (lines 11–58). Here, `RewardManager.compute` and `PPO.compute_returns` are wrapped to guard against numerical instability:

```python

# From mdp.py: patched reward computation for stability

from mjlab.managers import RewardManager

def nan_safe_compute(self, *args, **kwargs):
    rewards = self._original_compute(*args, **kwargs)
    return torch.nan_to_num(rewards, nan=0.0, posinf=1e6, neginf=-1e6)

```

## Curriculum and Domain Randomization

mjlab supplies **non-accumulating randomizers** and event management through `mjlab.managers.event_manager` and `dr.*` operations. Microduck RL builds custom domain randomization on these foundations.

### Custom DR implementations

The repository implements voltage sag, friction scaling, and motor parameter variations. These use mjlab's `requires_model_fields` decorator and event system:

```python
from mjlab.utils import requires_model_fields

@requires_model_fields("motor_model")
def randomize_voltage_sag(env, env_ids, cfg):
    """Applied via mjlab's event manager during reset."""
    env.motor_model.voltage_scale.uniform_(*cfg.voltage_range)

```

## Task Registration and Discovery

mjlab exposes a **plug-in entry point** at `mjlab.tasks` for automatic task discovery. The Microduck repository registers its task families in [`src/mjlab_microduck/tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/__init__.py):

```python

# Task registration for mjlab discovery

import mjlab
from .microduck_velocity_env_cfg import MicroduckVelocityEnvCfg

mjlab.tasks.register("Mjlab-Velocity-Flat-MicroDuck", MicroduckVelocityEnvCfg)

```

After registration, `uv run list-envs` displays available tasks:
- `Mjlab-Velocity-Flat-MicroDuck`
- `Mjlab-Velocity-Rough-MicroDuck`

This eliminates manual environment string parsing and enables type-safe configuration loading.

## ONNX Export for Real Robot Deployment

mjlab provides `mjlab.rl.exporter_utils` for policy extraction. The Microduck repository patches this system in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) (lines 68–84) to handle joint metadata correctly.

### Export command

```bash
uv run scripts/export.py Mjlab-Velocity-Flat-MicroDuck \
    --wandb-run-path pollen-robotics/microduck-rl/run_id

```

The patch filters **passive joints** before baking the observation normalizer. This ensures the exported ONNX graph accepts the exact 61-dimensional observation vector used on the real Microduck robot's embedded controller.

## Test Utilities and Stability Guarantees

mjlab includes **NaN-guarding**, reward statistics tracking, and environment sanity checks. The Microduck test suite (`tests/*.py`) leverages these to validate training stability across:
- NVIDIA and AMD GPUs
- Different CUDA versions
- Batch size variations

## Summary

- **mjlab** is the foundational engine that hosts all Microduck RL functionality—not a peripheral dependency.
- It provides **MuJoCo Warp physics**, **parallel environment management** via `ManagerBasedRlEnv`, and **PPO training** through rsl_rl integration.
- Custom Microduck code extends mjlab through **task registration**, **reward patches**, **domain randomization**, and **ONNX export modifications**.
- Key integration files: [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) (training patches), [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py) (task definition), [`train_cli.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/train_cli.py) (CLI forwarding), and [`tasks/__init__.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tasks/__init__.py) (entry-point registration).

## Frequently Asked Questions

### What does mjlab stand for in the Microduck RL repository?

mjlab denotes the internal simulation and learning framework developed for Pollen Robotics' MuJoCo-based RL workflows. It is built on **MuJoCo Warp** (hence "mj") and provides the laboratory ("lab") infrastructure for robot learning experiments. The README explicitly states the repo is "built on mjlab (MuJoCo Warp) with PPO."

### Can Microduck RL run without mjlab installed?

No. Every component—from environment stepping to policy optimization—imports from mjlab. The [`requirements.txt`](https://github.com/pollen-robotics/microduck_rl/blob/main/requirements.txt) or [`pyproject.toml`](https://github.com/pollen-robotics/microduck_rl/blob/main/pyproject.toml) lists mjlab as a mandatory dependency. Attempting to import `mjlab_microduck` modules without mjlab raises `ModuleNotFoundError` on `mjlab.envs`, `mjlab.managers`, or `mjlab.rl`.

### How does mjlab differ from Isaac Gym or other simulators?

mjlab specifically targets **MuJoCo Warp** for physics, whereas Isaac Gym uses PhysX. This choice enables exact model compatibility with MuJoCo MJCF files and deterministic simulation. The 50 Hz control frequency and motor model implementations in Microduck RL are tuned for mjlab's MuJoCo backend.

### Where is the PPO implementation located—in mjlab or Microduck RL?

The PPO algorithm resides in **mjlab's** `rsl_rl` integration. Microduck RL only patches specific methods (reward computation, returns calculation) in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) for numerical safety. The core policy network, value function, and PPO loss computation are imported from `mjlab.rl` or `rsl_rl.algorithms.ppo`.