# How IMU Orientation Is Randomized for the Microduck Robot: Two-Stage Domain Randomization Explained

> Discover how Microduck robot IMU orientation is randomized using two-stage domain randomization. Learn about per-environment misalignment and direct MuJoCo randomization for robust simulations.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-02

---

**The Microduck simulation randomizes IMU orientation through two complementary mechanisms: a per-environment constant misalignment applied to observations, and direct site-level quaternion randomization in MuJoCo at episode start.**

To train reinforcement learning policies that transfer reliably to the physical robot, the `pollen-robotics/microduck_rl` repository implements specialized domain randomization for IMU mounting errors. These techniques emulate real-world calibration offsets and sensor misalignments that occur during hardware assembly, forcing learned controllers to remain robust despite orientation uncertainty.

## Per-Environment Constant Misalignment

The first mechanism generates a **cached random rotation** that persists throughout an episode. This approach modifies how the policy perceives IMU data without altering the underlying simulation physics.

### How `_imu_misalignment_quat` Works

Located in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) (lines 18-27), this helper function creates a quaternion from a randomly sampled axis and angle:

- Samples a random 3D unit vector for the rotation axis
- Samples a random angle up to a configurable maximum (in radians)
- Stores the resulting quaternion on the environment as `env._imu_misalign_quat`

```python
def projected_gravity_imu_misaligned(env, max_angle_deg: float = 1.0):
    asset = env.scene["robot"]
    q = _imu_misalignment_quat(env, math.radians(max_angle_deg))
    return quat_apply(q, asset.data.projected_gravity_b)

```

The cached quaternion is then applied by observation functions `projected_gravity_imu_misaligned` and `base_ang_vel_imu_misaligned` to rotate the raw sensor readings before they reach the policy network. This creates the illusion of a slightly tilted IMU while the actual simulated sensor remains correctly aligned.

## Site-Level Orientation Randomization

The second mechanism **directly modifies the MuJoCo site quaternion** for the IMU sensor at episode initialization, affecting the actual physics simulation.

### The `randomize_imu_orientation` Implementation

Also in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) (lines 36-73), this function performs true geometric randomization:

```python
def randomize_imu_orientation(env, env_ids=None, max_angle_deg: float = 2.0):
    if env_ids is None:
        env_ids = torch.arange(env.num_envs, device=env.device, dtype=torch.int)
    asset = env.scene["robot"]
    site_id = 0  # IMU site index

    
    # Preserve original quaternion on first call

    if not hasattr(env, "_original_imu_quat"):
        env._original_imu_quat = env.sim.model.site_quat[0, site_id].clone()
    
    # Random rotation per environment

    n = len(env_ids)
    max_rad = max_angle_deg * torch.pi / 180.0
    angles = (torch.rand(n, 3, device=env.device) * 2 - 1) * max_rad
    
    # Small-angle quaternion approximation

    half = angles / 2.0
    delta = torch.zeros(n, 4, device=env.device)
    delta[:, 0] = 1.0
    delta[:, 1:] = half
    delta = delta / torch.norm(delta, dim=1, keepdim=True)
    
    # Apply rotation: q_new = q_delta * q_original

    original = env._original_imu_quat.unsqueeze(0).expand(n, -1)
    w1, x1, y1, z1 = delta[:, 0], delta[:, 1], delta[:, 2], delta[:, 3]
    w2, x2, y2, z2 = original[:, 0], original[:, 1], original[:, 2], original[:, 3]
    
    new_quat = torch.stack([
        w1*w2 - x1*x2 - y1*y2 - z1*z2,
        w1*x2 + x1*w2 + y1*z2 - z1*y2,
        w1*y2 - x1*z2 + y1*w2 + z1*x2,
        w1*z2 + x1*y2 - y1*x2 + z1*w2,
    ], dim=1)
    
    env.sim.model.site_quat[env_ids, site_id] = new_quat

```

This implementation uses a **small-angle quaternion approximation** for computational efficiency, then performs explicit Hamilton product multiplication to combine the random delta rotation with the original IMU site orientation.

## Key Design Decisions

**Separate caches for separate purposes**: The constant misalignment cache (`_imu_misalign_quat`) persists across steps in an episode, while the site randomization cache (`_original_imu_quat`) preserves the baseline orientation for repeated sampling.

**Different angle scales**: Typical configurations use `max_angle_deg=1.0` for observation-level noise and `max_angle_deg=2.0` for site-level randomization, reflecting that physical mounting errors are harder to correct than observation noise.

**Vectorized batching**: Both implementations operate on `env_ids` tensors, enabling efficient parallel environment execution on GPU.

## Summary

- **Observation randomization** (`_imu_misalignment_quat`): Cached per-episode quaternion that rotates gravity and angular velocity readings without physics changes
- **Site randomization** (`randomize_imu_orientation`): Direct MuJoCo site quaternion modification that alters actual sensor geometry at episode start
- **Implementation location**: Both mechanisms reside in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py)
- **Physical motivation**: Emulates realistic IMU mounting errors and calibration offsets from hardware assembly

## Frequently Asked Questions

### What is the difference between the two IMU randomization methods?

The **observation-level method** leaves the physics simulation unchanged and only rotates the sensor data the policy receives, making it computationally cheaper but less physically accurate. The **site-level method** actually modifies the MuJoCo site quaternion, so the IMU reports genuinely different physical measurements based on its new orientation in the world.

### Why not use only the site-level randomization?

Site-level changes require physics state updates and may interact with other contact dynamics or sensor placements. Observation-level randomization provides a lightweight alternative for scenarios where you want sensor noise without full geometric recomputation, and it allows finer-grained control over what the policy perceives versus what the physics engine computes.

### How do I adjust the randomization strength?

Both functions accept `max_angle_deg` parameters. In `projected_gravity_imu_misaligned` and related observation functions, this controls the cached misalignment magnitude. In `randomize_imu_orientation`, it bounds the Euler angle deviation applied to the MuJoCo site. Adjust these based on your hardware calibration confidence—typical values range from 0.5° to 5° depending on assembly precision.

### Where does the original IMU quaternion get stored?

On first invocation of `randomize_imu_orientation`, the baseline orientation is captured in `env._original_imu_quat` via `env.sim.model.site_quat[0, site_id].clone()`. This preserved quaternion serves as the reference for all subsequent randomization samples, ensuring accumulated errors don't drift unboundedly across episode resets.