How IMU Orientation Is Randomized for the Microduck Robot: Two-Stage Domain Randomization Explained

The Microduck simulation randomizes IMU orientation through two complementary mechanisms: a per-environment constant misalignment applied to observations, and direct site-level quaternion randomization in MuJoCo at episode start.

To train reinforcement learning policies that transfer reliably to the physical robot, the pollen-robotics/microduck_rl repository implements specialized domain randomization for IMU mounting errors. These techniques emulate real-world calibration offsets and sensor misalignments that occur during hardware assembly, forcing learned controllers to remain robust despite orientation uncertainty.

Per-Environment Constant Misalignment

The first mechanism generates a cached random rotation that persists throughout an episode. This approach modifies how the policy perceives IMU data without altering the underlying simulation physics.

How _imu_misalignment_quat Works

Located in src/mjlab_microduck/tasks/mdp.py (lines 18-27), this helper function creates a quaternion from a randomly sampled axis and angle:

  • Samples a random 3D unit vector for the rotation axis
  • Samples a random angle up to a configurable maximum (in radians)
  • Stores the resulting quaternion on the environment as env._imu_misalign_quat
def projected_gravity_imu_misaligned(env, max_angle_deg: float = 1.0):
    asset = env.scene["robot"]
    q = _imu_misalignment_quat(env, math.radians(max_angle_deg))
    return quat_apply(q, asset.data.projected_gravity_b)

The cached quaternion is then applied by observation functions projected_gravity_imu_misaligned and base_ang_vel_imu_misaligned to rotate the raw sensor readings before they reach the policy network. This creates the illusion of a slightly tilted IMU while the actual simulated sensor remains correctly aligned.

Site-Level Orientation Randomization

The second mechanism directly modifies the MuJoCo site quaternion for the IMU sensor at episode initialization, affecting the actual physics simulation.

The randomize_imu_orientation Implementation

Also in src/mjlab_microduck/tasks/mdp.py (lines 36-73), this function performs true geometric randomization:

def randomize_imu_orientation(env, env_ids=None, max_angle_deg: float = 2.0):
    if env_ids is None:
        env_ids = torch.arange(env.num_envs, device=env.device, dtype=torch.int)
    asset = env.scene["robot"]
    site_id = 0  # IMU site index

    
    # Preserve original quaternion on first call

    if not hasattr(env, "_original_imu_quat"):
        env._original_imu_quat = env.sim.model.site_quat[0, site_id].clone()
    
    # Random rotation per environment

    n = len(env_ids)
    max_rad = max_angle_deg * torch.pi / 180.0
    angles = (torch.rand(n, 3, device=env.device) * 2 - 1) * max_rad
    
    # Small-angle quaternion approximation

    half = angles / 2.0
    delta = torch.zeros(n, 4, device=env.device)
    delta[:, 0] = 1.0
    delta[:, 1:] = half
    delta = delta / torch.norm(delta, dim=1, keepdim=True)
    
    # Apply rotation: q_new = q_delta * q_original

    original = env._original_imu_quat.unsqueeze(0).expand(n, -1)
    w1, x1, y1, z1 = delta[:, 0], delta[:, 1], delta[:, 2], delta[:, 3]
    w2, x2, y2, z2 = original[:, 0], original[:, 1], original[:, 2], original[:, 3]
    
    new_quat = torch.stack([
        w1*w2 - x1*x2 - y1*y2 - z1*z2,
        w1*x2 + x1*w2 + y1*z2 - z1*y2,
        w1*y2 - x1*z2 + y1*w2 + z1*x2,
        w1*z2 + x1*y2 - y1*x2 + z1*w2,
    ], dim=1)
    
    env.sim.model.site_quat[env_ids, site_id] = new_quat

This implementation uses a small-angle quaternion approximation for computational efficiency, then performs explicit Hamilton product multiplication to combine the random delta rotation with the original IMU site orientation.

Key Design Decisions

Separate caches for separate purposes: The constant misalignment cache (_imu_misalign_quat) persists across steps in an episode, while the site randomization cache (_original_imu_quat) preserves the baseline orientation for repeated sampling.

Different angle scales: Typical configurations use max_angle_deg=1.0 for observation-level noise and max_angle_deg=2.0 for site-level randomization, reflecting that physical mounting errors are harder to correct than observation noise.

Vectorized batching: Both implementations operate on env_ids tensors, enabling efficient parallel environment execution on GPU.

Summary

  • Observation randomization (_imu_misalignment_quat): Cached per-episode quaternion that rotates gravity and angular velocity readings without physics changes
  • Site randomization (randomize_imu_orientation): Direct MuJoCo site quaternion modification that alters actual sensor geometry at episode start
  • Implementation location: Both mechanisms reside in src/mjlab_microduck/tasks/mdp.py
  • Physical motivation: Emulates realistic IMU mounting errors and calibration offsets from hardware assembly

Frequently Asked Questions

What is the difference between the two IMU randomization methods?

The observation-level method leaves the physics simulation unchanged and only rotates the sensor data the policy receives, making it computationally cheaper but less physically accurate. The site-level method actually modifies the MuJoCo site quaternion, so the IMU reports genuinely different physical measurements based on its new orientation in the world.

Why not use only the site-level randomization?

Site-level changes require physics state updates and may interact with other contact dynamics or sensor placements. Observation-level randomization provides a lightweight alternative for scenarios where you want sensor noise without full geometric recomputation, and it allows finer-grained control over what the policy perceives versus what the physics engine computes.

How do I adjust the randomization strength?

Both functions accept max_angle_deg parameters. In projected_gravity_imu_misaligned and related observation functions, this controls the cached misalignment magnitude. In randomize_imu_orientation, it bounds the Euler angle deviation applied to the MuJoCo site. Adjust these based on your hardware calibration confidence—typical values range from 0.5° to 5° depending on assembly precision.

Where does the original IMU quaternion get stored?

On first invocation of randomize_imu_orientation, the baseline orientation is captured in env._original_imu_quat via env.sim.model.site_quat[0, site_id].clone(). This preserved quaternion serves as the reference for all subsequent randomization samples, ensuring accumulated errors don't drift unboundedly across episode resets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →