How Encoder Biases Are Handled in Microduck RL Domain Randomization

Microduck RL treats encoder bias as a per-environment systematic offset sampled during startup randomization, exposing it only to the actor while keeping critic observations unbiased, then penalizes persistent bias through an EMA-based reward term with curriculum-driven weighting.

In the pollen-robotics/microduck_rl repository, encoder bias domain randomization (DR) simulates real-world sensor inaccuracies to train robust policies. The implementation spans task configuration files, observation pipelines, and specialized reward functions that together ensure the policy learns to operate despite uncertain joint position measurements.

The Encoder Bias Randomization Event

Domain randomization in Microduck RL uses an event-driven architecture where startup events configure simulation parameters before episode rollouts begin. The encoder_bias event is the central mechanism for injecting systematic joint position offsets.

Event Configuration and Sampling

Each task configuration enables encoder bias by initializing the encoder_bias event with a configurable range. In microduck_velocity_env_cfg.py, the velocity environment sets this range via a module-level constant:

cfg.events["encoder_bias"].params["bias_range"] = ENCODER_BIAS_RANGE

Source: [microduck_velocity_env_cfg.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L626-L633)

This event samples a constant offset for every joint from a uniform distribution defined by ENCODER_BIAS_RANGE, typically expressed in radians. The sampling occurs once per environment reset during the startup phase, creating environment-specific biases that persist throughout the episode.

Event Cleanup After Application

To prevent double-application on subsequent resets, the configuration removes the event immediately after parameterization:

cfg.events.pop("encoder_bias", None)

Source: [microduck_velocity_env_cfg.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L635)

This pattern ensures idempotent behavior—the bias is sampled and locked in, then the event infrastructure is stripped away before the training loop begins.

Actor-Critic Observation Split: The Biased/Unbiased Pattern

Microduck RL implements a critical asymmetry in how encoder bias affects learning: the policy network (actor) receives biased observations while the value network (critic) accesses ground-truth joint positions. This mirrors physical reality where the robot controller operates on potentially miscalibrated sensor readings, but the underlying dynamics—and thus the true state for value estimation—remain unbiased.

Observation Term Configuration

The joint position observation term is duplicated with contrasting biased flags:

cfg.observations["actor"].terms["joint_pos"].params["biased"] = True
cfg.observations["critic"].terms["joint_pos"].params["biased"] = False

Source: [microduck_velocity_env_cfg.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L628-L633)

  • Actor (biased=True): Receives joint_pos + encoder_bias, forcing the policy to learn robust behaviors invariant to systematic offset.
  • Critic (biased=False): Receives true joint positions, enabling accurate value estimation unaffected by sensor calibration errors.

This separation is essential for Sample Efficiency: the value function learns from clean signals while the policy adapts to realistic, noisy inputs.

Head Pose Bias Penalty: Reward-Based Bias Suppression

Beyond randomization, Microduck RL actively discourages persistent biases through a dedicated reward term. The head_pose_bias_penalty function in mdp.py detects and penalizes static offsets using exponential moving average (EMA) filtering.

EMA-Based Bias Extraction

The penalty logic isolates DC bias from oscillatory motion:

def head_pose_bias_penalty(env, tau_s=1.0, **kw):
    # Compute error between desired neck pose and current neck pose

    err = env.obs["head_pose"] - env.desired_head_pose
    # EMA to filter out high-frequency oscillations

    alpha = 1.0 / (tau_s * env.dt)
    if not hasattr(env, "_head_bias_ema"):
        env._head_bias_ema = torch.zeros_like(err)
    # Reset EMA on fresh episodes

    fresh = env.reset_mask
    env._head_bias_ema[fresh] = 0.0
    # Update EMA

    env._head_bias_ema = (1.0 - alpha) * env._head_bias_ema + alpha * err
    # Penalty is the L1 norm of the bias (negative reward)

    out = -env._head_bias_ema.abs().mean(dim=-1)
    return out

Source: [mdp.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/mdp.py#L5229-L5251)

The EMA time constant tau_s (default 1.0 seconds) determines the cutoff frequency. High-frequency tracking errors are attenuated, while persistent biases accumulate in _head_bias_ema and incur penalties proportional to their magnitude.

Reward Term Configuration

The bias penalty is registered as a standard reward term with initially zero weight:

cfg.rewards["head_pose_bias"] = RewardTermCfg(
    func=microduck_mdp.head_pose_bias_penalty,
    weight=0.0,  # ramped by curriculum

)

Source: [microduck_velocity_env_cfg.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L738-L743)

Curriculum-Driven Weight Scheduling

To prevent early training collapse, the bias penalty weight follows a curriculum schedule that phases in the regularization after initial skill acquisition:

cfg.curriculum["head_pose_bias_weight"] = CurriculumTermCfg(
    param_name="weight",
    schedule=[(600, 0.0), (1500, 3.0)],
)

Source: [microduck_velocity_env_cfg.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L894-L899)

The schedule specifies:

  • 0–600 iterations: Penalty weight held at 0.0, allowing free exploration of biased behaviors.
  • 600–1500 iterations: Linear ramp to 3.0, gradually enforcing bias suppression.
  • 1500+ iterations: Full penalty weight of 3.0 actively discourages persistent offsets.

This progression ensures the policy first learns fundamental locomotion skills before being constrained by additional regularization objectives.

Complete Implementation Example

To enable encoder bias domain randomization for a custom task:

from mjlab_microduck.tasks.cfg_base import make_microduck_velocity_env_cfg

def make_my_custom_env_cfg():
    cfg = make_microduck_velocity_env_cfg(play=False, rough=False)

    # 1. Configure encoder-bias DR range

    cfg.events["encoder_bias"].params["bias_range"] = 0.1  # ±0.1 rad

    # 2. Apply actor/critic observation split

    cfg.observations["actor"].terms["joint_pos"].params["biased"] = True
    cfg.observations["critic"].terms["joint_pos"].params["biased"] = False

    # 3. Clean up event to prevent re-sampling

    cfg.events.pop("encoder_bias", None)

    # 4. Optional: add bias penalty with curriculum

    cfg.rewards["head_pose_bias"] = RewardTermCfg(
        func=microduck_mdp.head_pose_bias_penalty,
        weight=0.0,
    )
    cfg.curriculum["head_pose_bias_weight"] = CurriculumTermCfg(
        param_name="weight",
        schedule=[(600, 0.0), (1500, 2.0)],
    )
    return cfg

During training, inspect sampled bias values through the MuJoCo model:


# After env.reset()

bias_vector = env.sim.mj_model.actuator_biasprm[:, 2]
print("Sampled encoder bias (rad):", bias_vector)

Key Source Files

File Purpose
[microduck_velocity_env_cfg.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) Encoder-bias event setup, observation flags, reward/curriculum configuration
[mdp.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/mdp.py) head_pose_bias_penalty implementation with EMA filtering
[microduck_standup_env_cfg.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_standup_env_cfg.py) Parallel implementation pattern for standup tasks
[tests/test_head_pose_bias.py](https://github.com/pollen-robotics/microduck_rl/blob/develop/tests/test_head_pose_bias.py) Unit tests verifying bias observation isolation and penalty behavior

Summary

  • Encoder bias in Microduck RL is a startup randomization event that samples per-environment offsets from a configurable range.
  • Actor-critic separation exposes bias only to the policy network while preserving ground-truth observations for value estimation.
  • EMA-based penalty in head_pose_bias_penalty isolates DC bias from dynamic motion and penalizes persistent offsets.
  • Curriculum scheduling phases in the penalty weight after initial skill acquisition to maintain training stability.
  • Event cleanup ensures biases are sampled once and locked, preventing re-randomization during episode resets.

Frequently Asked Questions

What is the typical magnitude of encoder bias randomization in Microduck RL?

The magnitude is task-dependent but typically falls within ±0.1 to ±0.5 radians, configured through the ENCODER_BIAS_RANGE constant in environment configuration files. For the velocity environment, this range is imported from base configuration modules and can be overridden per-task.

Why does the critic receive unbiased joint positions while the actor sees biased values?

This separation prevents value function corruption from sensor calibration errors. The critic estimates expected returns from true states, enabling accurate advantage estimation. The actor must learn robust policies that function despite systematic measurement offsets—matching real robot deployment where policies operate on potentially miscalibrated encoders.

How does the EMA filtering in head_pose_bias_penalty work?

The exponential moving average with time constant tau_s=1.0 seconds acts as a low-pass filter. The update rule ema = (1-alpha)*ema + alpha*error progressively accumulates persistent tracking errors (bias) while attenuating high-frequency oscillations. The resulting ema approximates the DC offset, which is penalized through its L1 norm.

Can encoder bias randomization be disabled for debugging or ablation studies?

Yes. Simply omit the encoder_bias event configuration or set bias_range=0.0. Additionally, setting both actor and critic biased=False ensures all observations use ground-truth joint positions. The modular event architecture allows clean removal without affecting other domain randomization features.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →