# How Encoder Biases Are Handled in Microduck RL Domain Randomization

> Learn how Microduck RL addresses encoder bias in domain randomization. Discover its unique approach to sampling, actor-critic observation, and EMA-based reward for unbiased performance.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: deep-dive
- Published: 2026-09-02

---

**Microduck RL treats encoder bias as a per-environment systematic offset sampled during startup randomization, exposing it only to the actor while keeping critic observations unbiased, then penalizes persistent bias through an EMA-based reward term with curriculum-driven weighting.**

In the `pollen-robotics/microduck_rl` repository, encoder bias domain randomization (DR) simulates real-world sensor inaccuracies to train robust policies. The implementation spans task configuration files, observation pipelines, and specialized reward functions that together ensure the policy learns to operate despite uncertain joint position measurements.

## The Encoder Bias Randomization Event

Domain randomization in Microduck RL uses an **event-driven architecture** where startup events configure simulation parameters before episode rollouts begin. The `encoder_bias` event is the central mechanism for injecting systematic joint position offsets.

### Event Configuration and Sampling

Each task configuration enables encoder bias by initializing the `encoder_bias` event with a configurable range. In [`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py), the velocity environment sets this range via a module-level constant:

```python
cfg.events["encoder_bias"].params["bias_range"] = ENCODER_BIAS_RANGE

```

*Source: [[`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L626-L633)*

This event samples a constant offset for every joint from a uniform distribution defined by `ENCODER_BIAS_RANGE`, typically expressed in radians. The sampling occurs once per environment reset during the startup phase, creating environment-specific biases that persist throughout the episode.

### Event Cleanup After Application

To prevent double-application on subsequent resets, the configuration removes the event immediately after parameterization:

```python
cfg.events.pop("encoder_bias", None)

```

*Source: [[`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L635)*

This pattern ensures idempotent behavior—the bias is sampled and locked in, then the event infrastructure is stripped away before the training loop begins.

## Actor-Critic Observation Split: The Biased/Unbiased Pattern

Microduck RL implements a **critical asymmetry** in how encoder bias affects learning: the policy network (actor) receives biased observations while the value network (critic) accesses ground-truth joint positions. This mirrors physical reality where the robot controller operates on potentially miscalibrated sensor readings, but the underlying dynamics—and thus the true state for value estimation—remain unbiased.

### Observation Term Configuration

The joint position observation term is duplicated with contrasting `biased` flags:

```python
cfg.observations["actor"].terms["joint_pos"].params["biased"] = True
cfg.observations["critic"].terms["joint_pos"].params["biased"] = False

```

*Source: [[`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L628-L633)*

- **Actor (`biased=True`)**: Receives `joint_pos + encoder_bias`, forcing the policy to learn robust behaviors invariant to systematic offset.
- **Critic (`biased=False`)**: Receives true joint positions, enabling accurate value estimation unaffected by sensor calibration errors.

This separation is essential for **Sample Efficiency**: the value function learns from clean signals while the policy adapts to realistic, noisy inputs.

## Head Pose Bias Penalty: Reward-Based Bias Suppression

Beyond randomization, Microduck RL actively discourages persistent biases through a dedicated reward term. The `head_pose_bias_penalty` function in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) detects and penalizes static offsets using exponential moving average (EMA) filtering.

### EMA-Based Bias Extraction

The penalty logic isolates DC bias from oscillatory motion:

```python
def head_pose_bias_penalty(env, tau_s=1.0, **kw):
    # Compute error between desired neck pose and current neck pose

    err = env.obs["head_pose"] - env.desired_head_pose
    # EMA to filter out high-frequency oscillations

    alpha = 1.0 / (tau_s * env.dt)
    if not hasattr(env, "_head_bias_ema"):
        env._head_bias_ema = torch.zeros_like(err)
    # Reset EMA on fresh episodes

    fresh = env.reset_mask
    env._head_bias_ema[fresh] = 0.0
    # Update EMA

    env._head_bias_ema = (1.0 - alpha) * env._head_bias_ema + alpha * err
    # Penalty is the L1 norm of the bias (negative reward)

    out = -env._head_bias_ema.abs().mean(dim=-1)
    return out

```

*Source: [[`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/mdp.py#L5229-L5251)*

The EMA time constant `tau_s` (default 1.0 seconds) determines the cutoff frequency. High-frequency tracking errors are attenuated, while persistent biases accumulate in `_head_bias_ema` and incur penalties proportional to their magnitude.

### Reward Term Configuration

The bias penalty is registered as a standard reward term with initially zero weight:

```python
cfg.rewards["head_pose_bias"] = RewardTermCfg(
    func=microduck_mdp.head_pose_bias_penalty,
    weight=0.0,  # ramped by curriculum

)

```

*Source: [[`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L738-L743)*

## Curriculum-Driven Weight Scheduling

To prevent early training collapse, the bias penalty weight follows a **curriculum schedule** that phases in the regularization after initial skill acquisition:

```python
cfg.curriculum["head_pose_bias_weight"] = CurriculumTermCfg(
    param_name="weight",
    schedule=[(600, 0.0), (1500, 3.0)],
)

```

*Source: [[`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py#L894-L899)*

The schedule specifies:
- **0–600 iterations**: Penalty weight held at 0.0, allowing free exploration of biased behaviors.
- **600–1500 iterations**: Linear ramp to 3.0, gradually enforcing bias suppression.
- **1500+ iterations**: Full penalty weight of 3.0 actively discourages persistent offsets.

This progression ensures the policy first learns fundamental locomotion skills before being constrained by additional regularization objectives.

## Complete Implementation Example

To enable encoder bias domain randomization for a custom task:

```python
from mjlab_microduck.tasks.cfg_base import make_microduck_velocity_env_cfg

def make_my_custom_env_cfg():
    cfg = make_microduck_velocity_env_cfg(play=False, rough=False)

    # 1. Configure encoder-bias DR range

    cfg.events["encoder_bias"].params["bias_range"] = 0.1  # ±0.1 rad

    # 2. Apply actor/critic observation split

    cfg.observations["actor"].terms["joint_pos"].params["biased"] = True
    cfg.observations["critic"].terms["joint_pos"].params["biased"] = False

    # 3. Clean up event to prevent re-sampling

    cfg.events.pop("encoder_bias", None)

    # 4. Optional: add bias penalty with curriculum

    cfg.rewards["head_pose_bias"] = RewardTermCfg(
        func=microduck_mdp.head_pose_bias_penalty,
        weight=0.0,
    )
    cfg.curriculum["head_pose_bias_weight"] = CurriculumTermCfg(
        param_name="weight",
        schedule=[(600, 0.0), (1500, 2.0)],
    )
    return cfg

```

During training, inspect sampled bias values through the MuJoCo model:

```python

# After env.reset()

bias_vector = env.sim.mj_model.actuator_biasprm[:, 2]
print("Sampled encoder bias (rad):", bias_vector)

```

## Key Source Files

| File | Purpose |
|------|---------|
| [[`microduck_velocity_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_velocity_env_cfg.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_velocity_env_cfg.py) | Encoder-bias event setup, observation flags, reward/curriculum configuration |
| [[`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/mdp.py) | `head_pose_bias_penalty` implementation with EMA filtering |
| [[`microduck_standup_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_standup_env_cfg.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/src/mjlab_microduck/tasks/microduck_standup_env_cfg.py) | Parallel implementation pattern for standup tasks |
| [[`tests/test_head_pose_bias.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_head_pose_bias.py)](https://github.com/pollen-robotics/microduck_rl/blob/develop/tests/test_head_pose_bias.py) | Unit tests verifying bias observation isolation and penalty behavior |

## Summary

- **Encoder bias in Microduck RL** is a startup randomization event that samples per-environment offsets from a configurable range.
- **Actor-critic separation** exposes bias only to the policy network while preserving ground-truth observations for value estimation.
- **EMA-based penalty** in `head_pose_bias_penalty` isolates DC bias from dynamic motion and penalizes persistent offsets.
- **Curriculum scheduling** phases in the penalty weight after initial skill acquisition to maintain training stability.
- **Event cleanup** ensures biases are sampled once and locked, preventing re-randomization during episode resets.

## Frequently Asked Questions

### What is the typical magnitude of encoder bias randomization in Microduck RL?

The magnitude is task-dependent but typically falls within ±0.1 to ±0.5 radians, configured through the `ENCODER_BIAS_RANGE` constant in environment configuration files. For the velocity environment, this range is imported from base configuration modules and can be overridden per-task.

### Why does the critic receive unbiased joint positions while the actor sees biased values?

This separation prevents value function corruption from sensor calibration errors. The critic estimates expected returns from true states, enabling accurate advantage estimation. The actor must learn robust policies that function despite systematic measurement offsets—matching real robot deployment where policies operate on potentially miscalibrated encoders.

### How does the EMA filtering in head_pose_bias_penalty work?

The exponential moving average with time constant `tau_s=1.0` seconds acts as a low-pass filter. The update rule `ema = (1-alpha)*ema + alpha*error` progressively accumulates persistent tracking errors (bias) while attenuating high-frequency oscillations. The resulting `ema` approximates the DC offset, which is penalized through its L1 norm.

### Can encoder bias randomization be disabled for debugging or ablation studies?

Yes. Simply omit the `encoder_bias` event configuration or set `bias_range=0.0`. Additionally, setting both actor and critic `biased=False` ensures all observations use ground-truth joint positions. The modular event architecture allows clean removal without affecting other domain randomization features.