# How Microduck RL Handles NaN Values to Prevent Training Crashes

> Learn how Microduck RL prevents training crashes from NaN values using four mechanisms: reward sanitisation, advantage sanitisation, observation NaN policies, and nan-state termination.

- Repository: [Pollen Robotics/microduck_rl](https://github.com/pollen-robotics/microduck_rl)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Microduck RL prevents training crashes from NaN propagation through four coordinated mechanisms: reward sanitisation, advantage/return sanitisation, observation NaN policies, and nan-state termination terms.**

Reinforcement learning training with physics simulators like MuJoCo or Isaac Sim frequently encounters *Not‑a‑Number* (NaN) values from unstable contact dynamics or joint constraint violations. The [`pollen-robotics/microduck_rl`](https://github.com/pollen-robotics/microduck_rl) repository implements a multi-layered defense strategy that keeps long-running PPO training stable without manual intervention. This article examines how NaN handling works in Microduck RL based on the source code in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) and related task configuration files.

## Reward Sanitisation in the MDP Layer

The first line of defense intercepts NaNs before they reach the PPO buffer. In [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) (lines 30‑40), Microduck RL patches the base `RewardManager.compute` method with `_nan_safe_reward_compute`.

This wrapper performs two sanitisation steps:

- **Episode sum buffers**: Zeroes out NaN entries in `self._episode_sums` for each reward term
- **Return value**: Replaces any NaN in the final reward tensor with `0.0`

```python
_orig_reward_compute = _RewardManager.compute

def _nan_safe_reward_compute(self, dt: float) -> torch.Tensor:
    result = _orig_reward_compute(self, dt)
    # Zero‑out NaNs in per‑term episode sums

    for key in self._episode_sums:
        torch.nan_to_num_(self._episode_sums[key], nan=0.0)
    # Return a NaN‑free reward tensor

    return torch.nan_to_num(result, nan=0.0)

_RewardManager.compute = _nan_safe_reward_compute

```

This ensures that a single corrupted physics step cannot poison the reward statistics used for PPO's advantage estimation.

## Advantage and Return Sanitisation for Stable Gradients

Even with clean rewards, numerical instability in advantage computation can introduce NaNs or infinities. The `_safe_compute_returns` function in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) (lines 50‑56) wraps the PPO algorithm's return calculation to sanitise `st.advantages` and `st.returns`:

```python
_orig_compute_returns = _PPO.compute_returns

def _safe_compute_returns(self, obs) -> None:
    _orig_compute_returns(self, obs)
    st = self.storage
    torch.nan_to_num_(st.advantages, nan=0.0, posinf=0.0, neginf=0.0)
    torch.nan_to_num_(st.returns,    nan=0.0, posinf=0.0, neginf=0.0)

_PPO.compute_returns = _safe_compute_returns

```

By handling `posinf` and `neginf` alongside NaN, this prevents the normalisation step from exploding gradients during backpropagation.

## Observation NaN Policies in Task Configurations

Microduck RL tasks explicitly configure observation handling to avoid NaN-related episode aborts. In configuration files such as [`microduck_roller_standup_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_roller_standup_env_cfg.py) and [`microduck_roller_slope_env_cfg.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/microduck_roller_slope_env_cfg.py), every observation group sets:

```python
for grp in ("actor", "critic"):
    cfg.observations[grp].nan_policy = "sanitize"

```

This **observation NaN policy** instructs the observation manager to replace NaN and Inf entries with `0` when constructing observation vectors. Without this setting, `rsl_rl`'s global `check_nan` would terminate episodes immediately upon detecting any NaN, disrupting training progress.

## NaN-State Termination for Clean Episode Boundaries

When NaNs indicate catastrophic simulator state corruption, Microduck RL terminates episodes gracefully rather than allowing error propagation. The `robot_state_is_nan` function in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) detects NaNs in:

- Joint positions
- Joint velocities  
- Contact forces (from specified sensors)

Task configurations register this as a termination condition:

```python
cfg.terminations["nan_state"] = TerminationTermCfg(
    func=microduck_mdp.robot_state_is_nan,
    params={"sensor_names": ("feet",)},
    time_out=False,
)

```

The `time_out=False` flag indicates this is a failure termination, not a time limit. This allows the training loop to reset the environment and continue without crashing.

## Unit Tests Guaranteeing NaN Safety

The repository includes targeted regression tests verifying NaN handling behavior:

| Test File | Coverage |
|-----------|----------|
| [`tests/test_obs_nan_guard.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_obs_nan_guard.py) | Validates observation NaN-policy sanitisation |
| [`tests/test_nan_guard.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_nan_guard.py) | Verifies `robot_state_is_nan` termination logic |
| [`tests/test_wheel_glide.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_wheel_glide.py) | Tests NaN-safe reward computation for specific terms |
| [`tests/test_descent_speed.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_descent_speed.py) | Tests NaN-safe reward computation for descent rewards |

These tests ensure that modifications to the MDP layer or task configurations do not compromise NaN protection.

## Summary

Microduck RL prevents training crashes from NaN values through a defense-in-depth strategy:

- **Reward sanitisation** via `_nan_safe_reward_compute` in [`mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/mdp.py) zeros NaN rewards and episode sums
- **Advantage/return sanitisation** via `_safe_compute_returns` protects gradient statistics
- **Observation NaN policies** (`nan_policy = "sanitize"`) prevent premature episode termination
- **NaN-state termination** detects and cleanly resets corrupted simulator states
- **Unit tests** maintain regression coverage for all NaN handling paths

Together, these mechanisms ensure that isolated physics instabilities do not corrupt the PPO training pipeline, enabling robust long-duration training runs.

## Frequently Asked Questions

### How does Microduck RL prevent NaN rewards from reaching the PPO algorithm?

The `_nan_safe_reward_compute` wrapper in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py) intercepts all reward computations, uses `torch.nan_to_num_` to sanitize episode sum buffers, and returns a NaN-free tensor via `torch.nan_to_num`. This prevents any NaN reward from entering the PPO buffer or influencing advantage estimates.

### What happens when an observation contains NaN values?

Task configurations set `cfg.observations[grp].nan_policy = "sanitize"` for both actor and critic observation groups. This triggers the observation manager to replace NaN and Inf values with `0` during vector construction, preventing `rsl_rl`'s native NaN checker from aborting the episode.

### Does Microduck RL terminate episodes when NaN states are detected?

Yes. The `robot_state_is_nan` termination term (registered in task configs as `cfg.terminations["nan_state"]`) monitors joint positions, velocities, and contact forces. When NaNs are detected, the episode terminates cleanly with `time_out=False`, allowing environment reset without training interruption.

### Where are the NaN handling mechanisms tested?

Regression tests in [`tests/test_obs_nan_guard.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_obs_nan_guard.py), [`tests/test_nan_guard.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/tests/test_nan_guard.py), and reward-specific tests ([`test_wheel_glide.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/test_wheel_glide.py), [`test_descent_speed.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/test_descent_speed.py)) verify that observation sanitisation, state termination, and reward computation all correctly handle NaN values according to the implementations in [`src/mjlab_microduck/tasks/mdp.py`](https://github.com/pollen-robotics/microduck_rl/blob/main/src/mjlab_microduck/tasks/mdp.py).