How Microduck RL Handles NaN Values to Prevent Training Crashes

Microduck RL prevents training crashes from NaN propagation through four coordinated mechanisms: reward sanitisation, advantage/return sanitisation, observation NaN policies, and nan-state termination terms.

Reinforcement learning training with physics simulators like MuJoCo or Isaac Sim frequently encounters Not‑a‑Number (NaN) values from unstable contact dynamics or joint constraint violations. The pollen-robotics/microduck_rl repository implements a multi-layered defense strategy that keeps long-running PPO training stable without manual intervention. This article examines how NaN handling works in Microduck RL based on the source code in src/mjlab_microduck/tasks/mdp.py and related task configuration files.

Reward Sanitisation in the MDP Layer

The first line of defense intercepts NaNs before they reach the PPO buffer. In src/mjlab_microduck/tasks/mdp.py (lines 30‑40), Microduck RL patches the base RewardManager.compute method with _nan_safe_reward_compute.

This wrapper performs two sanitisation steps:

  • Episode sum buffers: Zeroes out NaN entries in self._episode_sums for each reward term
  • Return value: Replaces any NaN in the final reward tensor with 0.0
_orig_reward_compute = _RewardManager.compute

def _nan_safe_reward_compute(self, dt: float) -> torch.Tensor:
    result = _orig_reward_compute(self, dt)
    # Zero‑out NaNs in per‑term episode sums

    for key in self._episode_sums:
        torch.nan_to_num_(self._episode_sums[key], nan=0.0)
    # Return a NaN‑free reward tensor

    return torch.nan_to_num(result, nan=0.0)

_RewardManager.compute = _nan_safe_reward_compute

This ensures that a single corrupted physics step cannot poison the reward statistics used for PPO's advantage estimation.

Advantage and Return Sanitisation for Stable Gradients

Even with clean rewards, numerical instability in advantage computation can introduce NaNs or infinities. The _safe_compute_returns function in mdp.py (lines 50‑56) wraps the PPO algorithm's return calculation to sanitise st.advantages and st.returns:

_orig_compute_returns = _PPO.compute_returns

def _safe_compute_returns(self, obs) -> None:
    _orig_compute_returns(self, obs)
    st = self.storage
    torch.nan_to_num_(st.advantages, nan=0.0, posinf=0.0, neginf=0.0)
    torch.nan_to_num_(st.returns,    nan=0.0, posinf=0.0, neginf=0.0)

_PPO.compute_returns = _safe_compute_returns

By handling posinf and neginf alongside NaN, this prevents the normalisation step from exploding gradients during backpropagation.

Observation NaN Policies in Task Configurations

Microduck RL tasks explicitly configure observation handling to avoid NaN-related episode aborts. In configuration files such as microduck_roller_standup_env_cfg.py and microduck_roller_slope_env_cfg.py, every observation group sets:

for grp in ("actor", "critic"):
    cfg.observations[grp].nan_policy = "sanitize"

This observation NaN policy instructs the observation manager to replace NaN and Inf entries with 0 when constructing observation vectors. Without this setting, rsl_rl's global check_nan would terminate episodes immediately upon detecting any NaN, disrupting training progress.

NaN-State Termination for Clean Episode Boundaries

When NaNs indicate catastrophic simulator state corruption, Microduck RL terminates episodes gracefully rather than allowing error propagation. The robot_state_is_nan function in mdp.py detects NaNs in:

  • Joint positions
  • Joint velocities
  • Contact forces (from specified sensors)

Task configurations register this as a termination condition:

cfg.terminations["nan_state"] = TerminationTermCfg(
    func=microduck_mdp.robot_state_is_nan,
    params={"sensor_names": ("feet",)},
    time_out=False,
)

The time_out=False flag indicates this is a failure termination, not a time limit. This allows the training loop to reset the environment and continue without crashing.

Unit Tests Guaranteeing NaN Safety

The repository includes targeted regression tests verifying NaN handling behavior:

Test File Coverage
tests/test_obs_nan_guard.py Validates observation NaN-policy sanitisation
tests/test_nan_guard.py Verifies robot_state_is_nan termination logic
tests/test_wheel_glide.py Tests NaN-safe reward computation for specific terms
tests/test_descent_speed.py Tests NaN-safe reward computation for descent rewards

These tests ensure that modifications to the MDP layer or task configurations do not compromise NaN protection.

Summary

Microduck RL prevents training crashes from NaN values through a defense-in-depth strategy:

  • Reward sanitisation via _nan_safe_reward_compute in mdp.py zeros NaN rewards and episode sums
  • Advantage/return sanitisation via _safe_compute_returns protects gradient statistics
  • Observation NaN policies (nan_policy = "sanitize") prevent premature episode termination
  • NaN-state termination detects and cleanly resets corrupted simulator states
  • Unit tests maintain regression coverage for all NaN handling paths

Together, these mechanisms ensure that isolated physics instabilities do not corrupt the PPO training pipeline, enabling robust long-duration training runs.

Frequently Asked Questions

How does Microduck RL prevent NaN rewards from reaching the PPO algorithm?

The _nan_safe_reward_compute wrapper in src/mjlab_microduck/tasks/mdp.py intercepts all reward computations, uses torch.nan_to_num_ to sanitize episode sum buffers, and returns a NaN-free tensor via torch.nan_to_num. This prevents any NaN reward from entering the PPO buffer or influencing advantage estimates.

What happens when an observation contains NaN values?

Task configurations set cfg.observations[grp].nan_policy = "sanitize" for both actor and critic observation groups. This triggers the observation manager to replace NaN and Inf values with 0 during vector construction, preventing rsl_rl's native NaN checker from aborting the episode.

Does Microduck RL terminate episodes when NaN states are detected?

Yes. The robot_state_is_nan termination term (registered in task configs as cfg.terminations["nan_state"]) monitors joint positions, velocities, and contact forces. When NaNs are detected, the episode terminates cleanly with time_out=False, allowing environment reset without training interruption.

Where are the NaN handling mechanisms tested?

Regression tests in tests/test_obs_nan_guard.py, tests/test_nan_guard.py, and reward-specific tests (test_wheel_glide.py, test_descent_speed.py) verify that observation sanitisation, state termination, and reward computation all correctly handle NaN values according to the implementations in src/mjlab_microduck/tasks/mdp.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →