# Debugging Neural Networks: Common Failure Modes and Systematic Remedies

> Debug neural networks effectively by understanding common failure modes. Learn systematic remedies for issues like vanishing gradients and data leakage using tests such as overfit-one-batch.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-26

---

**Neural networks often hide bugs that never raise exceptions yet silently degrade performance, requiring systematic techniques like overfit-one-batch tests and gradient monitoring to diagnose issues from vanishing gradients to data leakage.**

Unlike traditional software that crashes on null pointers or type mismatches, neural networks in the `ai-engineering-from-scratch` repository demonstrate that 60-70% of ML debugging time is spent on "silent" failures—models that run to completion and emit loss values while underlying optimization dynamics have quietly broken. The lesson *"Debugging Neural Networks"* (Phase 03 → Lesson 13) provides a battle-tested taxonomy of these failure modes and a practical toolbox for diagnosing each symptom.

## Why Neural Network Debugging Is Hard

Traditional software crashes immediately on errors like null pointers or type mismatches. Neural networks, however, may execute successfully, emit a loss value, and still be fundamentally wrong because the training loop contains a single misplaced line—such as a missing `zero_grad()`, a transposed weight matrix, or an inappropriate learning rate. According to the lesson documentation in [`/phases/03-deep-learning-core/13-debugging-neural-networks/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/03-deep-learning-core/13-debugging-neural-networks/docs/en.md), these silent bugs dominate debugging time and require specialized detection strategies.

## The Debugging Mindset

**Start simple, add complexity one piece at a time, and verify each piece independently.**

This mantra guides the systematic approach outlined in the repository. A decision-tree diagram in [`/phases/03-deep-learning-core/13-debugging-neural-networks/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/03-deep-learning-core/13-debugging-neural-networks/docs/en.md) provides a diagnostic path from "Loss not decreasing" → "Check learning rate" → "Check gradients" → "Check data pipeline" → "Check architecture". By isolating variables rather than adjusting multiple hyperparameters simultaneously, you can pinpoint whether the issue stems from the model architecture, data pipeline, or optimization configuration.

## Common Failure Modes and Symptoms

The lesson provides a symptom-to-cause mapping for the most prevalent neural network failure modes:

- **Loss flat or oscillating**: Typically indicates a learning rate that is too high or too low. First-try remedy involves sweeping the learning rate using an LR-finder.
- **NaN/Inf loss**: Caused by learning rates that are too high, `log(0)` operations, division-by-zero, or numerical overflow. Fix by clamping predictions, adding epsilon values, or lowering the learning rate.
- **Training ≈ Test ≈ Chance**: Suggests the model is not learning at all, often due to implementation bugs. Run the **overfit-one-batch** test to verify the model can memorize a tiny dataset.
- **Training high, Test low**: Indicates overfitting or data leakage. Implement dropout, add weight decay, or verify your data split integrity.
- **Gradient all zero**: Signals dead ReLUs or detached computation graphs. Switch to LeakyReLU and ensure tensors have `.requires_grad=True`.
- **OOM during training**: Batch size too large or computation graph not freed. Reduce batch size and use `torch.no_grad()` for evaluation.
- **Slow or no loss decrease**: Often caused by forgetting `optimizer.zero_grad()`, causing gradients to accumulate across steps.

## Core Debugging Techniques

The [`/phases/03-deep-learning-core/13-debugging-neural-networks/code/debug_neural_nets.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/03-deep-learning-core/13-debugging-neural-networks/code/debug_neural_nets.py) file implements four essential diagnostic tools:

### NetworkDebugger Class

The `NetworkDebugger` class hooks into each layer to record activation and gradient statistics, flagging dead neurons, exploding gradients, or vanishing gradients. Instantiate it with your model, record loss values during training, and generate diagnostic reports:

```python
from phases.03_deep_learning_core.13_debugging_neural_networks.code.debug_neural_nets import NetworkDebugger

debugger = NetworkDebugger(model)

# ... training loop ...

debugger.record_loss(loss.item())
debugger.print_report()  # Shows per-layer health metrics

debugger.remove_hooks()

```

### Overfit-One-Batch Test

This sanity check validates that your model and training loop can actually learn. The `overfit_one_batch()` function trains on a single tiny batch until loss approaches zero, confirming that gradients flow correctly and the optimizer updates parameters:

```python
from phases.03_deep_learning_core.13_debugging_neural_networks.code.debug_neural_nets import overfit_one_batch

assert overfit_one_batch(model, x_batch, y_batch, criterion)

```

### Learning-Rate Finder

Instead of guessing learning rates, the `find_learning_rate()` function performs an exponential sweep over one epoch, selecting the learning rate just before loss diverges. This automated approach eliminates manual tuning guesswork.

### Gradient Check

The `gradient_check()` function compares analytical gradients computed via backpropagation against numerical finite-difference approximations. This catches subtle backpropagation implementation bugs where manual gradient calculations might differ from autograd results.

## PyTorch-Specific Pitfalls

The lesson documents frequent implementation errors specific to PyTorch that cause silent failures:

- **Forgetting `optimizer.zero_grad()`**: Gradients accumulate across steps, causing loss oscillation. Always call before `loss.backward()`.
- **Missing `model.eval()`**: Dropout and batch normalization behave differently during training and inference, causing fluctuating test accuracy. Wrap evaluation in `model.eval()` and `torch.no_grad()`.
- **Shape mismatches broadcasting silently**: Operations with incorrect tensor shapes produce wrong results without raising errors. Print tensor shapes after each operation during development.
- **CPU/GPU mismatch**: Runtime errors occur when model and data reside on different devices. Use `.to(device)` consistently for both.
- **In-place operations breaking autograd**: Statements like `x += 1` modify tensors in-place and break the computation graph. Replace with `x = x + 1`.
- **Un-normalized inputs**: Loss remains stuck at chance level when input features lack zero mean and unit variance. Normalize inputs to mean ≈ 0, std ≈ 1.
- **Wrong label dtype**: Cross-entropy loss expects `Long` tensors; passing other types causes silent failures. Cast labels using `labels.long()`.

## A Complete Debug Workflow

Following the systematic approach from the repository, here is a complete debugging workflow that combines multiple techniques:

```python
import torch
import torch.nn as nn
from phases.03_deep_learning_core.13_debugging_neural_networks.code.debug_neural_nets import (
    NetworkDebugger, overfit_one_batch, find_learning_rate
)

# Step 1: Build a simple baseline model

model = nn.Sequential(
    nn.Linear(10, 32),
    nn.ReLU(),
    nn.Linear(32, 2)
)

# Step 2: Verify the model can learn (overfit-one-batch test)

x_batch = torch.randn(8, 10)
y_batch = (x_batch[:, 0] > 0).long()
criterion = nn.CrossEntropyLoss()

assert overfit_one_batch(model, x_batch, y_batch, criterion), \
    "Model cannot overfit one batch - check architecture/loss"

# Step 3: Find appropriate learning rate

optimal_lr = find_learning_rate(model, x_batch, y_batch, criterion)

# Step 4: Attach debugger and train with monitoring

debugger = NetworkDebugger(model)
optimizer = torch.optim.Adam(model.parameters(), lr=optimal_lr)

for step in range(20):
    optimizer.zero_grad()  # Critical: prevent gradient accumulation

    out = model(x_batch)
    loss = criterion(out, y_batch)
    debugger.record_loss(loss.item())
    loss.backward()
    optimizer.step()

debugger.print_report()  # Check for DEAD_NEURONS or NAN_OR_INF

debugger.remove_hooks()

```

This workflow enforces the recommended order: **overfit-one-batch → attach debugger → monitor loss/gradients**. If any check flags an issue—such as "DEAD_NEURONS" or "NAN_OR_INF"—you can immediately narrow the root cause before scaling to full-dataset training.

## Summary

- **Silent failures dominate**: 60-70% of ML debugging involves bugs that don't crash but degrade performance, like missing `zero_grad()` or inappropriate learning rates.
- **Start simple**: Use the overfit-one-batch test to verify your model can learn before training on full datasets.
- **Monitor internals**: The `NetworkDebugger` class in [`/phases/03-deep-learning-core/13-debugging-neural-networks/code/debug_neural_nets.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/03-deep-learning-core/13-debugging-neural-networks/code/debug_neural_nets.py) provides per-layer activation and gradient statistics to catch dead neurons and exploding gradients.
- **Automate hyperparameters**: Use the learning-rate finder to eliminate guesswork in selecting optimization parameters.
- **Verify gradients**: The `gradient_check()` function validates backpropagation correctness against numerical approximations.

## Frequently Asked Questions

### Why does my neural network loss remain flat even after many epochs?

A flat loss curve typically indicates a learning rate that is too small, dead ReLU neurons causing zero gradients, or forgotten `optimizer.zero_grad()` calls that accumulate stale gradients. According to the `ai-engineering-from-scratch` curriculum, run the **overfit-one-batch** test first to confirm the model can learn at all, then check for dead neurons using the `NetworkDebugger` class.

### What causes NaN or Inf values in my loss during training?

NaN or Inf losses usually stem from learning rates that are too high causing numerical overflow, `log(0)` operations in loss functions like cross-entropy, or division-by-zero in custom layers. The lesson recommends clamping predictions to avoid zero values, adding epsilon constants to logarithms, and using the learning-rate finder to identify stable training ranges.

### How can I detect if my PyTorch model has dead ReLU neurons?

Dead ReLUs occur when neurons output zero for all inputs, stopping gradient flow. The `NetworkDebugger` class in [`/phases/03-deep-learning-core/13-debugging-neural-networks/code/debug_neural_nets.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main//phases/03-deep-learning-core/13-debugging-neural-networks/code/debug_neural_nets.py) tracks the zero-fraction of activations per layer. If you detect dead neurons, switch to LeakyReLU or PReLU activations to maintain small gradients for negative inputs.

### What is the overfit-one-batch test and when should I use it?

The **overfit-one-batch** test trains your model on a single small batch of data until the loss approaches zero, verifying that your architecture, loss function, and optimizer can actually reduce loss. You should run this immediately after implementing a new model or training loop—before wasting compute on full datasets—to catch bugs like incorrect loss functions or detached computation graphs.