# How to Diagnose Overfitting vs Underfitting in Neural Networks: A Practical Guide with nn-zero-to-hero

> Easily diagnose overfitting vs underfitting in neural networks by comparing training and validation loss. Learn practical tips from nn-zero-to-hero to build better models.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: how-to-guide
- Published: 2026-05-23

---

**Diagnose overfitting versus underfitting by comparing training loss against validation loss: if training loss falls while validation loss plateaus or rises, the model is overfitting (high variance); if both losses remain high and decrease together, the model is underfitting (high bias).**

Neural networks learn by minimizing loss on training data, but minimizing training error alone guarantees poor generalization. The **karpathy/nn-zero-to-hero** repository demonstrates how to diagnose these pathologies through explicit data splits and rigorous loss monitoring in the *makemore* lecture series. By analyzing the training loop implementations in `lectures/makemore/makemore_part2_mlp.ipynb`, you can apply a systematic workflow to detect whether your model suffers from insufficient capacity or excessive memorization.

## Establishing the Three-Way Data Split

Reliable diagnosis requires partitioning data into three distinct sets. In the *makemore* notebooks, Andrej Karpathy explicitly constructs these splits using the `build_dataset` function:

- **Training set** (`Xtr`, `Ytr`): Drives gradient updates via backpropagation.
- **Validation (development) set** (`Xdev`, `Ydev`): Monitors generalization during training without contaminating hyperparameter choices.
- **Test set** (`Xte`, `Yte`): Provides a final, unbiased estimate of performance after tuning.

The repository implements an 80/10/10 split:

```python

# From lectures/makemore/makemore_part2_mlp.ipynb

random.shuffle(words)
n1 = int(0.8 * len(words))
n2 = int(0.9 * len(words))
Xtr, Ytr = build_dataset(words[:n1])
Xdev, Ydev = build_dataset(words[n1:n2])
Xte, Yte   = build_dataset(words[n2:])

```

Tracking loss on both `Xtr` and `Xdev` during the training loop is the critical first step for distinguishing fitting regimes.

## Interpreting Training vs. Validation Loss Curves

The shape of the loss curves reveals the model's fitting status. You must log the iteration index (`stepi`) and loss values for both splits to visualize these patterns:

- **Both losses high and decreasing together**: Indicates **underfitting**. The model lacks sufficient capacity or has not trained long enough to capture data patterns.
- **Training loss continues falling while validation loss plateaus or rises**: Signals **overfitting**. The network memorizes training examples and fails to generalize to unseen data.
- **Training loss low, validation loss low, with a minimal gap**: Represents a **good fit**. Model capacity matches data complexity.

As implemented in `makemore_part2_mlp.ipynb`, you record these metrics inside the training loop:

```python
stepi, train_loss_i, dev_loss_i = [], [], []

for i in range(200000):
    # ... forward pass and backward pass ...

    stepi.append(i)
    train_loss_i.append(loss.item())
    
    if i % 1000 == 0:
        with torch.no_grad():
            dev_logits = (torch.tanh(C[Xdev].view(-1,30) @ W1 + b1) @ W2 + b2)
            dev_loss = F.cross_entropy(dev_logits, Ydev)
            dev_loss_i.append(dev_loss.item())

```

Plotting `train_loss_i` against `dev_loss_i` at regular intervals exposes the generalization gap immediately.

## Quantifying the Generalization Gap

Beyond visual inspection, calculate the numerical gap between splits to automate diagnostics:

- **Gap size**: Compute `gap = train_loss - dev_loss`. A gap exceeding approximately 10% of the training loss typically indicates overfitting.
- **Early stopping**: Halt training when `dev_loss` fails to improve for a defined patience period (e.g., 5000 steps). This prevents the model from entering the overfitting regime where validation performance degrades.

The repository demonstrates early stopping logic manually in Lecture 4, where inspection of the loss curve determines the optimal stopping point before divergence occurs.

## Hyperparameter Levers for Correcting Fit

Once diagnosed, adjust these specific levers to move the model toward the "good fit" zone:

- **Model capacity** (embedding dimension, hidden layer width): Increase to combat underfitting; decrease to reduce overfitting.
- **Regularization techniques**: Implement **BatchNorm** (shown in `lectures/makemore/makemore_part3_bn.ipynb`) to stabilize activations and reduce overfitting, or add dropout and weight decay.
- **Learning rate schedule**: A learning rate that is too large causes erratic convergence (possible underfitting), while a rate too small may prevent reaching low training loss entirely.
- **Dataset size**: Expanding training data naturally shrinks the train-dev gap by forcing the model to learn generalizable patterns rather than memorizing specific examples.

## Step-by-Step Diagnostic Workflow

Apply this reproducible protocol derived from the nn-zero-to-hero source code:

1. **Initialize splits** using `build_dataset` to create `Xtr`, `Xdev`, and `Xte` with an 80/10/10 distribution.
2. **Run a short training loop** (e.g., 10,000 iterations) logging `loss.item()` for training and evaluating `F.cross_entropy` on `Xdev` every 1,000 steps.
3. **Visualize the curves** by plotting `stepi` against `train_loss_i` and `dev_loss_i` using `matplotlib`.
4. **Inspect the gap**: If `dev_loss` consistently exceeds `train_loss` by a large margin, apply regularization or reduce model size. If both curves plateau at high values, increase hidden dimensions or train longer.
5. **Apply early stopping** when `dev_loss` stops decreasing, then evaluate final performance on `Xte` for an unbiased metric.

## Summary

- **Three splits are mandatory**: Use `build_dataset` in `makemore_part2_mlp.ipynb` to create distinct training, validation, and test sets.
- **Loss divergence reveals overfitting**: When training loss drops but validation loss rises, the model has high variance.
- **Parallel high losses indicate underfitting**: Both curves remaining elevated signals insufficient model capacity.
- **Quantify the gap**: Calculate `train_loss - dev_loss` to automate detection of generalization failure.
- **Act on levers**: Adjust hidden layer width, add BatchNorm (as shown in `makemore_part3_bn.ipynb`), or implement early stopping to correct the fit.

## Frequently Asked Questions

### How do I formally calculate the generalization gap in PyTorch?

Compute the difference between the average training loss and the average validation loss over a stable window of iterations. In the nn-zero-to-hero notebooks, this is implemented by accumulating `loss.item()` during training and periodically running a no-gradient evaluation pass on `Xdev` using `F.cross_entropy`, then subtracting the dev loss from the train loss.

### What is the difference between the validation set and the test set in this context?

The **validation set** (`Xdev`) monitors generalization during training and guides hyperparameter tuning, while the **test set** (`Xte`) provides a final performance estimate only after all tuning is complete. According to the repository structure, you must never use `Xte` to make training decisions, as this would leak information and invalidate your generalization estimate.

### When should I stop training to prevent overfitting?

Implement **early stopping** by tracking `dev_loss` at regular intervals (e.g., every 1,000 steps) and halting when the validation loss fails to improve for a predetermined patience period (typically 5,000 to 10,000 steps). This technique is demonstrated conceptually in Lecture 4 of the makemore series within the nn-zero-to-hero repository.

### Can BatchNorm help with overfitting, and where is it shown in the codebase?

Yes, **Batch Normalization** acts as a regularizer by reducing internal covariate shift and adding noise to activations, which helps prevent overfitting. The implementation and diagnostic visualization of BatchNorm are covered in `lectures/makemore/makemore_part3_bn.ipynb`, where activation statistics are tracked to ensure stable gradient flow.