How to Diagnose Overfitting vs Underfitting in Neural Networks: A Practical Guide with nn-zero-to-hero

Diagnose overfitting versus underfitting by comparing training loss against validation loss: if training loss falls while validation loss plateaus or rises, the model is overfitting (high variance); if both losses remain high and decrease together, the model is underfitting (high bias).

Neural networks learn by minimizing loss on training data, but minimizing training error alone guarantees poor generalization. The karpathy/nn-zero-to-hero repository demonstrates how to diagnose these pathologies through explicit data splits and rigorous loss monitoring in the makemore lecture series. By analyzing the training loop implementations in lectures/makemore/makemore_part2_mlp.ipynb, you can apply a systematic workflow to detect whether your model suffers from insufficient capacity or excessive memorization.

Establishing the Three-Way Data Split

Reliable diagnosis requires partitioning data into three distinct sets. In the makemore notebooks, Andrej Karpathy explicitly constructs these splits using the build_dataset function:

  • Training set (Xtr, Ytr): Drives gradient updates via backpropagation.
  • Validation (development) set (Xdev, Ydev): Monitors generalization during training without contaminating hyperparameter choices.
  • Test set (Xte, Yte): Provides a final, unbiased estimate of performance after tuning.

The repository implements an 80/10/10 split:


# From lectures/makemore/makemore_part2_mlp.ipynb

random.shuffle(words)
n1 = int(0.8 * len(words))
n2 = int(0.9 * len(words))
Xtr, Ytr = build_dataset(words[:n1])
Xdev, Ydev = build_dataset(words[n1:n2])
Xte, Yte   = build_dataset(words[n2:])

Tracking loss on both Xtr and Xdev during the training loop is the critical first step for distinguishing fitting regimes.

Interpreting Training vs. Validation Loss Curves

The shape of the loss curves reveals the model's fitting status. You must log the iteration index (stepi) and loss values for both splits to visualize these patterns:

  • Both losses high and decreasing together: Indicates underfitting. The model lacks sufficient capacity or has not trained long enough to capture data patterns.
  • Training loss continues falling while validation loss plateaus or rises: Signals overfitting. The network memorizes training examples and fails to generalize to unseen data.
  • Training loss low, validation loss low, with a minimal gap: Represents a good fit. Model capacity matches data complexity.

As implemented in makemore_part2_mlp.ipynb, you record these metrics inside the training loop:

stepi, train_loss_i, dev_loss_i = [], [], []

for i in range(200000):
    # ... forward pass and backward pass ...

    stepi.append(i)
    train_loss_i.append(loss.item())
    
    if i % 1000 == 0:
        with torch.no_grad():
            dev_logits = (torch.tanh(C[Xdev].view(-1,30) @ W1 + b1) @ W2 + b2)
            dev_loss = F.cross_entropy(dev_logits, Ydev)
            dev_loss_i.append(dev_loss.item())

Plotting train_loss_i against dev_loss_i at regular intervals exposes the generalization gap immediately.

Quantifying the Generalization Gap

Beyond visual inspection, calculate the numerical gap between splits to automate diagnostics:

  • Gap size: Compute gap = train_loss - dev_loss. A gap exceeding approximately 10% of the training loss typically indicates overfitting.
  • Early stopping: Halt training when dev_loss fails to improve for a defined patience period (e.g., 5000 steps). This prevents the model from entering the overfitting regime where validation performance degrades.

The repository demonstrates early stopping logic manually in Lecture 4, where inspection of the loss curve determines the optimal stopping point before divergence occurs.

Hyperparameter Levers for Correcting Fit

Once diagnosed, adjust these specific levers to move the model toward the "good fit" zone:

  • Model capacity (embedding dimension, hidden layer width): Increase to combat underfitting; decrease to reduce overfitting.
  • Regularization techniques: Implement BatchNorm (shown in lectures/makemore/makemore_part3_bn.ipynb) to stabilize activations and reduce overfitting, or add dropout and weight decay.
  • Learning rate schedule: A learning rate that is too large causes erratic convergence (possible underfitting), while a rate too small may prevent reaching low training loss entirely.
  • Dataset size: Expanding training data naturally shrinks the train-dev gap by forcing the model to learn generalizable patterns rather than memorizing specific examples.

Step-by-Step Diagnostic Workflow

Apply this reproducible protocol derived from the nn-zero-to-hero source code:

  1. Initialize splits using build_dataset to create Xtr, Xdev, and Xte with an 80/10/10 distribution.
  2. Run a short training loop (e.g., 10,000 iterations) logging loss.item() for training and evaluating F.cross_entropy on Xdev every 1,000 steps.
  3. Visualize the curves by plotting stepi against train_loss_i and dev_loss_i using matplotlib.
  4. Inspect the gap: If dev_loss consistently exceeds train_loss by a large margin, apply regularization or reduce model size. If both curves plateau at high values, increase hidden dimensions or train longer.
  5. Apply early stopping when dev_loss stops decreasing, then evaluate final performance on Xte for an unbiased metric.

Summary

  • Three splits are mandatory: Use build_dataset in makemore_part2_mlp.ipynb to create distinct training, validation, and test sets.
  • Loss divergence reveals overfitting: When training loss drops but validation loss rises, the model has high variance.
  • Parallel high losses indicate underfitting: Both curves remaining elevated signals insufficient model capacity.
  • Quantify the gap: Calculate train_loss - dev_loss to automate detection of generalization failure.
  • Act on levers: Adjust hidden layer width, add BatchNorm (as shown in makemore_part3_bn.ipynb), or implement early stopping to correct the fit.

Frequently Asked Questions

How do I formally calculate the generalization gap in PyTorch?

Compute the difference between the average training loss and the average validation loss over a stable window of iterations. In the nn-zero-to-hero notebooks, this is implemented by accumulating loss.item() during training and periodically running a no-gradient evaluation pass on Xdev using F.cross_entropy, then subtracting the dev loss from the train loss.

What is the difference between the validation set and the test set in this context?

The validation set (Xdev) monitors generalization during training and guides hyperparameter tuning, while the test set (Xte) provides a final performance estimate only after all tuning is complete. According to the repository structure, you must never use Xte to make training decisions, as this would leak information and invalidate your generalization estimate.

When should I stop training to prevent overfitting?

Implement early stopping by tracking dev_loss at regular intervals (e.g., every 1,000 steps) and halting when the validation loss fails to improve for a predetermined patience period (typically 5,000 to 10,000 steps). This technique is demonstrated conceptually in Lecture 4 of the makemore series within the nn-zero-to-hero repository.

Can BatchNorm help with overfitting, and where is it shown in the codebase?

Yes, Batch Normalization acts as a regularizer by reducing internal covariate shift and adding noise to activations, which helps prevent overfitting. The implementation and diagnostic visualization of BatchNorm are covered in lectures/makemore/makemore_part3_bn.ipynb, where activation statistics are tracked to ensure stable gradient flow.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →