# Common Pitfalls When Training Deep Neural Networks: Lessons from nn-zero-to-hero

> Avoid common pitfalls when training deep neural networks. Learn how vanishing/exploding gradients, poor initialization, and learning rate issues destabilize training and hinder convergence. Master nn-zero-to-hero lessons for ef...

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**Training deep neural networks fails most often due to vanishing or exploding gradients, poor weight initialization, improperly scaled learning rates, and incorrect BatchNorm usage, all of which destabilize gradient flow and prevent effective convergence.**

Training deep neural networks is notoriously fragile, requiring meticulous attention to gradient flow, activation scaling, and hyperparameter tuning. The **nn-zero-to-hero** repository by Andrej Karpathy provides hands-on demonstrations of these failure modes across the *makemore* lecture series. Recognizing these common pitfalls when training deep neural networks is essential for building stable models that converge reliably and generalize beyond the training set.

## Vanishing and Exploding Gradients

In deep architectures, gradients shrink to near-zero or explode toward infinity as they propagate backward through successive layers. According to the repository source, this makes "training deep neural nets... fragile" because improper scaling of activations compounds exponentially with depth [README.md#L38-L41]. When gradients vanish, neurons effectively stop learning; when they explode, numerical overflow corrupts weight updates and destabilizes training.

### The Role of Proper Initialization

Poor weight initialization directly causes these scaling catastrophes. The BatchNorm lecture highlights that random weights with inappropriate variance either saturate nonlinearities or yield gradients too small to update effectively [README.md#L38-L41]. The solution lies in **Kaiming (He) initialization**, which preserves activation variance across deep stacks of layers.

```python
import torch
import torch.nn as nn

class SimpleMLP(nn.Module):
    def __init__(self, dim_in, dim_hidden, dim_out):
        super().__init__()
        self.fc1 = nn.Linear(dim_in, dim_hidden)
        self.fc2 = nn.Linear(dim_hidden, dim_out)
        
        # He/Kaiming init for ReLU activations

        nn.init.kaiming_normal_(self.fc1.weight, nonlinearity='relu')
        nn.init.kaiming_normal_(self.fc2.weight, nonlinearity='relu')
    
    def forward(self, x):
        x = torch.relu(self.fc1(x))
        return self.fc2(x)

```

## Unbalanced Learning Rates and Optimization

An inappropriate learning rate represents one of the most common pitfalls when training deep neural networks. The MLP lecture demonstrates that rates set too high cause divergence and loss oscillation, while rates set too low result in painfully slow convergence that wastes computational resources [README.md#L30-L33]. The "Backprop Ninja" segment further recommends experimenting with adaptive optimizers only after manually verifying gradient computations [README.md#L48-L52].

### Learning Rate Schedules and Gradient Clipping

Implementing a learning rate schedule alongside gradient clipping prevents optimization instability:

```python
optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=1e-2)
scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, T_max=100)

for epoch in range(200):
    for xb, yb in train_loader:
        logits = model(xb)
        loss = F.cross_entropy(logits, yb)
        
        optimizer.zero_grad()
        loss.backward()
        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
        optimizer.step()
    scheduler.step()

```

This configuration uses **AdamW** for adaptive per-parameter step sizes, **weight decay** for regularization, and **cosine annealing** to smoothly decay the learning rate and prevent overshooting near convergence.

## Overfitting and Underfitting

The MLP lecture explicitly addresses the bias-variance trade-off through proper train/dev/test splits [README.md#L30-L33]. Models with excessive capacity or insufficient regularization memorize training noise, while undersized models fail to capture the underlying data distribution.

### Early Stopping Implementation

Prevent overfitting by monitoring validation loss and halting training when generalization deteriorates:

```python
best_val_loss = float('inf')
patience = 5
epochs_no_improve = 0

for epoch in range(200):
    train_one_epoch(model, train_loader)
    val_loss = evaluate(model, val_loader)
    
    if val_loss < best_val_loss:
        best_val_loss = val_loss
        epochs_no_improve = 0
        torch.save(model.state_dict(), 'best.pt')
    else:
        epochs_no_improve += 1
        if epochs_no_improve >= patience:
            print('Early stopping triggered')
            break

```

## Improper Batch Normalization Usage

Batch normalization stabilizes deep training but introduces specific failure modes when misapplied. The nn-zero-to-hero course warns about "pitfalls when they are improperly scaled" and demonstrates correct architectural placement [README.md#L38-L41]. Critical errors include applying BatchNorm before the non-linearity, using prohibitively small batch sizes, or forgetting to set `model.eval()` during inference.

### Correct BatchNorm Placement and Inference Mode

Place BatchNorm after the linear transformation but before the activation function, and strictly enforce evaluation mode during prediction:

```python
class BatchNormMLP(nn.Module):
    def __init__(self, dim_in, dim_hidden, dim_out):
        super().__init__()
        self.fc1 = nn.Linear(dim_in, dim_hidden)
        self.bn1 = nn.BatchNorm1d(dim_hidden)  # BN after linear

        self.fc2 = nn.Linear(dim_hidden, dim_out)
        
        nn.init.kaiming_normal_(self.fc1.weight, nonlinearity='relu')
    
    def forward(self, x):
        x = torch.relu(self.bn1(self.fc1(x)))  # Activation after BN

        return self.fc2(x)

# Inference mode prevents updating running statistics

model.eval()
with torch.no_grad():
    preds = model(test_tensor)

```

Using `model.eval()` ensures the network uses stored training statistics rather than computing new means and variances from test batches, which would corrupt predictions.

## Data Preprocessing and Regularization

While less emphasized in specific lecture titles, the makemore series underscores that inconsistent input scaling or poorly shuffled minibatches create non-stationary data distributions that destabilize training [README.md#L18-L21]. The "Backprop Ninja" segment further encourages using regularization techniques to keep gradients under control and prevent overfitting [README.md#L48-L52].

## Key Source Files for Deeper Study

The following notebooks in the nn-zero-to-hero repository demonstrate these pitfalls through concrete implementations:

- `lectures/makemore/makemore_part3_bn.ipynb`: Hands-on demonstration of BatchNorm pitfalls and correct usage patterns.
- `lectures/makemore/makemore_part4_backprop.ipynb`: Manual backpropagation through a 2-layer MLP, highlighting gradient flow and vanishing/exploding issues.
- `lectures/makemore/makemore_part2_mlp.ipynb`: Learning rate tuning and overfitting/underfitting trade-offs.
- `lectures/makemore/makemore_part5_cnn1.ipynb`: Deep convolutional network scaling problems and architectural solutions.

## Summary

- **Vanishing and exploding gradients** require Kaiming (He) initialization and proper activation scaling to maintain stable variances through deep networks.
- **Learning rate mismatches** cause divergence or stagnation; use adaptive optimizers like AdamW with cosine annealing schedules and gradient clipping.
- **Overfitting** is prevented through validation monitoring and early stopping, while **underfitting** indicates insufficient model capacity.
- **BatchNorm placement** must follow linear layers and precede non-linearities, with strict `eval()` mode enforcement during inference to prevent statistical mismatch.
- **Data preprocessing** and regularization (weight decay, dropout) stabilize training dynamics and improve generalization performance.

## Frequently Asked Questions

### What causes gradients to vanish or explode in deep neural networks?

Gradients vanish or explode due to the multiplicative nature of backpropagation through many layers. When weights are initialized with variances too small or too large, repeated matrix multiplications cause gradients to shrink exponentially toward zero or grow toward infinity. The nn-zero-to-hero course emphasizes that improper activation scaling across depth creates this compounding effect, which BatchNorm and careful initialization mitigate by normalizing layer inputs.

### Should BatchNorm be placed before or after the activation function?

BatchNorm should be placed **after** the linear transformation but **before** the non-linearity (activation function). As demonstrated in `lectures/makemore/makemore_part3_bn.ipynb`, this placement prevents activation saturation by normalizing inputs to the non-linearity. Placing BatchNorm after the activation often results in improperly scaled activations that destabilize gradient flow.

### How do I know if my learning rate is too high or too low?

A learning rate that is **too high** causes the loss to oscillate or diverge to infinity, often producing NaN values within a few iterations. A learning rate that is **too low** shows negligible loss reduction over many epochs. According to the MLP lecture in nn-zero-to-hero, you should observe steady loss decrease without oscillation; use learning rate schedulers like cosine annealing to automate the decay process as training progresses.

### Why is gradient clipping necessary when training deep networks?

Gradient clipping prevents gradient explosion by capping the global L2 norm of gradients to a maximum threshold, typically between 0.5 and 1.0. This is critical for deep networks or recurrent architectures where backpropagated gradients accumulate multiplicatively across layers. As suggested in the "Backprop Ninja" segment, clipping ensures stable weight updates without arbitrarily changing the optimizer's direction, preserving training stability while permitting aggressive learning rates.