# Learning Rate Tuning: Why It Matters and How to Find Optimal Hyperparameters in Neural Networks

> Master learning rate tuning in neural networks. Discover why optimal hyperparameters are crucial for fast yet stable convergence and prevent common optimization pitfalls.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**Learning rate tuning is the process of selecting and scheduling the step size for weight updates to balance rapid convergence against stability, preventing both explosive divergence and sluggish optimization.**

The learning rate represents the single most influential hyperparameter controlling how aggressively a neural network adjusts its weights during backpropagation. In the `karpathy/nn-zero-to-hero` repository, Andrej Karpathy demonstrates that improper tuning causes training loss to either explode into `NaN` values or plateau prematurely, while strategic schedules enable smooth traversal of complex loss landscapes. Understanding how to find optimal hyperparameters distinguishes production-grade models from failed experiments.

## Why Learning Rate Tuning Is Critical

The learning rate directly determines the magnitude of weight updates applied after each backward pass. When the learning rate is **too high**, the optimizer overshoots local minima, causing loss to bounce chaotically or diverge entirely with exploding gradients. When the learning rate is **too low**, convergence proceeds painfully slowly, wasting compute cycles and risking entrapment in suboptimal local minima that full-scale training might otherwise escape.

In the `nn-zero-to-hero` notebooks, Karpathy illustrates that effective **learning rate tuning** requires navigating a precise trade-off between speed and stability. A well-tuned rate initially drives rapid loss reduction while the loss surface remains smooth, then transitions to smaller steps for fine-grained optimization as the model approaches convergence regions.

## How to Find Optimal Learning Rate Hyperparameters

Systematic discovery of optimal values follows a practical workflow grounded in empirical observation rather than guesswork.

### Start with a Wide Logarithmic Range

Begin by testing values spaced logarithmically across several orders of magnitude, such as `1e-4`, `1e-3`, `1e-2`, and `1e-1`. This approach efficiently brackets the feasible region without wasting compute on clearly catastrophic values.

### Identify the Sweet Spot from Loss Curves

Run short experimental training runs and monitor loss curves closely. The optimal learning rate produces a curve that drops steeply during early iterations while continuing smooth descent, rather than flattening prematurely (indicating a rate too low) or spiking upward (indicating a rate too high).

### Implement Learning Rate Decay Schedules

Once a base rate is identified, apply a decay strategy to improve final performance. The `nn-zero-to-hero` repository demonstrates **step decay** in `lectures/makemore/makemore_part5_cnn1.ipynb` (lines 385-387), where the learning rate drops from `0.1` to `0.01` after 150,000 iterations:

Alternative decay patterns include:
- **Step decay**: Sudden drops at fixed milestones (e.g., `lr = 0.1 if i < 100000 else 0.01` as shown in `lectures/makemore/makemore_part4_backprop.ipynb`)
- **Exponential decay**: Continuous reduction following `lr = initial_lr * decay_rate ** (epoch / decay_steps)`
- **Cosine annealing**: Smooth oscillation following a cosine curve, particularly effective for transformer architectures

### Validate on Held-Out Data and Automate

Select the learning rate that yields the lowest validation loss rather than the lowest training loss. For comprehensive sweeps, integrate automated optimization libraries like **Optuna**, **Ray Tune**, or **Keras Tuner**, which employ Bayesian optimization to efficiently explore the hyperparameter space.

## Practical Implementation in nn-zero-to-hero

The following patterns appear throughout the `makemore` lecture series, providing concrete implementations for learning rate management.

### Step Decay Schedule

This example mirrors the implementation in `lectures/makemore/makemore_part5_cnn1.ipynb`, using a hard threshold at 150,000 iterations:

```python
def get_lr(iteration):
    # Aggressive learning early, fine-tuning later

    return 0.1 if iteration < 150_000 else 0.01

# Training loop integration

for i in range(num_iterations):
    lr = get_lr(i)
    optimizer.param_groups[0]["lr"] = lr  # Update optimizer state

    loss = model(data)
    loss.backward()
    optimizer.step()
    optimizer.zero_grad()

```

### Exponential Decay Schedule

For smoother reduction, particularly when using Adam optimizer, exponential decay provides continuous adjustment:

```python
def exp_decay_lr(initial_lr, decay_rate, epoch, decay_step):
    return initial_lr * (decay_rate ** (epoch / decay_step))

# Configuration for gradual slowdown

initial_lr = 1e-3
decay_rate = 0.5
decay_step = 10  # Halve every 10 epochs

for epoch in range(num_epochs):
    lr = exp_decay_lr(initial_lr, decay_rate, epoch, decay_step)
    for param_group in optimizer.param_groups:
        param_group["lr"] = lr
    # Training steps follow...

```

The `lectures/makemore/makemore_part3_bn.ipynb` notebook reinforces these concepts by demonstrating how learning rate interacts with batch normalization layers, further emphasizing that optimal hyperparameters depend on the full training pipeline architecture.

## Summary

- **Learning rate tuning** determines whether training converges efficiently, diverges, or stagnates.
- Start with logarithmic sweeps (`1e-4` to `1e-1`) to bracket viable values.
- Implement **step decay** (as shown in `makemore_part5_cnn1.ipynb`) to combine fast initial progress with stable final convergence.
- Validate choices using held-out validation loss curves rather than training metrics alone.
- Automate extensive searches using Bayesian optimization tools when computing resources permit.

## Frequently Asked Questions

### What happens if the learning rate is too high?

When the learning rate exceeds the stability threshold, the optimizer takes steps that overshoot local minima in the loss landscape, often causing loss values to explode into `NaN` or oscillate chaotically without converging. According to the `nn-zero-to-hero` source code, this manifests as immediate training failure where weights grow unbounded.

### How do I know when to decay the learning rate?

Decay the learning rate when the training loss plateaus for a sustained number of epochs or iterations, indicating the optimizer is bouncing around a minimum but unable to settle into it. The repository demonstrates hard-coded milestones (e.g., 100,000 or 150,000 iterations) as a simple heuristic that works well for consistent batch sizes and dataset sizes.

### Should I use the same learning rate for all layers?

While the `nn-zero-to-hero` examples use global learning rates for simplicity, different layers—particularly pretrained embeddings versus randomly initialized classification heads—often benefit from **discriminative learning rates** where deeper layers receive smaller updates. This prevents catastrophic forgetting in transfer learning scenarios.

### Can I automate learning rate tuning?

Yes, frameworks like **Optuna** and **Ray Tune** can automate the search using algorithms like Bayesian optimization or Hyperband to efficiently explore the space of learning rates and complementary hyperparameters such as batch size and weight decay. However, manual inspection of loss curves remains essential for understanding model behavior and debugging training failures.