Learning Rate Tuning: Why It Matters and How to Find Optimal Hyperparameters in Neural Networks

Learning rate tuning is the process of selecting and scheduling the step size for weight updates to balance rapid convergence against stability, preventing both explosive divergence and sluggish optimization.

The learning rate represents the single most influential hyperparameter controlling how aggressively a neural network adjusts its weights during backpropagation. In the karpathy/nn-zero-to-hero repository, Andrej Karpathy demonstrates that improper tuning causes training loss to either explode into NaN values or plateau prematurely, while strategic schedules enable smooth traversal of complex loss landscapes. Understanding how to find optimal hyperparameters distinguishes production-grade models from failed experiments.

Why Learning Rate Tuning Is Critical

The learning rate directly determines the magnitude of weight updates applied after each backward pass. When the learning rate is too high, the optimizer overshoots local minima, causing loss to bounce chaotically or diverge entirely with exploding gradients. When the learning rate is too low, convergence proceeds painfully slowly, wasting compute cycles and risking entrapment in suboptimal local minima that full-scale training might otherwise escape.

In the nn-zero-to-hero notebooks, Karpathy illustrates that effective learning rate tuning requires navigating a precise trade-off between speed and stability. A well-tuned rate initially drives rapid loss reduction while the loss surface remains smooth, then transitions to smaller steps for fine-grained optimization as the model approaches convergence regions.

How to Find Optimal Learning Rate Hyperparameters

Systematic discovery of optimal values follows a practical workflow grounded in empirical observation rather than guesswork.

Start with a Wide Logarithmic Range

Begin by testing values spaced logarithmically across several orders of magnitude, such as 1e-4, 1e-3, 1e-2, and 1e-1. This approach efficiently brackets the feasible region without wasting compute on clearly catastrophic values.

Identify the Sweet Spot from Loss Curves

Run short experimental training runs and monitor loss curves closely. The optimal learning rate produces a curve that drops steeply during early iterations while continuing smooth descent, rather than flattening prematurely (indicating a rate too low) or spiking upward (indicating a rate too high).

Implement Learning Rate Decay Schedules

Once a base rate is identified, apply a decay strategy to improve final performance. The nn-zero-to-hero repository demonstrates step decay in lectures/makemore/makemore_part5_cnn1.ipynb (lines 385-387), where the learning rate drops from 0.1 to 0.01 after 150,000 iterations:

Alternative decay patterns include:

  • Step decay: Sudden drops at fixed milestones (e.g., lr = 0.1 if i < 100000 else 0.01 as shown in lectures/makemore/makemore_part4_backprop.ipynb)
  • Exponential decay: Continuous reduction following lr = initial_lr * decay_rate ** (epoch / decay_steps)
  • Cosine annealing: Smooth oscillation following a cosine curve, particularly effective for transformer architectures

Validate on Held-Out Data and Automate

Select the learning rate that yields the lowest validation loss rather than the lowest training loss. For comprehensive sweeps, integrate automated optimization libraries like Optuna, Ray Tune, or Keras Tuner, which employ Bayesian optimization to efficiently explore the hyperparameter space.

Practical Implementation in nn-zero-to-hero

The following patterns appear throughout the makemore lecture series, providing concrete implementations for learning rate management.

Step Decay Schedule

This example mirrors the implementation in lectures/makemore/makemore_part5_cnn1.ipynb, using a hard threshold at 150,000 iterations:

def get_lr(iteration):
    # Aggressive learning early, fine-tuning later

    return 0.1 if iteration < 150_000 else 0.01

# Training loop integration

for i in range(num_iterations):
    lr = get_lr(i)
    optimizer.param_groups[0]["lr"] = lr  # Update optimizer state

    loss = model(data)
    loss.backward()
    optimizer.step()
    optimizer.zero_grad()

Exponential Decay Schedule

For smoother reduction, particularly when using Adam optimizer, exponential decay provides continuous adjustment:

def exp_decay_lr(initial_lr, decay_rate, epoch, decay_step):
    return initial_lr * (decay_rate ** (epoch / decay_step))

# Configuration for gradual slowdown

initial_lr = 1e-3
decay_rate = 0.5
decay_step = 10  # Halve every 10 epochs

for epoch in range(num_epochs):
    lr = exp_decay_lr(initial_lr, decay_rate, epoch, decay_step)
    for param_group in optimizer.param_groups:
        param_group["lr"] = lr
    # Training steps follow...

The lectures/makemore/makemore_part3_bn.ipynb notebook reinforces these concepts by demonstrating how learning rate interacts with batch normalization layers, further emphasizing that optimal hyperparameters depend on the full training pipeline architecture.

Summary

  • Learning rate tuning determines whether training converges efficiently, diverges, or stagnates.
  • Start with logarithmic sweeps (1e-4 to 1e-1) to bracket viable values.
  • Implement step decay (as shown in makemore_part5_cnn1.ipynb) to combine fast initial progress with stable final convergence.
  • Validate choices using held-out validation loss curves rather than training metrics alone.
  • Automate extensive searches using Bayesian optimization tools when computing resources permit.

Frequently Asked Questions

What happens if the learning rate is too high?

When the learning rate exceeds the stability threshold, the optimizer takes steps that overshoot local minima in the loss landscape, often causing loss values to explode into NaN or oscillate chaotically without converging. According to the nn-zero-to-hero source code, this manifests as immediate training failure where weights grow unbounded.

How do I know when to decay the learning rate?

Decay the learning rate when the training loss plateaus for a sustained number of epochs or iterations, indicating the optimizer is bouncing around a minimum but unable to settle into it. The repository demonstrates hard-coded milestones (e.g., 100,000 or 150,000 iterations) as a simple heuristic that works well for consistent batch sizes and dataset sizes.

Should I use the same learning rate for all layers?

While the nn-zero-to-hero examples use global learning rates for simplicity, different layers—particularly pretrained embeddings versus randomly initialized classification heads—often benefit from discriminative learning rates where deeper layers receive smaller updates. This prevents catastrophic forgetting in transfer learning scenarios.

Can I automate learning rate tuning?

Yes, frameworks like Optuna and Ray Tune can automate the search using algorithms like Bayesian optimization or Hyperband to efficiently explore the space of learning rates and complementary hyperparameters such as batch size and weight decay. However, manual inspection of loss curves remains essential for understanding model behavior and debugging training failures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →