Learning Rate Tuning: Why It Matters and How to Find Optimal Hyperparameters in Neural Networks
Learning rate tuning is the process of selecting and scheduling the step size for weight updates to balance rapid convergence against stability, preventing both explosive divergence and sluggish optimization.
The learning rate represents the single most influential hyperparameter controlling how aggressively a neural network adjusts its weights during backpropagation. In the karpathy/nn-zero-to-hero repository, Andrej Karpathy demonstrates that improper tuning causes training loss to either explode into NaN values or plateau prematurely, while strategic schedules enable smooth traversal of complex loss landscapes. Understanding how to find optimal hyperparameters distinguishes production-grade models from failed experiments.
Why Learning Rate Tuning Is Critical
The learning rate directly determines the magnitude of weight updates applied after each backward pass. When the learning rate is too high, the optimizer overshoots local minima, causing loss to bounce chaotically or diverge entirely with exploding gradients. When the learning rate is too low, convergence proceeds painfully slowly, wasting compute cycles and risking entrapment in suboptimal local minima that full-scale training might otherwise escape.
In the nn-zero-to-hero notebooks, Karpathy illustrates that effective learning rate tuning requires navigating a precise trade-off between speed and stability. A well-tuned rate initially drives rapid loss reduction while the loss surface remains smooth, then transitions to smaller steps for fine-grained optimization as the model approaches convergence regions.
How to Find Optimal Learning Rate Hyperparameters
Systematic discovery of optimal values follows a practical workflow grounded in empirical observation rather than guesswork.
Start with a Wide Logarithmic Range
Begin by testing values spaced logarithmically across several orders of magnitude, such as 1e-4, 1e-3, 1e-2, and 1e-1. This approach efficiently brackets the feasible region without wasting compute on clearly catastrophic values.
Identify the Sweet Spot from Loss Curves
Run short experimental training runs and monitor loss curves closely. The optimal learning rate produces a curve that drops steeply during early iterations while continuing smooth descent, rather than flattening prematurely (indicating a rate too low) or spiking upward (indicating a rate too high).
Implement Learning Rate Decay Schedules
Once a base rate is identified, apply a decay strategy to improve final performance. The nn-zero-to-hero repository demonstrates step decay in lectures/makemore/makemore_part5_cnn1.ipynb (lines 385-387), where the learning rate drops from 0.1 to 0.01 after 150,000 iterations:
Alternative decay patterns include:
- Step decay: Sudden drops at fixed milestones (e.g.,
lr = 0.1 if i < 100000 else 0.01as shown inlectures/makemore/makemore_part4_backprop.ipynb) - Exponential decay: Continuous reduction following
lr = initial_lr * decay_rate ** (epoch / decay_steps) - Cosine annealing: Smooth oscillation following a cosine curve, particularly effective for transformer architectures
Validate on Held-Out Data and Automate
Select the learning rate that yields the lowest validation loss rather than the lowest training loss. For comprehensive sweeps, integrate automated optimization libraries like Optuna, Ray Tune, or Keras Tuner, which employ Bayesian optimization to efficiently explore the hyperparameter space.
Practical Implementation in nn-zero-to-hero
The following patterns appear throughout the makemore lecture series, providing concrete implementations for learning rate management.
Step Decay Schedule
This example mirrors the implementation in lectures/makemore/makemore_part5_cnn1.ipynb, using a hard threshold at 150,000 iterations:
def get_lr(iteration):
# Aggressive learning early, fine-tuning later
return 0.1 if iteration < 150_000 else 0.01
# Training loop integration
for i in range(num_iterations):
lr = get_lr(i)
optimizer.param_groups[0]["lr"] = lr # Update optimizer state
loss = model(data)
loss.backward()
optimizer.step()
optimizer.zero_grad()
Exponential Decay Schedule
For smoother reduction, particularly when using Adam optimizer, exponential decay provides continuous adjustment:
def exp_decay_lr(initial_lr, decay_rate, epoch, decay_step):
return initial_lr * (decay_rate ** (epoch / decay_step))
# Configuration for gradual slowdown
initial_lr = 1e-3
decay_rate = 0.5
decay_step = 10 # Halve every 10 epochs
for epoch in range(num_epochs):
lr = exp_decay_lr(initial_lr, decay_rate, epoch, decay_step)
for param_group in optimizer.param_groups:
param_group["lr"] = lr
# Training steps follow...
The lectures/makemore/makemore_part3_bn.ipynb notebook reinforces these concepts by demonstrating how learning rate interacts with batch normalization layers, further emphasizing that optimal hyperparameters depend on the full training pipeline architecture.
Summary
- Learning rate tuning determines whether training converges efficiently, diverges, or stagnates.
- Start with logarithmic sweeps (
1e-4to1e-1) to bracket viable values. - Implement step decay (as shown in
makemore_part5_cnn1.ipynb) to combine fast initial progress with stable final convergence. - Validate choices using held-out validation loss curves rather than training metrics alone.
- Automate extensive searches using Bayesian optimization tools when computing resources permit.
Frequently Asked Questions
What happens if the learning rate is too high?
When the learning rate exceeds the stability threshold, the optimizer takes steps that overshoot local minima in the loss landscape, often causing loss values to explode into NaN or oscillate chaotically without converging. According to the nn-zero-to-hero source code, this manifests as immediate training failure where weights grow unbounded.
How do I know when to decay the learning rate?
Decay the learning rate when the training loss plateaus for a sustained number of epochs or iterations, indicating the optimizer is bouncing around a minimum but unable to settle into it. The repository demonstrates hard-coded milestones (e.g., 100,000 or 150,000 iterations) as a simple heuristic that works well for consistent batch sizes and dataset sizes.
Should I use the same learning rate for all layers?
While the nn-zero-to-hero examples use global learning rates for simplicity, different layers—particularly pretrained embeddings versus randomly initialized classification heads—often benefit from discriminative learning rates where deeper layers receive smaller updates. This prevents catastrophic forgetting in transfer learning scenarios.
Can I automate learning rate tuning?
Yes, frameworks like Optuna and Ray Tune can automate the search using algorithms like Bayesian optimization or Hyperband to efficiently explore the space of learning rates and complementary hyperparameters such as batch size and weight decay. However, manual inspection of loss curves remains essential for understanding model behavior and debugging training failures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →