How d2l-zh Explains Learning Rate Optimization for Model Convergence
d2l-zh explains learning rate optimization as a critical balance between convergence speed and stability, teaching that while too-small values cause slow progress and too-large values trigger divergence, dynamic schedules like cosine decay with warm-up in chapter_optimization/lr-scheduler.md enable reliable convergence for deep networks.
The Chinese edition of Dive into Deep Learning (d2l-zh) dedicates multiple chapters to learning rate optimization, treating the learning rate (denoted as η) as the single most influential hyper-parameter for ensuring training algorithms converge to optimal solutions. Through implementations in chapter_optimization/gd.md and related files, the repository provides both theoretical explanations and runnable code demonstrating how proper scheduling prevents oscillation and accelerates convergence across different optimizer types.
Why the Learning Rate Matters for Convergence
In chapter_optimization/gd.md (lines 47-66), the authors establish that η controls the step size during parameter updates. A value that is too small results in painfully slow progress that may never reach the global minimum within practical time limits, while an excessively large value causes the loss function to bounce or diverge entirely. The book uses a simple quadratic objective f(x) = x² to visualize these effects, showing that stable convergence requires balancing step size against gradient magnitude.
def gd(eta, f_grad):
x = 10.0
for _ in range(10):
x -= eta * f_grad(x)
return x
def f_grad(x):
return 2 * x # gradient of f(x)=x²
# Too small – slow progress
print(gd(0.05, f_grad)) # → ~5.0 after 10 steps
# Reasonable – converges
print(gd(0.2, f_grad)) # → ~0.0 after 10 steps
# Too large – diverges
print(gd(1.1, f_grad)) # → grows quickly, loss explodes
Static vs. Dynamic Learning Rate Schedules
While constant learning rates suffice for convex problems, d2l-zh argues in chapter_optimization/sgd.md (lines 107-119) that deep learning requires dynamic schedules that adapt η during training. The repository implements several strategies in chapter_optimization/lr-scheduler.md to handle non-convex loss landscapes where aggressive early steps and conservative later steps improve final model quality.
Piecewise-Constant Decay
The FactorScheduler class demonstrates step-wise reduction. As shown in chapter_optimization/lr-scheduler.md (lines 5-14), this approach multiplies the learning rate by a decay factor (e.g., 0.5) every fixed number of epochs, allowing the optimizer to take smaller, more precise steps as it approaches a minimum.
from d2l import mxnet as d2l
# Define a scheduler that halves the LR every 5 epochs
lr_sched = d2l.lr_scheduler.FactorScheduler(step=5, factor=0.5, base_lr=0.3)
for epoch in range(15):
lr = lr_sched(epoch)
print(f'epoch {epoch+1}: lr = {lr:.4f}')
Cosine Decay with Warm-Up
For deep networks, d2l-zh recommends combining an initial warm-up period with cosine decay. According to chapter_optimization/lr-scheduler.md (lines 378-386), the CosineScheduler gradually decreases the learning rate following a cosine curve, providing smoother convergence than abrupt step changes. The warm-up phase stabilizes early training when gradients are noisy, particularly effective when using large batch sizes.
from d2l import mxnet as d2l
# Warm-up for 3 epochs, then cosine decay to 0.01
lr_sched = d2l.lr_scheduler.CosineScheduler(base_lr=0.3,
warmup_steps=3,
total_steps=30,
final_lr=0.01)
for step in range(30):
print(f'step {step+1}: lr = {lr_sched(step):.5f}')
Practical Guidelines for Learning Rate Selection
The authors provide concrete experimental workflows in chapter_optimization/sgd.md (lines 130-144). They recommend starting with a modest constant η (such as 0.1 or 0.01), monitoring the loss curve for plateaus or explosions, and implementing a schedule only after confirming the initial value produces stable gradients. The book emphasizes that no universal learning rate exists across different architectures or datasets, requiring empirical testing for each new problem.
Learning Rate in Adaptive Optimizers
Even adaptive methods like Adam require careful tuning of the base learning rate. In chapter_optimization/adam.md (lines 55-68), d2l-zh demonstrates that while Adam adjusts per-parameter learning rates automatically based on gradient history, the global η (typically defaulting to 0.01) significantly impacts convergence speed and final model quality.
from d2l import mxnet as d2l
trainer = d2l.Adam(learning_rate=0.01) # default η = 0.01
# … training loop …
Summary
- d2l-zh identifies the learning rate η as the most critical hyper-parameter for model convergence in
chapter_optimization/gd.md, where improper values cause either slow progress or divergence. - Static learning rates often fail for deep networks; dynamic schedules like cosine decay with warm-up in
chapter_optimization/lr-scheduler.mdprovide superior stability. - The repository implements practical schedulers including
FactorSchedulerfor step decay andCosineSchedulerfor smooth annealing to zero. - Adaptive optimizers like Adam still require base learning rate tuning, typically starting at 0.01, as shown in
chapter_optimization/adam.md.
Frequently Asked Questions
What happens if the learning rate is too large in d2l-zh examples?
According to the gradient descent demonstrations in chapter_optimization/gd.md (lines 47-66), an excessively large η causes the loss to oscillate or diverge rather than converge, as the parameter updates overshoot the minimum and potentially increase the loss indefinitely. The code example with eta=1.1 demonstrates this explosive behavior explicitly.
Does d2l-zh recommend warm-up for all deep learning models?
While not mandatory for every architecture, d2l-zh specifically recommends warm-up periods followed by cosine decay for deep networks and large batch sizes, as implemented in chapter_optimization/lr-scheduler.md (lines 378-386). This combination prevents early training instability when gradients are noisy and learning is most fragile.
How does d2l-zh suggest finding the optimal initial learning rate?
The text in chapter_optimization/sgd.md (lines 130-144) advises testing several η values (such as 0.1, 0.01, and 0.001) with a constant schedule first, observing whether the loss decreases smoothly or explodes. Practitioners should select the largest value that produces stable convergence before applying dynamic scheduling to refine the training process.
Do adaptive optimizers like Adam eliminate the need for learning rate scheduling in d2l-zh?
No, the Adam implementation in chapter_optimization/adam.md (lines 55-68) still relies on a base learning rate that requires tuning, and the book demonstrates that combining Adam with cosine decay schedules often yields better convergence than using a fixed rate throughout training. The adaptive moment estimation adjusts per-parameter scales but does not replace the need for global learning rate annealing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →