# How d2l-zh Explains Learning Rate Optimization for Model Convergence

> d2l-zh explains learning rate optimization for model convergence detailing optimal ranges and dynamic schedules like cosine decay with warm-up for stable, fast deep network training.

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: tutorial
- Published: 2026-03-01

---

**d2l-zh explains learning rate optimization as a critical balance between convergence speed and stability, teaching that while too-small values cause slow progress and too-large values trigger divergence, dynamic schedules like cosine decay with warm-up in [`chapter_optimization/lr-scheduler.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/lr-scheduler.md) enable reliable convergence for deep networks.**

The Chinese edition of *Dive into Deep Learning* (d2l-zh) dedicates multiple chapters to learning rate optimization, treating the learning rate (denoted as **η**) as the single most influential hyper-parameter for ensuring training algorithms converge to optimal solutions. Through implementations in [`chapter_optimization/gd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/gd.md) and related files, the repository provides both theoretical explanations and runnable code demonstrating how proper scheduling prevents oscillation and accelerates convergence across different optimizer types.

## Why the Learning Rate Matters for Convergence

In [`chapter_optimization/gd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/gd.md) (lines 47-66), the authors establish that **η** controls the step size during parameter updates. A value that is too small results in painfully slow progress that may never reach the global minimum within practical time limits, while an excessively large value causes the loss function to bounce or diverge entirely. The book uses a simple quadratic objective **f(x) = x²** to visualize these effects, showing that stable convergence requires balancing step size against gradient magnitude.

```python
def gd(eta, f_grad):
    x = 10.0
    for _ in range(10):
        x -= eta * f_grad(x)
    return x

def f_grad(x):
    return 2 * x          # gradient of f(x)=x²

# Too small – slow progress

print(gd(0.05, f_grad))   # → ~5.0 after 10 steps

# Reasonable – converges

print(gd(0.2, f_grad))    # → ~0.0 after 10 steps

# Too large – diverges

print(gd(1.1, f_grad))    # → grows quickly, loss explodes

```

## Static vs. Dynamic Learning Rate Schedules

While constant learning rates suffice for convex problems, d2l-zh argues in [`chapter_optimization/sgd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/sgd.md) (lines 107-119) that deep learning requires **dynamic schedules** that adapt η during training. The repository implements several strategies in [`chapter_optimization/lr-scheduler.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/lr-scheduler.md) to handle non-convex loss landscapes where aggressive early steps and conservative later steps improve final model quality.

### Piecewise-Constant Decay

The `FactorScheduler` class demonstrates step-wise reduction. As shown in [`chapter_optimization/lr-scheduler.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/lr-scheduler.md) (lines 5-14), this approach multiplies the learning rate by a decay factor (e.g., 0.5) every fixed number of epochs, allowing the optimizer to take smaller, more precise steps as it approaches a minimum.

```python
from d2l import mxnet as d2l

# Define a scheduler that halves the LR every 5 epochs

lr_sched = d2l.lr_scheduler.FactorScheduler(step=5, factor=0.5, base_lr=0.3)

for epoch in range(15):
    lr = lr_sched(epoch)
    print(f'epoch {epoch+1}: lr = {lr:.4f}')

```

### Cosine Decay with Warm-Up

For deep networks, d2l-zh recommends combining an initial **warm-up** period with cosine decay. According to [`chapter_optimization/lr-scheduler.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/lr-scheduler.md) (lines 378-386), the `CosineScheduler` gradually decreases the learning rate following a cosine curve, providing smoother convergence than abrupt step changes. The warm-up phase stabilizes early training when gradients are noisy, particularly effective when using large batch sizes.

```python
from d2l import mxnet as d2l

# Warm-up for 3 epochs, then cosine decay to 0.01

lr_sched = d2l.lr_scheduler.CosineScheduler(base_lr=0.3,
                                            warmup_steps=3,
                                            total_steps=30,
                                            final_lr=0.01)

for step in range(30):
    print(f'step {step+1}: lr = {lr_sched(step):.5f}')

```

## Practical Guidelines for Learning Rate Selection

The authors provide concrete experimental workflows in [`chapter_optimization/sgd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/sgd.md) (lines 130-144). They recommend starting with a modest constant η (such as 0.1 or 0.01), monitoring the loss curve for plateaus or explosions, and implementing a schedule only after confirming the initial value produces stable gradients. The book emphasizes that no universal learning rate exists across different architectures or datasets, requiring empirical testing for each new problem.

## Learning Rate in Adaptive Optimizers

Even adaptive methods like **Adam** require careful tuning of the base learning rate. In [`chapter_optimization/adam.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/adam.md) (lines 55-68), d2l-zh demonstrates that while Adam adjusts per-parameter learning rates automatically based on gradient history, the global **η** (typically defaulting to 0.01) significantly impacts convergence speed and final model quality.

```python
from d2l import mxnet as d2l

trainer = d2l.Adam(learning_rate=0.01)   # default η = 0.01

# … training loop …

```

## Summary

- d2l-zh identifies the learning rate η as the most critical hyper-parameter for model convergence in [`chapter_optimization/gd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/gd.md), where improper values cause either slow progress or divergence.
- Static learning rates often fail for deep networks; dynamic schedules like cosine decay with warm-up in [`chapter_optimization/lr-scheduler.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/lr-scheduler.md) provide superior stability.
- The repository implements practical schedulers including `FactorScheduler` for step decay and `CosineScheduler` for smooth annealing to zero.
- Adaptive optimizers like Adam still require base learning rate tuning, typically starting at 0.01, as shown in [`chapter_optimization/adam.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/adam.md).

## Frequently Asked Questions

### What happens if the learning rate is too large in d2l-zh examples?

According to the gradient descent demonstrations in [`chapter_optimization/gd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/gd.md) (lines 47-66), an excessively large η causes the loss to oscillate or diverge rather than converge, as the parameter updates overshoot the minimum and potentially increase the loss indefinitely. The code example with `eta=1.1` demonstrates this explosive behavior explicitly.

### Does d2l-zh recommend warm-up for all deep learning models?

While not mandatory for every architecture, d2l-zh specifically recommends warm-up periods followed by cosine decay for deep networks and large batch sizes, as implemented in [`chapter_optimization/lr-scheduler.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/lr-scheduler.md) (lines 378-386). This combination prevents early training instability when gradients are noisy and learning is most fragile.

### How does d2l-zh suggest finding the optimal initial learning rate?

The text in [`chapter_optimization/sgd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/sgd.md) (lines 130-144) advises testing several η values (such as 0.1, 0.01, and 0.001) with a constant schedule first, observing whether the loss decreases smoothly or explodes. Practitioners should select the largest value that produces stable convergence before applying dynamic scheduling to refine the training process.

### Do adaptive optimizers like Adam eliminate the need for learning rate scheduling in d2l-zh?

No, the Adam implementation in [`chapter_optimization/adam.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/adam.md) (lines 55-68) still relies on a base learning rate that requires tuning, and the book demonstrates that combining Adam with cosine decay schedules often yields better convergence than using a fixed rate throughout training. The adaptive moment estimation adjusts per-parameter scales but does not replace the need for global learning rate annealing.