# How d2l-zh Explains Optimization Strategies for Deep Learning Models: From Gradient Descent to Adam

> Explore deep learning optimization strategies from Gradient Descent to Adam with d2l-zh. Learn theory and code in chapter_optimization.

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: deep-dive
- Published: 2026-03-01

---

**The d2l-zh repository systematically teaches optimization strategies for deep learning models through a pedagogical progression from basic gradient descent to adaptive Adam optimizers, combining theoretical analysis with reproducible code in `chapter_optimization/`.**

The **d2l-zh** (Dive into Deep Learning Chinese edition) textbook treats optimization as the fundamental engine driving deep learning training. In [[`chapter_optimization/index.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/index.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/index.md), the authors establish minimizing the **objective function** (loss) on the training set as the core goal. The chapter then presents a family of increasingly sophisticated algorithms that mirror the historical development of modern deep learning, backed by theoretical convergence analysis and visual experiments using a canonical 2-D quadratic objective function.

## The Foundation: Gradient Descent and Stochastic Methods

The optimization narrative in d2l-zh begins with classical gradient-based methods before introducing the stochastic approximations that make large-scale deep learning feasible.

### Full-Batch Gradient Descent in [`gd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/gd.md)

According to the source code in [[`chapter_optimization/gd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/gd.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/gd.md), **Gradient Descent (GD)** computes the full gradient of the loss with respect to all parameters and performs the update **x ← x – η ∇f(x)**. The book emphasizes that this requires careful selection of a single learning rate η and demonstrates the algorithm using the toy objective **f(x₁,x₂)=0.1x₁²+2x₂²**. While conceptually simple, full-batch GD incurs **O(n)** computational cost per iteration, where n is the training set size, making it impractical for modern deep learning datasets.

```python
def gd(x1, x2, s1, s2):
    eta = 0.4
    x1 -= eta * 0.2 * x1          # ∂f/∂x1 = 0.2·x1

    x2 -= eta * 4   * x2          # ∂f/∂x2 = 4·x2

    return x1, x2, 0, 0           # no state needed

```

### Stochastic Gradient Descent and Mini-Batches in [`sgd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/sgd.md)

The transition to **Stochastic Gradient Descent (SGD)** in [[`chapter_optimization/sgd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/sgd.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/sgd.md) replaces the full gradient with an estimate from a single randomly sampled example: **x ← x – η ∇f_i(x)**. This reduces per-iteration cost from **O(n)** to **O(1)**. The book explains that the introduced noise actually helps escape shallow local minima. The text further refines this with **Mini-batch SGD**, which uses a small batch of examples to compute a less noisy gradient estimate, balancing variance reduction with parallel computational efficiency.

```python
import math, random

def sgd(x1, x2, s1, s2):
    eta = 0.4
    # noisy gradient ≈ true gradient + N(0,1) noise

    g1 = 0.2 * x1 + random.gauss(0, 1)
    g2 = 4   * x2 + random.gauss(0, 1)
    x1 -= eta * g1
    x2 -= eta * g2
    return x1, x2, 0, 0

```

### Learning Rate Schedules

D2l-zh addresses the limitation of constant learning rates through **learning rate schedules** detailed in the *Dynamic Learning Rate* section of [[`sgd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/sgd.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/sgd.md) and further expanded in [[`lr-scheduler.md`](https://github.com/d2l-ai/d2l-zh/blob/main/lr-scheduler.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/lr-scheduler.md). The book presents time-dependent schedules η(t) including piecewise constant, exponential decay, and polynomial decay (η ∝ t⁻¹ᐟ²), explaining how these "cool down" the optimizer during training to ensure convergence.

## Accelerating Convergence: Momentum Methods

### The Momentum Algorithm in [`momentum.md`](https://github.com/d2l-ai/d2l-zh/blob/main/momentum.md)

The [[`chapter_optimization/momentum.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/momentum.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/momentum.md) file introduces **momentum** as a strategy to accelerate progress along long, shallow directions while dampening oscillations. The algorithm maintains a velocity vector updated as **v_t = β v_{t-1} + g_t** where β ∈ (0,1), then applies **x ← x – η v_t**. By accumulating past gradients, momentum smooths noisy gradient estimates and helps navigate ravines in the loss landscape common in deep networks.

```python
beta = 0.9
v1 = v2 = 0.0

def momentum(x1, x2, v1, v2):
    eta = 0.6
    g1 = 0.2 * x1
    g2 = 4   * x2
    v1 = beta * v1 + g1
    v2 = beta * v2 + g2
    x1 -= eta * v1
    x2 -= eta * v2
    return x1, x2, v1, v2

```

## Adaptive Learning Rate Strategies

D2l-zh dedicates significant attention to **adaptive optimization algorithms** that automatically adjust learning rates per parameter or per coordinate, addressing the challenge of ill-conditioned optimization problems and sparse data.

### AdaGrad for Sparse Features in [`adagrad.md`](https://github.com/d2l-ai/d2l-zh/blob/main/adagrad.md)

The [[`chapter_optimization/adagrad.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/adagrad.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/adagrad.md) implementation adapts the learning rate per coordinate using the sum of squared past gradients: **s_t = s_{t-1} + g_t²**. The update rule becomes **x ← x – η g_t / √(s_t+ε)**. The book highlights that AdaGrad works particularly well for sparse features because rarely-seen dimensions maintain a larger effective step size, while frequently-updated dimensions see their learning rates decay rapidly.

```python
eps = 1e-6
s1 = s2 = 0.0

def adagrad(x1, x2, s1, s2):
    eta = 0.4
    g1 = 0.2 * x1
    g2 = 4   * x2
    s1 += g1 ** 2
    s2 += g2 ** 2
    x1 -= eta / math.sqrt(s1 + eps) * g1
    x2 -= eta / math.sqrt(s2 + eps) * g2
    return x1, x2, s1, s2

```

### RMSProp for Non-Convex Optimization in [`rmsprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/rmsprop.md)

As implemented in [[`chapter_optimization/rmsprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/rmsprop.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/rmsprop.md), **RMSProp** modifies AdaGrad by using an **exponentially weighted moving average (EWMA)** of squared gradients: **s_t = β s_{t-1} + (1-β) g_t²**. This prevents the denominator from growing indefinitely, making the method suitable for non-convex deep learning objectives where AdaGrad's aggressive learning rate decay would stall training prematurely.

```python
beta = 0.9
s1 = s2 = 0.0

def rmsprop(x1, x2, s1, s2):
    eta = 0.01
    eps = 1e-6
    g1 = 0.2 * x1
    g2 = 4   * x2
    s1 = beta * s1 + (1 - beta) * g1 ** 2
    s2 = beta * s2 + (1 - beta) * g2 ** 2
    x1 -= eta / math.sqrt(s1 + eps) * g1
    x2 -= eta / math.sqrt(s2 + eps) * g2
    return x1, x2, s1, s2

```

### Adam and Robust Variants in [`adam.md`](https://github.com/d2l-ai/d2l-zh/blob/main/adam.md)

The [[`chapter_optimization/adam.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/adam.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/adam.md) file presents **Adam** (Adaptive Moment Estimation) as the combination of momentum (first-moment estimate) and RMSProp (second-moment estimate) with bias correction. The update rules are **v_t ← β₁v_{t-1}+(1-β₁)g_t** and **s_t ← β₂s_{t-1}+(1-β₂)g_t²**, followed by bias-corrected estimates **v̂_t** and **ŝ_t**, and finally **x ← x – η v̂_t / (√ŝ_t+ε)**. D2l-zh presents Adam as the default optimizer for many deep learning tasks because it adapts both direction and scale automatically. The book also covers the **Yogi** variant, which modifies the second-moment update with a sign-based correction to reduce the risk of exploding s_t when Adam diverges.

```python
beta1, beta2 = 0.9, 0.999
v1 = v2 = s1 = s2 = 0.0
t = 1

def adam(x1, x2, v1, v2, s1, s2):
    global t
    eta = 0.01
    eps = 1e-6
    g1 = 0.2 * x1
    g2 = 4   * x2
    v1 = beta1 * v1 + (1 - beta1) * g1
    v2 = beta1 * v2 + (1 - beta1) * g2
    s1 = beta2 * s1 + (1 - beta2) * g1 ** 2
    s2 = beta2 * s2 + (1 - beta2) * g2 ** 2
    v1_hat = v1 / (1 - beta1 ** t)
    v2_hat = v2 / (1 - beta1 ** t)
    s1_hat = s1 / (1 - beta2 ** t)
    s2_hat = s2 / (1 - beta2 ** t)
    x1 -= eta * v1_hat / (math.sqrt(s1_hat) + eps)
    x2 -= eta * v2_hat / (math.sqrt(s2_hat) + eps)
    t += 1
    return x1, x2, v1, v2, s1, s2

```

## Theoretical Foundations and Visual Validation

D2l-zh reinforces these algorithms with theoretical analysis found in [[`convexity.md`](https://github.com/d2l-ai/d2l-zh/blob/main/convexity.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/convexity.md), covering convex convergence properties and eigenvalue perspectives that explain why certain methods succeed or fail. All code examples use the toy objective **f(x₁,x₂)=0.1x₁²+2x₂²**, allowing readers to visualize optimizer trajectories using `d2l.show_trace_2d`:

```python
def f_2d(x1, x2):
    return 0.1 * x1 ** 2 + 2 * x2 ** 2

# Example usage with any trainer function:

d2l.show_trace_2d(f_2d, d2l.train_2d(momentum))

```

## Summary

- D2l-zh presents optimization as a progression from **O(n) full-batch gradient descent** to **O(1) stochastic methods**, then to variance-reduction techniques and adaptive learning rates.
- The repository explains **why** each strategy matters: SGD introduces beneficial noise, momentum accelerates along flat directions, and adaptive methods (AdaGrad, RMSProp, Adam) handle sparse data and non-convex landscapes.
- Every algorithm is implemented from scratch in the `chapter_optimization/` directory with consistent API patterns and visual debugging tools.
- The **Adam optimizer** is positioned as the default modern choice, with the **Yogi variant** offered as a robust alternative for unstable training scenarios.

## Frequently Asked Questions

### What is the difference between batch gradient descent and SGD in d2l-zh?

According to the source code in [[`gd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/gd.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/gd.md) and [[`sgd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/sgd.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/sgd.md), batch gradient descent computes the full gradient over all n training examples (cost **O(n)**), while SGD approximates the gradient using a single random example (cost **O(1)**). The book notes that SGD's gradient noise helps escape shallow local minima, making it essential for large-scale deep learning despite the approximation error.

### Why does Adam use bias correction?

As explained in [[`adam.md`](https://github.com/d2l-ai/d2l-zh/blob/main/adam.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/adam.md), Adam initializes the first and second moment estimates (v and s) to zero. Without bias correction, these estimates are biased toward zero during early iterations, particularly with high β values. The correction terms **v̂ = v / (1-β^t)** and **ŝ = s / (1-β^t)** compensate for this initialization, ensuring accurate early steps in training.

### When should I use RMSProp instead of AdaGrad?

D2l-zh recommends **RMSProp** (implemented in [[`rmsprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/rmsprop.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/rmsprop.md)) over **AdaGrad** (from [[`adagrad.md`](https://github.com/d2l-ai/d2l-zh/blob/main/adagrad.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/adagrad.md)) for deep neural networks with non-convex loss surfaces. While AdaGrad's aggressive learning rate decay suits sparse convex problems, its denominator growth causes premature convergence stalling in deep learning. RMSProp's exponentially weighted moving average prevents this indefinite growth, maintaining stable learning rates throughout training.

### How does momentum help optimization in deep learning?

The [[`momentum.md`](https://github.com/d2l-ai/d2l-zh/blob/main/momentum.md)](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/momentum.md) file explains that momentum accumulates past gradients into a velocity vector, effectively smoothing noisy stochastic gradients. This accumulation accelerates progress along directions of consistent gradient sign (such as long, shallow ravines) while dampening oscillations in directions of alternating sign, allowing the optimizer to navigate ill-conditioned loss landscapes common in high-dimensional deep networks.