How d2l-zh Explains Optimization Strategies for Deep Learning Models: From Gradient Descent to Adam

The d2l-zh repository systematically teaches optimization strategies for deep learning models through a pedagogical progression from basic gradient descent to adaptive Adam optimizers, combining theoretical analysis with reproducible code in chapter_optimization/.

The d2l-zh (Dive into Deep Learning Chinese edition) textbook treats optimization as the fundamental engine driving deep learning training. In [chapter_optimization/index.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/index.md), the authors establish minimizing the objective function (loss) on the training set as the core goal. The chapter then presents a family of increasingly sophisticated algorithms that mirror the historical development of modern deep learning, backed by theoretical convergence analysis and visual experiments using a canonical 2-D quadratic objective function.

The Foundation: Gradient Descent and Stochastic Methods

The optimization narrative in d2l-zh begins with classical gradient-based methods before introducing the stochastic approximations that make large-scale deep learning feasible.

Full-Batch Gradient Descent in gd.md

According to the source code in [chapter_optimization/gd.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/gd.md), Gradient Descent (GD) computes the full gradient of the loss with respect to all parameters and performs the update x ← x – η ∇f(x). The book emphasizes that this requires careful selection of a single learning rate η and demonstrates the algorithm using the toy objective f(x₁,x₂)=0.1x₁²+2x₂². While conceptually simple, full-batch GD incurs O(n) computational cost per iteration, where n is the training set size, making it impractical for modern deep learning datasets.

def gd(x1, x2, s1, s2):
    eta = 0.4
    x1 -= eta * 0.2 * x1          # ∂f/∂x1 = 0.2·x1

    x2 -= eta * 4   * x2          # ∂f/∂x2 = 4·x2

    return x1, x2, 0, 0           # no state needed

Stochastic Gradient Descent and Mini-Batches in sgd.md

The transition to Stochastic Gradient Descent (SGD) in [chapter_optimization/sgd.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/sgd.md) replaces the full gradient with an estimate from a single randomly sampled example: x ← x – η ∇f_i(x). This reduces per-iteration cost from O(n) to O(1). The book explains that the introduced noise actually helps escape shallow local minima. The text further refines this with Mini-batch SGD, which uses a small batch of examples to compute a less noisy gradient estimate, balancing variance reduction with parallel computational efficiency.

import math, random

def sgd(x1, x2, s1, s2):
    eta = 0.4
    # noisy gradient ≈ true gradient + N(0,1) noise

    g1 = 0.2 * x1 + random.gauss(0, 1)
    g2 = 4   * x2 + random.gauss(0, 1)
    x1 -= eta * g1
    x2 -= eta * g2
    return x1, x2, 0, 0

Learning Rate Schedules

D2l-zh addresses the limitation of constant learning rates through learning rate schedules detailed in the Dynamic Learning Rate section of [sgd.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/sgd.md) and further expanded in [lr-scheduler.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/lr-scheduler.md). The book presents time-dependent schedules η(t) including piecewise constant, exponential decay, and polynomial decay (η ∝ t⁻¹ᐟ²), explaining how these "cool down" the optimizer during training to ensure convergence.

Accelerating Convergence: Momentum Methods

The Momentum Algorithm in momentum.md

The [chapter_optimization/momentum.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/momentum.md) file introduces momentum as a strategy to accelerate progress along long, shallow directions while dampening oscillations. The algorithm maintains a velocity vector updated as v_t = β v_{t-1} + g_t where β ∈ (0,1), then applies x ← x – η v_t. By accumulating past gradients, momentum smooths noisy gradient estimates and helps navigate ravines in the loss landscape common in deep networks.

beta = 0.9
v1 = v2 = 0.0

def momentum(x1, x2, v1, v2):
    eta = 0.6
    g1 = 0.2 * x1
    g2 = 4   * x2
    v1 = beta * v1 + g1
    v2 = beta * v2 + g2
    x1 -= eta * v1
    x2 -= eta * v2
    return x1, x2, v1, v2

Adaptive Learning Rate Strategies

D2l-zh dedicates significant attention to adaptive optimization algorithms that automatically adjust learning rates per parameter or per coordinate, addressing the challenge of ill-conditioned optimization problems and sparse data.

AdaGrad for Sparse Features in adagrad.md

The [chapter_optimization/adagrad.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/adagrad.md) implementation adapts the learning rate per coordinate using the sum of squared past gradients: s_t = s_{t-1} + g_t². The update rule becomes x ← x – η g_t / √(s_t+ε). The book highlights that AdaGrad works particularly well for sparse features because rarely-seen dimensions maintain a larger effective step size, while frequently-updated dimensions see their learning rates decay rapidly.

eps = 1e-6
s1 = s2 = 0.0

def adagrad(x1, x2, s1, s2):
    eta = 0.4
    g1 = 0.2 * x1
    g2 = 4   * x2
    s1 += g1 ** 2
    s2 += g2 ** 2
    x1 -= eta / math.sqrt(s1 + eps) * g1
    x2 -= eta / math.sqrt(s2 + eps) * g2
    return x1, x2, s1, s2

RMSProp for Non-Convex Optimization in rmsprop.md

As implemented in [chapter_optimization/rmsprop.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/rmsprop.md), RMSProp modifies AdaGrad by using an exponentially weighted moving average (EWMA) of squared gradients: s_t = β s_{t-1} + (1-β) g_t². This prevents the denominator from growing indefinitely, making the method suitable for non-convex deep learning objectives where AdaGrad's aggressive learning rate decay would stall training prematurely.

beta = 0.9
s1 = s2 = 0.0

def rmsprop(x1, x2, s1, s2):
    eta = 0.01
    eps = 1e-6
    g1 = 0.2 * x1
    g2 = 4   * x2
    s1 = beta * s1 + (1 - beta) * g1 ** 2
    s2 = beta * s2 + (1 - beta) * g2 ** 2
    x1 -= eta / math.sqrt(s1 + eps) * g1
    x2 -= eta / math.sqrt(s2 + eps) * g2
    return x1, x2, s1, s2

Adam and Robust Variants in adam.md

The [chapter_optimization/adam.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/adam.md) file presents Adam (Adaptive Moment Estimation) as the combination of momentum (first-moment estimate) and RMSProp (second-moment estimate) with bias correction. The update rules are v_t ← β₁v_{t-1}+(1-β₁)g_t and s_t ← β₂s_{t-1}+(1-β₂)g_t², followed by bias-corrected estimates v̂_t and ŝ_t, and finally x ← x – η v̂_t / (√ŝ_t+ε). D2l-zh presents Adam as the default optimizer for many deep learning tasks because it adapts both direction and scale automatically. The book also covers the Yogi variant, which modifies the second-moment update with a sign-based correction to reduce the risk of exploding s_t when Adam diverges.

beta1, beta2 = 0.9, 0.999
v1 = v2 = s1 = s2 = 0.0
t = 1

def adam(x1, x2, v1, v2, s1, s2):
    global t
    eta = 0.01
    eps = 1e-6
    g1 = 0.2 * x1
    g2 = 4   * x2
    v1 = beta1 * v1 + (1 - beta1) * g1
    v2 = beta1 * v2 + (1 - beta1) * g2
    s1 = beta2 * s1 + (1 - beta2) * g1 ** 2
    s2 = beta2 * s2 + (1 - beta2) * g2 ** 2
    v1_hat = v1 / (1 - beta1 ** t)
    v2_hat = v2 / (1 - beta1 ** t)
    s1_hat = s1 / (1 - beta2 ** t)
    s2_hat = s2 / (1 - beta2 ** t)
    x1 -= eta * v1_hat / (math.sqrt(s1_hat) + eps)
    x2 -= eta * v2_hat / (math.sqrt(s2_hat) + eps)
    t += 1
    return x1, x2, v1, v2, s1, s2

Theoretical Foundations and Visual Validation

D2l-zh reinforces these algorithms with theoretical analysis found in [convexity.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/convexity.md), covering convex convergence properties and eigenvalue perspectives that explain why certain methods succeed or fail. All code examples use the toy objective f(x₁,x₂)=0.1x₁²+2x₂², allowing readers to visualize optimizer trajectories using d2l.show_trace_2d:

def f_2d(x1, x2):
    return 0.1 * x1 ** 2 + 2 * x2 ** 2

# Example usage with any trainer function:

d2l.show_trace_2d(f_2d, d2l.train_2d(momentum))

Summary

  • D2l-zh presents optimization as a progression from O(n) full-batch gradient descent to O(1) stochastic methods, then to variance-reduction techniques and adaptive learning rates.
  • The repository explains why each strategy matters: SGD introduces beneficial noise, momentum accelerates along flat directions, and adaptive methods (AdaGrad, RMSProp, Adam) handle sparse data and non-convex landscapes.
  • Every algorithm is implemented from scratch in the chapter_optimization/ directory with consistent API patterns and visual debugging tools.
  • The Adam optimizer is positioned as the default modern choice, with the Yogi variant offered as a robust alternative for unstable training scenarios.

Frequently Asked Questions

What is the difference between batch gradient descent and SGD in d2l-zh?

According to the source code in [gd.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/gd.md) and [sgd.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/sgd.md), batch gradient descent computes the full gradient over all n training examples (cost O(n)), while SGD approximates the gradient using a single random example (cost O(1)). The book notes that SGD's gradient noise helps escape shallow local minima, making it essential for large-scale deep learning despite the approximation error.

Why does Adam use bias correction?

As explained in [adam.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/adam.md), Adam initializes the first and second moment estimates (v and s) to zero. Without bias correction, these estimates are biased toward zero during early iterations, particularly with high β values. The correction terms v̂ = v / (1-β^t) and ŝ = s / (1-β^t) compensate for this initialization, ensuring accurate early steps in training.

When should I use RMSProp instead of AdaGrad?

D2l-zh recommends RMSProp (implemented in [rmsprop.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/rmsprop.md)) over AdaGrad (from [adagrad.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/adagrad.md)) for deep neural networks with non-convex loss surfaces. While AdaGrad's aggressive learning rate decay suits sparse convex problems, its denominator growth causes premature convergence stalling in deep learning. RMSProp's exponentially weighted moving average prevents this indefinite growth, maintaining stable learning rates throughout training.

How does momentum help optimization in deep learning?

The [momentum.md](https://github.com/d2l-ai/d2l-zh/blob/master/chapter_optimization/momentum.md) file explains that momentum accumulates past gradients into a velocity vector, effectively smoothing noisy stochastic gradients. This accumulation accelerates progress along directions of consistent gradient sign (such as long, shallow ravines) while dampening oscillations in directions of alternating sign, allowing the optimizer to navigate ill-conditioned loss landscapes common in high-dimensional deep networks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →