How d2l-zh Explains the Backpropagation Algorithm for Training Neural Networks

The d2l-zh textbook explains backpropagation as an application of the chain rule on a computational graph, computing gradients from the output layer back to the input layer to enable efficient parameter updates in neural networks.

The Chinese edition of Dive into Deep Learning (d2l-zh) provides a comprehensive, mathematically rigorous yet practically grounded explanation of backpropagation in Chapter 12: "Forward Propagation, Backpropagation, and Computational Graphs" (chapter_multilayer-perceptrons/backprop.md). The text bridges theoretical derivations with framework implementations across MXNet, PyTorch, TensorFlow, and Paddle.

Forward Propagation and the Computational Graph

Before gradients can flow backward, the network must first compute the forward pass. D2l-zh establishes this foundation by defining the computational graph where each node represents a mathematical operation.

Forward Propagation Mechanics

In backprop.md, the text defines forward propagation for a single-hidden-layer MLP as a sequence of matrix operations and activations:

  1. Affine transformation: $\mathbf{z} = W^{(1)}\mathbf{x}$
  2. Activation: $\mathbf{h} = \phi(\mathbf{z})$
  3. Output layer: $\mathbf{o} = W^{(2)}\mathbf{h}$

The implementation emphasizes that all intermediate variables ($\mathbf{z}$, $\mathbf{h}$) must be retained in memory because they are required for gradient computation during the backward pass.

Objective Function and Regularization

The textbook defines the complete loss function $J$ as the sum of the empirical loss $L$ (e.g., cross-entropy or MSE) and the regularization term $s$ (typically $L_2$ penalty):

$$J = L + s \quad \text{where} \quad s = \frac{\lambda}{2} \left(|W^{(1)}|_F^2 + |W^{(2)}|_F^2\right)$$

This decomposition is critical because backpropagation must compute gradients with respect to $J$, incorporating both the data-dependent error and the weight decay terms.

The Chain Rule and Gradient Flow

D2l-zh explains backpropagation as the systematic application of the chain rule to traverse the computational graph in reverse topological order.

Mathematical Derivation of Backpropagation

The text introduces the abstract chain rule operator $\text{prod}$ to handle vector/matrix derivatives:

$$\frac{\partial Z}{\partial X} = \text{prod}\left(\frac{\partial Z}{\partial Y}, \frac{\partial Y}{\partial X}\right)$$

For the MLP example, the gradient computation follows this specific sequence:

  1. Output layer gradient: Compute $\frac{\partial J}{\partial \mathbf{o}}$ (derivative of loss with respect to network output)
  2. Second weight matrix: $\frac{\partial J}{\partial W^{(2)}} = \text{prod}\left(\frac{\partial J}{\partial \mathbf{o}}, \frac{\partial \mathbf{o}}{\partial W^{(2)}}\right) = \frac{\partial J}{\partial \mathbf{o}} \mathbf{h}^\top$
  3. Hidden layer gradient: $\frac{\partial J}{\partial \mathbf{h}} = W^{(2)\top} \frac{\partial J}{\partial \mathbf{o}}$
  4. Activation derivative: $\frac{\partial J}{\partial \mathbf{z}} = \frac{\partial J}{\partial \mathbf{h}} \odot \phi'(\mathbf{z})$
  5. First weight matrix: $\frac{\partial J}{\partial W^{(1)}} = \text{prod}\left(\frac{\partial J}{\partial \mathbf{z}}, \frac{\partial \mathbf{z}}{\partial W^{(1)}}\right) = \frac{\partial J}{\partial \mathbf{z}} \mathbf{x}^\top$

Regularization Gradients

The textbook explicitly handles the regularization term $s$ by noting that its gradient with respect to weights is simply $\lambda W$. Therefore, the final gradient for each weight matrix combines the data gradient and the regularization gradient:

$$\frac{\partial J}{\partial W^{(i)}} = \frac{\partial L}{\partial W^{(i)}} + \lambda W^{(i)}$$

This detail is crucial for correct implementation of weight decay in optimization loops.

Practical Implementation in Deep Learning Frameworks

D2l-zh transitions from mathematical derivation to code by demonstrating how modern frameworks automate the backward pass through automatic differentiation (autograd).

The Standard Training Loop Pattern

Regardless of the backend (MXNet, PyTorch, TensorFlow, or Paddle), the textbook establishes a consistent training pattern:

  1. Forward pass: Compute predictions and loss inside an autograd context
  2. Backward pass: Call .backward() to compute gradients
  3. Update step: Optimizer adjusts parameters using computed gradients

This pattern appears in chapter_multilayer-perceptrons/backprop.md under the "Training Neural Networks" section, emphasizing that the framework handles the chain rule application internally.

Framework-Specific Examples

The textbook provides interchangeable implementations across frameworks. Here are the PyTorch and MXNet variants:

PyTorch Implementation (from chapter_linear-networks/linear-regression-concise.md):

import torch
from torch import nn, optim

net = nn.Sequential(nn.Linear(2, 8), nn.ReLU(), nn.Linear(8, 1))
loss = nn.MSELoss()
trainer = optim.SGD(net.parameters(), lr=0.03)

for epoch in range(3):
    for X, y in data_iter:
        trainer.zero_grad()      # Clear gradients

        l = loss(net(X), y)      # Forward

        l.backward()             # Backpropagation

        trainer.step()           # Parameter update

MXNet/Gluon Implementation (from chapter_linear-networks/linear-regression-concise.md):

from mxnet import gluon, autograd, np, npx
npx.set_np()

net = gluon.nn.Sequential()
with net.name_scope():
    net.add(gluon.nn.Dense(8, activation='relu'), gluon.nn.Dense(1))
net.initialize()

loss = gluon.loss.L2Loss()
trainer = gluon.Trainer(net.collect_params(), 'sgd', {'learning_rate': 0.03})

for epoch in range(3):
    for X, y in data_iter:
        with autograd.record():      # Forward context

            l = loss(net(X), y)
        l.backward()                 # Compute gradients

        trainer.step(batch_size)     # Update parameters

Both examples demonstrate that backpropagation is invoked via a single .backward() call, which internally traverses the computational graph constructed during the forward pass, applying the chain rule exactly as derived in the mathematical sections.

Summary

  • D2l-zh explains backpropagation in chapter_multilayer-perceptrons/backprop.md as the application of the chain rule on a computational graph, computing gradients from output to input layers.
  • The textbook decomposes the process into forward propagation (computing and storing intermediate values), loss computation (including regularization), and backward propagation (gradient calculation via chain rule).
  • Mathematical derivations explicitly show how gradients flow through weight matrices ($W^{(1)}$, $W^{(2)}$), activations ($\phi$), and hidden states ($\mathbf{h}$), including the addition of regularization gradients ($\lambda W$).
  • Practical implementation demonstrates that modern frameworks (MXNet, PyTorch, TensorFlow, Paddle) automate the backward pass through automatic differentiation, requiring only a backward() call after recording the forward computation.

Frequently Asked Questions

What is the difference between backpropagation and automatic differentiation?

Backpropagation specifically refers to the algorithm for computing gradients in neural networks by applying the chain rule from the output layer back to the input layer. Automatic differentiation (autograd) is the broader programming framework technique that implements backpropagation (and other differentiation methods) automatically. In d2l-zh, backpropagation is the mathematical concept explained in backprop.md, while autograd is the practical mechanism used in the code examples to execute that mathematics.

Why does backpropagation require saving intermediate values during the forward pass?

According to the d2l-zh explanation in chapter_multilayer-perceptrons/backprop.md, backpropagation requires intermediate values (such as $\mathbf{z}$ and $\mathbf{h}$ in a hidden layer) because the chain rule derivatives depend on them. For example, computing $\frac{\partial J}{\partial W^{(2)}}$ requires the hidden activation $\mathbf{h}^\top$, and computing gradients through the activation function requires $\phi'(\mathbf{z})$. The textbook notes that this memory requirement creates a trade-off with batch size and network depth.

How does d2l-zh handle backpropagation across different deep learning frameworks?

D2l-zh implements a framework-agnostic pedagogy by teaching the mathematical principles once in backprop.md, then providing parallel code implementations in MXNet, PyTorch, TensorFlow, and Paddle. The training loop pattern remains consistent across frameworks: (1) record forward computation, (2) call backward(), (3) step the optimizer. This approach allows readers to understand that backpropagation is a universal algorithm while seeing the specific API differences (e.g., autograd.record() in MXNet vs. zero_grad() in PyTorch).

What role does the chain rule play in the backpropagation algorithm?

In d2l-zh, the chain rule is the mathematical foundation that makes backpropagation possible. As detailed in chapter_multilayer-perceptrons/backprop.md, the chain rule allows the decomposition of complex derivatives (like $\frac{\partial J}{\partial W^{(1)}}$) into products of simpler, local derivatives. The textbook explicitly uses the notation $\text{prod}\left(\frac{\partial Z}{\partial Y}, \frac{\partial Y}{\partial X}\right)$ to represent this multiplicative propagation of gradients through each layer of the network, enabling efficient computation from output to input.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →