# How d2l-zh Explains the Backpropagation Algorithm for Training Neural Networks

> Learn how d2l-zh explains backpropagation using computational graphs and the chain rule for efficient neural network training. Understand gradient computation from output to input.

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: deep-dive
- Published: 2026-03-01

---

**The d2l-zh textbook explains backpropagation as an application of the chain rule on a computational graph, computing gradients from the output layer back to the input layer to enable efficient parameter updates in neural networks.**

The Chinese edition of *Dive into Deep Learning* (d2l-zh) provides a comprehensive, mathematically rigorous yet practically grounded explanation of backpropagation in **Chapter 12: "Forward Propagation, Backpropagation, and Computational Graphs"** ([`chapter_multilayer-perceptrons/backprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_multilayer-perceptrons/backprop.md)). The text bridges theoretical derivations with framework implementations across MXNet, PyTorch, TensorFlow, and Paddle.

## Forward Propagation and the Computational Graph

Before gradients can flow backward, the network must first compute the forward pass. D2l-zh establishes this foundation by defining the computational graph where each node represents a mathematical operation.

### Forward Propagation Mechanics

In [`backprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/backprop.md), the text defines forward propagation for a single-hidden-layer MLP as a sequence of matrix operations and activations:

1.  **Affine transformation**: $\mathbf{z} = W^{(1)}\mathbf{x}$
2.  **Activation**: $\mathbf{h} = \phi(\mathbf{z})$
3.  **Output layer**: $\mathbf{o} = W^{(2)}\mathbf{h}$

The implementation emphasizes that **all intermediate variables** ($\mathbf{z}$, $\mathbf{h}$) must be retained in memory because they are required for gradient computation during the backward pass.

### Objective Function and Regularization

The textbook defines the complete loss function $J$ as the sum of the empirical loss $L$ (e.g., cross-entropy or MSE) and the regularization term $s$ (typically $L_2$ penalty):

$$J = L + s \quad \text{where} \quad s = \frac{\lambda}{2} \left(\|W^{(1)}\|_F^2 + \|W^{(2)}\|_F^2\right)$$

This decomposition is critical because backpropagation must compute gradients with respect to $J$, incorporating both the data-dependent error and the weight decay terms.

## The Chain Rule and Gradient Flow

D2l-zh explains backpropagation as the systematic application of the **chain rule** to traverse the computational graph in reverse topological order.

### Mathematical Derivation of Backpropagation

The text introduces the abstract chain rule operator $\text{prod}$ to handle vector/matrix derivatives:

$$\frac{\partial Z}{\partial X} = \text{prod}\left(\frac{\partial Z}{\partial Y}, \frac{\partial Y}{\partial X}\right)$$

For the MLP example, the gradient computation follows this specific sequence:

1.  **Output layer gradient**: Compute $\frac{\partial J}{\partial \mathbf{o}}$ (derivative of loss with respect to network output)
2.  **Second weight matrix**: $\frac{\partial J}{\partial W^{(2)}} = \text{prod}\left(\frac{\partial J}{\partial \mathbf{o}}, \frac{\partial \mathbf{o}}{\partial W^{(2)}}\right) = \frac{\partial J}{\partial \mathbf{o}} \mathbf{h}^\top$
3.  **Hidden layer gradient**: $\frac{\partial J}{\partial \mathbf{h}} = W^{(2)\top} \frac{\partial J}{\partial \mathbf{o}}$
4.  **Activation derivative**: $\frac{\partial J}{\partial \mathbf{z}} = \frac{\partial J}{\partial \mathbf{h}} \odot \phi'(\mathbf{z})$
5.  **First weight matrix**: $\frac{\partial J}{\partial W^{(1)}} = \text{prod}\left(\frac{\partial J}{\partial \mathbf{z}}, \frac{\partial \mathbf{z}}{\partial W^{(1)}}\right) = \frac{\partial J}{\partial \mathbf{z}} \mathbf{x}^\top$

### Regularization Gradients

The textbook explicitly handles the regularization term $s$ by noting that its gradient with respect to weights is simply $\lambda W$. Therefore, the final gradient for each weight matrix combines the data gradient and the regularization gradient:

$$\frac{\partial J}{\partial W^{(i)}} = \frac{\partial L}{\partial W^{(i)}} + \lambda W^{(i)}$$

This detail is crucial for correct implementation of weight decay in optimization loops.

## Practical Implementation in Deep Learning Frameworks

D2l-zh transitions from mathematical derivation to code by demonstrating how modern frameworks automate the backward pass through **automatic differentiation** (autograd).

### The Standard Training Loop Pattern

Regardless of the backend (MXNet, PyTorch, TensorFlow, or Paddle), the textbook establishes a consistent training pattern:

1.  **Forward pass**: Compute predictions and loss inside an autograd context
2.  **Backward pass**: Call `.backward()` to compute gradients
3.  **Update step**: Optimizer adjusts parameters using computed gradients

This pattern appears in [`chapter_multilayer-perceptrons/backprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_multilayer-perceptrons/backprop.md) under the "Training Neural Networks" section, emphasizing that the framework handles the chain rule application internally.

### Framework-Specific Examples

The textbook provides interchangeable implementations across frameworks. Here are the PyTorch and MXNet variants:

**PyTorch Implementation** (from [`chapter_linear-networks/linear-regression-concise.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_linear-networks/linear-regression-concise.md)):

```python
import torch
from torch import nn, optim

net = nn.Sequential(nn.Linear(2, 8), nn.ReLU(), nn.Linear(8, 1))
loss = nn.MSELoss()
trainer = optim.SGD(net.parameters(), lr=0.03)

for epoch in range(3):
    for X, y in data_iter:
        trainer.zero_grad()      # Clear gradients

        l = loss(net(X), y)      # Forward

        l.backward()             # Backpropagation

        trainer.step()           # Parameter update

```

**MXNet/Gluon Implementation** (from [`chapter_linear-networks/linear-regression-concise.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_linear-networks/linear-regression-concise.md)):

```python
from mxnet import gluon, autograd, np, npx
npx.set_np()

net = gluon.nn.Sequential()
with net.name_scope():
    net.add(gluon.nn.Dense(8, activation='relu'), gluon.nn.Dense(1))
net.initialize()

loss = gluon.loss.L2Loss()
trainer = gluon.Trainer(net.collect_params(), 'sgd', {'learning_rate': 0.03})

for epoch in range(3):
    for X, y in data_iter:
        with autograd.record():      # Forward context

            l = loss(net(X), y)
        l.backward()                 # Compute gradients

        trainer.step(batch_size)     # Update parameters

```

Both examples demonstrate that **backpropagation** is invoked via a single `.backward()` call, which internally traverses the computational graph constructed during the forward pass, applying the chain rule exactly as derived in the mathematical sections.

## Summary

- **D2l-zh** explains backpropagation in [`chapter_multilayer-perceptrons/backprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_multilayer-perceptrons/backprop.md) as the application of the **chain rule** on a computational graph, computing gradients from output to input layers.
- The textbook decomposes the process into **forward propagation** (computing and storing intermediate values), **loss computation** (including regularization), and **backward propagation** (gradient calculation via chain rule).
- Mathematical derivations explicitly show how gradients flow through weight matrices ($W^{(1)}$, $W^{(2)}$), activations ($\phi$), and hidden states ($\mathbf{h}$), including the addition of regularization gradients ($\lambda W$).
- Practical implementation demonstrates that modern frameworks (MXNet, PyTorch, TensorFlow, Paddle) automate the backward pass through **automatic differentiation**, requiring only a `backward()` call after recording the forward computation.

## Frequently Asked Questions

### What is the difference between backpropagation and automatic differentiation?

**Backpropagation** specifically refers to the algorithm for computing gradients in neural networks by applying the chain rule from the output layer back to the input layer. **Automatic differentiation** (autograd) is the broader programming framework technique that implements backpropagation (and other differentiation methods) automatically. In d2l-zh, backpropagation is the mathematical concept explained in [`backprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/backprop.md), while autograd is the practical mechanism used in the code examples to execute that mathematics.

### Why does backpropagation require saving intermediate values during the forward pass?

According to the d2l-zh explanation in [`chapter_multilayer-perceptrons/backprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_multilayer-perceptrons/backprop.md), backpropagation requires intermediate values (such as $\mathbf{z}$ and $\mathbf{h}$ in a hidden layer) because the chain rule derivatives depend on them. For example, computing $\frac{\partial J}{\partial W^{(2)}}$ requires the hidden activation $\mathbf{h}^\top$, and computing gradients through the activation function requires $\phi'(\mathbf{z})$. The textbook notes that this memory requirement creates a trade-off with batch size and network depth.

### How does d2l-zh handle backpropagation across different deep learning frameworks?

D2l-zh implements a **framework-agnostic pedagogy** by teaching the mathematical principles once in [`backprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/backprop.md), then providing parallel code implementations in MXNet, PyTorch, TensorFlow, and Paddle. The training loop pattern remains consistent across frameworks: (1) record forward computation, (2) call `backward()`, (3) step the optimizer. This approach allows readers to understand that backpropagation is a universal algorithm while seeing the specific API differences (e.g., `autograd.record()` in MXNet vs. `zero_grad()` in PyTorch).

### What role does the chain rule play in the backpropagation algorithm?

In d2l-zh, the **chain rule** is the mathematical foundation that makes backpropagation possible. As detailed in [`chapter_multilayer-perceptrons/backprop.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_multilayer-perceptrons/backprop.md), the chain rule allows the decomposition of complex derivatives (like $\frac{\partial J}{\partial W^{(1)}}$) into products of simpler, local derivatives. The textbook explicitly uses the notation $\text{prod}\left(\frac{\partial Z}{\partial Y}, \frac{\partial Y}{\partial X}\right)$ to represent this multiplicative propagation of gradients through each layer of the network, enabling efficient computation from output to input.