Vanishing Gradient Problem in Deep Learning: How d2l-zh Mitigates It

TLDR: The vanishing gradient problem occurs when backpropagated gradients become infinitesimally small in deep networks, preventing early layers from learning, and d2l-zh addresses this through ReLU activations, Xavier initialization, and LSTM gating mechanisms.

The vanishing gradient problem is a fundamental challenge in training deep neural networks, where gradients shrink exponentially as they propagate backward through many layers. In the d2l-zh repository (the Chinese edition of Dive into Deep Learning), this issue is examined across multiple chapters, from multilayer perceptrons to recurrent neural networks. This article explores the root causes of vanishing gradients as documented in d2l-zh and the specific architectural solutions implemented to maintain healthy gradient flow.

Understanding the Vanishing Gradient Problem

The vanishing gradient problem occurs when the gradients that are back-propagated through a deep network become extremely small, often close to zero. As a result, the parameters of the earlier (lower) layers receive negligible updates and the network cannot learn long-range dependencies.

Sigmoid and Tanh Activation Functions

In chapter_multilayer-perceptrons/numerical-stability-and-init.md, d2l-zh identifies sigmoid and tanh activations as primary culprits. Their derivatives are ≤ 0.25 and shrink rapidly for large-magnitude inputs. When many such layers are stacked, the product of the derivatives quickly approaches zero. As noted in the text: "当sigmoid函数的输入很大或是很小时,它的梯度都会消失… 当网络有很多层时…梯度可能会消失" (When the input to the sigmoid function is very large or very small, its gradients vanish... when the network has many layers... gradients may vanish).

Deep Linear Chains and Eigenvalue Decay

The problem extends beyond activation functions. In chapter_optimization/optimization-intro.md, d2l-zh explains that deep linear chains—repeated multiplication of weight matrices—can produce eigenvalues that are < 1, causing the norm of the gradient to shrink exponentially with depth. The text states: "高次幂的矩阵会导致特征值趋于 0 或无穷,表现为梯度消失或爆炸" (High powers of matrices cause eigenvalues to tend toward 0 or infinity, manifesting as vanishing or exploding gradients).

Improper Weight Initialization

Finally, chapter_multilayer-perceptrons/numerical-stability-and-init.md highlights poor weight initialization as a contributing factor. If the variance of the initial weights is not scaled properly, the forward activations and backward gradients shrink (or explode) across layers. The documentation warns: "使用默认的随机初始化可能导致梯度消失/爆炸" (Using default random initialization may lead to vanishing/exploding gradients).

Mitigation Strategies in d2l-zh

The d2l-zh repository implements several architectural and initialization strategies to combat vanishing gradients.

ReLU Activation Functions

The most direct solution appears in chapter_multilayer-perceptrons/mlp.md: replacing sigmoid with ReLU (Rectified Linear Unit). ReLU’s derivative is 1 for positive inputs and 0 for negatives, so positive signals propagate without attenuation. This prevents the multiplicative decay that plagues sigmoid networks.

Xavier (Glorot) Initialization

To address initialization-related vanishing, d2l-zh adopts Xavier initialization (also called Glorot initialization). As detailed in chapter_multilayer-perceptrons/numerical-stability-and-init.md, this method sets the variance of weights to 2/(n_in + n_out), keeping both forward activations and backward gradients at a healthy scale across deep stacks.

LSTM and GRU Architectures

For recurrent networks, where vanishing gradients impede learning long-range temporal dependencies, chapter_recurrent-modern/lstm.md introduces Long Short-Term Memory (LSTM) networks. The LSTM architecture introduces input, forget, and output gates that create a separate memory cell (C_t). Because the cell state is updated by element-wise multiplication rather than repeated matrix products, gradients can flow unchanged over many time steps, dramatically alleviating vanishing gradients in sequence models.

Gradient Clipping for Stability

While primarily targeting exploding gradients, chapter_recurrent-neural-networks/bptt.md mentions gradient clipping as a stabilizing technique during backpropagation through time (BPTT). This provides training stability when both vanishing and exploding gradient risks appear in deep recurrent stacks.

Practical Code Examples from d2l-zh

The following examples demonstrate how d2l-zh implements these solutions using popular deep learning frameworks.

Visualizing Vanishing Gradients: Sigmoid vs. ReLU

This MXNet example from chapter_multilayer-perceptrons/mlp.md contrasts gradient behavior:

%matplotlib inline
from d2l import mxnet as d2l
from mxnet import autograd, np, npx
npx.set_np()

x = np.arange(-8.0, 8.0, 0.1)
x.attach_grad()

# Sigmoid

with autograd.record():
    y_sig = npx.sigmoid(x)
y_sig.backward()
d2l.plot(x, [y_sig, x.grad], legend=['sigmoid', 'grad(sigmoid)'],
         figsize=(5, 2.5))

# ReLU

x.grad[:] = 0  # reset gradients

with autograd.record():
    y_relu = npx.relu(x)
y_relu.backward()
d2l.plot(x, [y_relu, x.grad], legend=['ReLU', 'grad(ReLU)'],
         figsize=(5, 2.5))

The sigmoid plot shows gradients near zero for large |x|, while ReLU maintains a constant gradient of 1 for positive x, preventing vanishing.

Implementing Xavier Initialization in PyTorch

This example from chapter_multilayer-perceptrons/numerical-stability-and-init.md applies Xavier initialization to a two-layer MLP:

import torch
from d2l import torch as d2l

def init_xavier(m):
    if isinstance(m, torch.nn.Linear):
        torch.nn.init.xavier_uniform_(m.weight)
        if m.bias is not None:
            torch.nn.init.zeros_(m.bias)

net = torch.nn.Sequential(
    torch.nn.Linear(784, 256),
    torch.nn.ReLU(),
    torch.nn.Linear(256, 10)
)
net.apply(init_xavier)          # Apply Xavier init

d2l.train_ch5(net, train_iter, test_iter,
               loss, num_epochs=5, trainer=torch.optim.SGD(net.parameters(), lr=0.1))

Xavier scaling keeps the variance of activations and gradients stable across layers, preventing the exponential shrinkage that causes vanishing gradients.

LSTM Cell Propagation and Gradient Flow

This TensorFlow example from chapter_recurrent-modern/lstm.md demonstrates how LSTM architectures maintain gradient flow:

import tensorflow as tf
from d2l import tensorflow as d2l

# Load a tiny sequence dataset (time-machine)

batch_size, num_steps = 4, 35
train_iter, vocab = d2l.load_data_time_machine(batch_size, num_steps)

lstm_layer = tf.keras.layers.LSTM(256, return_sequences=True, return_state=True)
inputs = tf.random.normal((batch_size, num_steps, len(vocab)))
with tf.GradientTape() as tape:
    outputs, h, c = lstm_layer(inputs)
    loss = tf.reduce_mean(outputs)                # dummy loss

grads = tape.gradient(loss, lstm_layer.trainable_variables)
print([g.numpy().mean() for g in grads])          # gradients stay non-zero

The LSTM's memory cell (c) enables gradients to pass unchanged across many time steps, mitigating the vanishing gradient problem in recurrent networks.

Summary

  • The vanishing gradient problem arises when backpropagated gradients become exponentially small in deep networks, preventing early layers from updating effectively.
  • Sigmoid and tanh activations are primary causes because their derivatives are bounded by 0.25, causing multiplicative decay across layers.
  • d2l-zh mitigates this through ReLU activations (constant gradient of 1), Xavier initialization (variance scaling), and LSTM architectures (gated memory cells).
  • Proper initialization and architectural choices are essential for training deep networks and recurrent models without losing the learning signal.

Frequently Asked Questions

What causes the vanishing gradient problem in deep neural networks?

The vanishing gradient problem occurs when gradients become infinitesimally small during backpropagation through many layers. According to chapter_multilayer-perceptrons/numerical-stability-and-init.md, this happens primarily because sigmoid and tanh activation functions have derivatives bounded by 0.25, and when multiplied across many layers, the gradient product approaches zero. Additionally, repeated multiplication of weight matrices with eigenvalues less than 1 causes exponential shrinkage of gradient norms.

Why does ReLU help prevent vanishing gradients?

ReLU (Rectified Linear Unit) helps prevent vanishing gradients because its derivative is 1 for all positive inputs and 0 for negative inputs, as demonstrated in chapter_multilayer-perceptrons/mlp.md. Unlike sigmoid, which saturates and produces near-zero gradients for large positive or negative inputs, ReLU allows positive signals to propagate backward through the network without attenuation. This prevents the multiplicative decay that causes gradients to vanish in deep stacks of sigmoid layers.

How does Xavier initialization solve the vanishing gradient problem?

Xavier (Glorot) initialization solves the vanishing gradient problem by scaling the variance of initial weights to 2/(n_in + n_out), as detailed in chapter_multilayer-perceptrons/numerical-stability-and-init.md. This scaling ensures that the variance of activations remains approximately constant during forward propagation, and critically, that the variance of gradients remains stable during backpropagation. Without this careful variance scaling, default random initialization causes gradients to shrink exponentially as they propagate backward through deep networks.

What role do LSTM networks play in addressing vanishing gradients?

LSTM (Long Short-Term Memory) networks address vanishing gradients in recurrent architectures by introducing a memory cell state (C_t) and gating mechanisms, as explained in chapter_recurrent-modern/lstm.md. Unlike standard RNNs where gradients must flow through repeated matrix multiplications, LSTMs update the cell state through element-wise operations and gates that can maintain a constant error flow. This architectural design allows gradients to propagate unchanged across many time steps, effectively solving the vanishing gradient problem in sequence modeling tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →