Preventing Overfitting in Deep Learning Models: Techniques from d2l-zh

The d2l-zh repository addresses preventing overfitting in deep learning models through three core techniques: controlling model capacity, applying L₂ regularization (weight decay), and implementing dropout regularization.

The Chinese edition of Dive into Deep Learning (d2l-zh) provides a comprehensive treatment of preventing overfitting in deep learning models within its multilayer perceptrons chapter. The tutorial moves from theoretical foundations to practical implementations, offering both mathematical intuition and runnable code examples across major frameworks.

Understanding Overfitting and Model Capacity

The foundation for preventing overfitting in deep learning models begins with understanding the bias-variance trade-off. In chapter_multilayer-perceptrons/underfit-overfit.md, d2l-zh defines overfitting (过拟合) as the phenomenon where a model fits the training data closer than the underlying true distribution, resulting in training error substantially lower than validation error.

The text emphasizes that model capacity determines the range of functions a model can fit. High-capacity models (e.g., high-degree polynomials or deep networks with many parameters) can memorize training noise, while insufficient capacity leads to underfitting. The tutorial demonstrates this empirically by training polynomial models of varying degrees on synthetic data, showing that validation error first decreases then increases as model complexity grows—a classic signature of overfitting.

L₂ Regularization and Weight Decay

The Mathematics of Weight Decay

The primary mechanism for preventing overfitting in deep learning models discussed in d2l-zh is L₂ regularization, commonly called weight decay. According to chapter_multilayer-perceptrons/weight-decay.md, this technique adds a penalty term $\lambda |\mathbf{w}|^2 / 2$ to the loss function, where $\lambda$ controls regularization strength.

This penalty drives the optimizer to minimize both the original loss and the squared magnitude of weights. During gradient descent, the weight update becomes $\mathbf{w} \leftarrow (1-\eta\lambda) \mathbf{w} - \frac{\eta}{|\mathcal{B}|} \sum_{i \in \mathcal{B}} \partial_{\mathbf{w}} l^{(i)}$, effectively decaying weights by factor $(1-\eta\lambda)$ at each step.

Manual Implementation in MXNet

D2l-zh provides explicit implementations showing how weight decay prevents overfitting in deep learning models. The manual implementation demonstrates adding the L₂ penalty directly inside the training loop:

from d2l import mxnet as d2l
from mxnet import autograd, gluon, np, npx
npx.set_np()

def l2_penalty(w):
    return (w**2).sum() / 2

def train(lambd):
    w, b = np.random.normal(scale=1, size=(num_inputs, 1)), np.zeros(1)
    w.attach_grad(); b.attach_grad()
    loss = gluon.loss.L2Loss()
    trainer = gluon.Trainer([w, b], 'sgd', {'learning_rate': 0.003})
    
    for epoch in range(100):
        for X, y in train_iter:
            with autograd.record():
                # L2 penalty added to loss

                l = loss(np.dot(X, w) + b, y) + lambd * l2_penalty(w)
            l.backward()
            trainer.step(batch_size=5)
    print('final L2 norm of w:', np.linalg.norm(w))

Setting lambd=0 yields large weight norms and overfitting, while lambd=3 demonstrates how weight decay keeps weights small and improves generalization.

Dropout Regularization

Stochastic Neuron Dropping

Beyond weight decay, d2l-zh presents dropout as a stochastic approach to preventing overfitting in deep learning models. According to chapter_multilayer-perceptrons/dropout.md, dropout randomly zeros out hidden unit activations during training with probability $p$, effectively creating an ensemble of thinned networks within a single model.

The tutorial explains inverted dropout, where surviving neurons are scaled by $1/(1-p)$ during training to maintain expected activation magnitudes. This ensures that at test time, when dropout is disabled, the network uses its full capacity without requiring output scaling.

Low-Level PyTorch Implementation

D2l-zh provides a from-scratch implementation clarifying dropout mechanics:

import torch

def dropout_layer(X, dropout):
    assert 0 <= dropout <= 1
    if dropout == 1:
        return torch.zeros_like(X)
    if dropout == 0:
        return X
    mask = (torch.rand(X.shape) > dropout).float()
    return mask * X / (1.0 - dropout)

This implementation demonstrates the unbiased scaling (/(1.0 - dropout)) that keeps expected values consistent regardless of dropout rate.

High-Level API Usage

For production use, d2l-zh demonstrates framework-native dropout layers. In MXNet:

from mxnet.gluon import nn

net = nn.Sequential()
net.add(
    nn.Dense(256, activation='relu'),
    nn.Dropout(0.2),  # 20% dropout after first hidden layer

    nn.Dense(256, activation='relu'),
    nn.Dropout(0.5),  # 50% dropout after second hidden layer

    nn.Dense(10)
)
net.initialize()

The tutorial emphasizes that dropout is only active during training (training=True), automatically disabled during evaluation, and typically combined with weight decay for maximum regularization effect.

Combining Techniques for Robust Regularization

D2l-zh emphasizes that preventing overfitting in deep learning models rarely relies on a single technique. The tutorial demonstrates that weight decay and dropout are frequently combined: weight decay constrains the overall magnitude of parameters, while dropout breaks co-adaptation between neurons by injecting stochastic noise.

The text also notes that while early stopping (halting training when validation error plateaus) is another valid technique, the Chinese edition focuses primarily on explicit regularization methods that modify the loss function or network architecture rather than training procedures.

Summary

  • Model capacity control forms the foundation for preventing overfitting in deep learning models, requiring balance between underfitting and overfitting through appropriate architecture choices.
  • L₂ regularization (weight decay) adds $\lambda |\mathbf{w}|^2/2$ to the loss function, shrinking weights continuously during optimization to reduce model complexity without feature selection.
  • Dropout randomly zeros hidden units during training with inverted scaling ($1/(1-p)$), creating an implicit ensemble that breaks neuron co-adaptation and is only active during the training phase.
  • Combined approaches using both weight decay and dropout provide stronger regularization than either technique alone, as demonstrated in the d2l-zh multilayer perceptron chapters.

Frequently Asked Questions

What is the difference between weight decay and L₂ regularization?

Weight decay and L₂ regularization are mathematically equivalent in standard SGD, but differ in implementation details. L₂ regularization explicitly adds the penalty term $\lambda |\mathbf{w}|^2/2$ to the loss function, while weight decay decouples the regularization from the loss calculation by directly multiplying weights by $(1-\eta\lambda)$ during the update step. In modern optimizers like Adam, the distinction becomes important because decoupled weight decay (AdamW) often outperforms standard L₂ regularization.

When should I use dropout versus weight decay?

Use weight decay as a baseline regularizer for all parametric models, as it prevents any single weight from growing too large without complicating the forward pass. Add dropout when your network shows signs of co-adaptation—when neurons rely heavily on specific other neurons rather than learning robust features independently. Dropout is particularly effective for large fully-connected layers in multilayer perceptrons and convolutional networks, but is less common in batch-normalized architectures or recurrent layers where it requires specialized variants.

Why does dropout use inverted scaling (dividing by 1-p)?

Inverted dropout scales surviving activations by $1/(1-p)$ during training to maintain the expected value of each neuron's output. Without this scaling, the expected output during training would be $(1-p) \cdot h$, but at test time (when dropout is disabled), the output would be $h$, causing a distribution shift. By scaling during training, the network sees the same magnitude of signals during training and inference, eliminating the need to scale outputs at test time and ensuring consistent behavior between phases.

Can I use both weight decay and dropout together?

Yes, weight decay and dropout are complementary and frequently combined in deep learning practice. Weight decay constrains the overall magnitude of the weight vectors, preventing any single parameter from dominating, while dropout adds stochastic noise to the activations, breaking co-adaptation between neurons. In chapter_multilayer-perceptrons/dropout.md, d2l-zh demonstrates networks that use both nn.Dropout layers and optimizer weight decay (wd parameter) simultaneously, showing improved generalization over using either technique alone.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →