How d2l-zh Introduces and Explains Multilayer Perceptrons (MLPs): Theory, Math, and Code

The d2l-zh textbook introduces multilayer perceptrons through a two-tiered pedagogical approach: first establishing theoretical foundations covering architecture, activation functions, and universal approximation properties in mlp.md, then providing hands-on from-scratch and high-level API implementations across MXNet, PyTorch, TensorFlow, and PaddlePaddle in mlp-scratch.md and mlp-concise.md.

The Chinese edition of Dive into Deep Learning (d2l-zh) delivers a comprehensive introduction to multilayer perceptrons (MLPs) that bridges rigorous mathematical formulation with executable code. This educational resource demonstrates how adding fully-connected hidden layers and non-linear activations transforms simple linear models into universal function approximators. Through a structured progression from conceptual explanations to framework-specific implementations, d2l-zh provides canonical explanations of MLPs backed by working examples in four major deep learning frameworks.

Conceptual Foundation: The MLP Architecture

The theoretical introduction to multilayer perceptrons resides in chapter_multilayer-perceptrons/mlp.md, where the text establishes the architectural principles distinguishing MLPs from linear models.

From Linear Models to Hidden Layers

In mlp.md (lines 58-66), d2l-zh defines an MLP by contrasting it with softmax regression, explaining that "adding one or more fully-connected hidden layers between the input and output" constitutes the fundamental architectural change. The chapter visualizes a single-hidden-layer MLP with 4 inputs, 5 hidden units, and 3 outputs, explicitly noting that this architecture counts as 2 layers (input-to-hidden and hidden-to-output) according to standard conventions cited in lines 71-73.

Mathematical Formulation and Forward Pass

The forward computation for a one-hidden-layer MLP appears in Equations (1)-(2) of mlp.md (lines 94-105). The text introduces:

  • Weight matrix W¹ and bias b¹ for the hidden layer
  • Weight matrix W² and bias b² for the output layer
  • Hidden representation H computed as an intermediate affine transformation

This mathematical notation establishes the blueprint for the subsequent code implementations, showing how input data flows through successive linear transformations separated by activation functions.

The Critical Role of Non-Linear Activation

A crucial theoretical point emphasized in mlp.md (lines 126-133) is the necessity of non-linear activation functions (denoted as σ). The text explicitly warns that without applying a non-linear activation after the first affine transform, the network collapses back to a linear model regardless of depth. This concept motivates the ReLU implementations found in the practical sections.

Universal Approximation Properties

Extending beyond single hidden layers, mlp.md (lines 150-162) discusses how deeper stacks of hidden layers provide universal approximation capabilities, forming the theoretical justification for using MLPs to model complex, non-linear relationships in data.

Hands-On Implementation: From-Scratch vs. High-Level APIs

D2l-zh reinforces theoretical concepts with dual implementation tracks: manual construction using fundamental operations and streamlined versions using framework abstractions.

Building an MLP from Scratch

The chapter chapter_multilayer-perceptrons/mlp-scratch.md provides a detailed manual implementation that exposes every component of the MLP computation graph.

Parameter Initialization (lines 63-70):

import d2l.mxnet as d2l
from mxnet import np, npx
npx.set_np()

num_inputs, num_hiddens, num_outputs = 784, 256, 10

W1 = np.random.normal(scale=0.01, size=(num_inputs, num_hiddens))
b1 = np.zeros(num_hiddens)
W2 = np.random.normal(scale=0.01, size=(num_hiddens, num_outputs))
b2 = np.zeros(num_outputs)
params = [W1, b1, W2, b2]
for p in params:
    p.attach_grad()

Activation Function (lines 27-29):

def relu(X): 
    return np.maximum(X, 0)

Forward Pass Implementation (lines 58-62):

def net(X):
    X = d2l.reshape(X, (-1, num_inputs))
    H = relu(np.dot(X, W1) + b1)
    return np.dot(H, W2) + b2

Training Loop Reuse: Notably, mlp-scratch.md (lines 12-23) demonstrates that the training procedure remains identical to softmax regression, utilizing the d2l.train_ch3 helper function provided by the d2l utility modules (d2l/torch.py, d2l/mxnet.py, etc.). This architectural decision emphasizes that MLPs and linear models share the same optimization interface despite their representational differences.

Concise Implementation with Framework APIs

For production-oriented workflows, chapter_multilayer-perceptrons/mlp-concise.md demonstrates compact implementations using high-level APIs:

PyTorch Implementation:

import d2l.torch as d2l
from torch import nn

net = nn.Sequential(
    nn.Flatten(),
    nn.Linear(784, 256),      # hidden layer

    nn.ReLU(),
    nn.Linear(256, 10)        # output layer

)

# Training uses the same d2l.train_ch3 helper

# d2l.train_ch3(net, train_iter, test_iter,

#               d2l.loss_cross_entropy, num_epochs=10,

#               optimizer=lambda lr: torch.optim.SGD(net.parameters(), lr=0.1))

TensorFlow Keras Implementation:

import tensorflow as tf
from d2l import tensorflow as d2l

net = tf.keras.Sequential([
    tf.keras.layers.Flatten(),
    tf.keras.layers.Dense(256, activation='relu'),
    tf.keras.layers.Dense(10)  # logits

])

loss = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)
optimizer = tf.keras.optimizers.SGD(learning_rate=0.1)

d2l.train_ch3(net, train_iter, test_iter,
              loss, num_epochs=10,
              optimizer=optimizer)

Both approaches implement the same three-layer architecture (input → hidden → output) introduced in the theoretical chapter, demonstrating the equivalence between explicit matrix operations and framework abstractions.

Framework Agnostic Training Infrastructure

A distinctive feature of d2l-zh's MLP introduction is the use of the d2l utility library (available in d2l/torch.py, d2l/mxnet.py, d2l/tensorflow.py, and d2l/paddle.py) to abstract training procedures. This design allows readers to focus on architectural differences between models while using identical training loops, such as d2l.train_ch3, whether implementing softmax regression or complex MLPs.

Summary

  • D2l-zh introduces multilayer perceptrons in chapter_multilayer-perceptrons/mlp.md by contrasting them with linear models and defining them as networks with one or more fully-connected hidden layers between input and output.
  • The mathematical formulation in mlp.md (lines 94-105) introduces weight matrices W¹, b¹, W², b² and emphasizes that non-linear activations (σ) are essential to prevent the network from collapsing into a linear model.
  • From-scratch implementation in mlp-scratch.md manually constructs the ReLU activation, parameter tensors, and forward pass using basic tensor operations across four frameworks.
  • High-level API implementations in mlp-concise.md demonstrate concise MLP construction using nn.Sequential (PyTorch) and tf.keras.Sequential (TensorFlow).
  • The training loop remains consistent across model types through the d2l.train_ch3 utility function, highlighting that MLPs and linear regression share identical optimization interfaces despite different representational capacities.

Frequently Asked Questions

What is the difference between the MLP explanations in mlp.md and mlp-scratch.md?

mlp.md provides the theoretical and mathematical foundation for multilayer perceptrons, explaining concepts like hidden layers, weight matrices, non-linear activations, and universal approximation properties with supporting equations and diagrams. In contrast, mlp-scratch.md offers a practical, step-by-step implementation where you manually initialize parameters W₁, b₁, W₂, b₂, implement the ReLU function, and construct the forward pass using fundamental tensor operations in MXNet, PyTorch, TensorFlow, or PaddlePaddle.

Why does d2l-zh emphasize non-linear activation functions for MLPs?

According to mlp.md (lines 126-133), d2l-zh emphasizes non-linear activations because without them, composing multiple affine transformations collapses mathematically into a single linear transformation. The text demonstrates that applying a non-linear function σ (such as ReLU) after the hidden layer's affine transform enables the network to learn non-linear decision boundaries, which is the primary advantage of MLPs over simple linear models like softmax regression.

Can I use the same training loop for MLPs as for softmax regression in d2l-zh?

Yes. As explicitly shown in mlp-scratch.md (lines 12-23), d2l-zh uses the identical d2l.train_ch3 helper function for training MLPs that was introduced in the softmax regression chapter. This design emphasizes that while the model architecture and forward pass differ between linear models and MLPs, the optimization procedure—computing gradients and updating parameters—remains consistent, allowing readers to focus on architectural changes rather than boilerplate training code.

Which deep learning frameworks does d2l-zh support for MLP implementation?

D2l-zh provides complete MLP implementations for four major frameworks: MXNet, PyTorch, TensorFlow, and PaddlePaddle. Each framework has corresponding utility modules (d2l/mxnet.py, d2l/torch.py, d2l/tensorflow.py, d2l/paddle.py) that provide consistent APIs for data loading, training loops, and loss functions, enabling readers to learn MLP concepts while working in their preferred environment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →