Core Concepts of Deep Learning in d2l-zh: The Four-Component Framework Explained

The d2l-zh repository introduces deep learning through four fundamental components—data, model, objective function, and optimization algorithm—combined with two key neural network principles: layered architectures and back-propagation.

The d2l-zh project (the Chinese edition of Dive into Deep Learning) establishes a unified framework for understanding the core concepts of deep learning. According to the source code in chapter_introduction/index.md, every machine learning system relies on four interlocking elements, while deep neural networks specifically depend on two architectural principles that enable efficient learning from complex data.

The Four Fundamental Components of Machine Learning

The foundation of any machine learning system in d2l-zh rests on four essential components detailed in the introduction chapter. These elements form a complete pipeline that remains consistent from simple linear regression to advanced transformer architectures.

Data

Data represents the raw observations that the model learns from. In d2l-zh, datasets serve as the empirical foundation for training, validation, and testing. The repository emphasizes that high-quality data preprocessing directly impacts model performance across all subsequent deep learning applications.

Model

The model is the parametric mapping that transforms input data into predictions. Whether implementing a simple linear function in chapter_linear-networks/linear-regression.md or complex attention mechanisms, the model defines the mathematical relationship between inputs and outputs through learnable parameters.

Objective Function

Also called the loss function, this scalar quantifies how well the model fits the data. The objective function provides a measurable target for optimization, converting prediction errors into a single numerical value that guides parameter updates.

Optimization Algorithm

The optimization algorithm adjusts model parameters to minimize the loss. Methods such as SGD (Stochastic Gradient Descent) and Adam iterate through the training data, computing gradients and updating weights to progressively improve model accuracy.

The Two Core Principles of Neural Networks

Beyond the four components, d2l-zh emphasizes two architectural principles in the "神经网络的核心" (Core of Neural Networks) section that specifically enable deep learning.

Layered Structure

Deep neural networks utilize alternating linear and nonlinear processing units called layers. This layered architecture allows models to learn hierarchical representations, where early layers detect simple features and deeper layers combine these into complex patterns. The chapter_multilayer-perceptrons/mlp.md file demonstrates how stacking Linear layers with activation functions like ReLU creates this depth.

Back-Propagation

Back-propagation applies the chain rule to compute gradients for all parameters in a single backward pass. This efficient algorithm makes training deep networks computationally feasible by calculating how each parameter contributes to the final loss, enabling simultaneous updates across millions of weights.

From Theory to Practice: Code Examples

The following runnable examples demonstrate these core concepts using PyTorch-based implementations from the repository.

Linear Regression: The Simplest Model

Located in chapter_linear-networks/linear-regression.md, this example illustrates the complete data → model → loss → optimization loop:

import torch
from d2l import torch as d2l

# synthetic data

x = torch.arange(0, 10, dtype=torch.float32).unsqueeze(1)
y = 2 * x + 1 + torch.randn_like(x) * 0.5

# model: ŷ = w·x + b

net = torch.nn.Linear(1, 1)

# loss & optimizer

loss = torch.nn.MSELoss()
trainer = torch.optim.SGD(net.parameters(), lr=0.01)

def train():
    for epoch in range(10):
        trainer.zero_grad()
        l = loss(net(x), y)
        l.backward()
        trainer.step()
        print(f'epoch {epoch}: loss {l.item():.4f}')

train()

Multilayer Perceptrons: Adding Depth and Non-Linearity

The MLP example from chapter_multilayer-perceptrons/mlp.md demonstrates how stacking layers with activation functions enables complex pattern recognition:

net = torch.nn.Sequential(
    torch.nn.Linear(784, 256), torch.nn.ReLU(),
    torch.nn.Linear(256, 10)
)

loss = torch.nn.CrossEntropyLoss()
trainer = torch.optim.SGD(net.parameters(), lr=0.1)

# a single training step on a mini‑batch (X, y)

def train_step(X, y):
    trainer.zero_grad()
    l = loss(net(X), y)
    l.backward()
    trainer.step()

Back-Propagation: Automatic Gradient Computation

This example shows how PyTorch's backward() method implements the chain rule described in the introduction:


# forward pass

y_hat = net(X)

# compute loss

l = loss(y_hat, y)

# backward pass (chain rule)

l.backward()

# inspect a gradient

print(net[0].weight.grad)   # gradient of first Linear layer

Optimization: Stochastic Gradient Descent with Momentum

From chapter_optimization/sgd.md, this example shows how momentum accelerates convergence:

trainer = torch.optim.SGD(net.parameters(),
                         lr=0.01,
                         momentum=0.9)   # momentum term

Regularization: Dropout

The dropout implementation in chapter_multilayer-perceptrons/dropout.md prevents overfitting by randomly zeroing neurons during training:

net = torch.nn.Sequential(
    torch.nn.Linear(784, 256), torch.nn.ReLU(),
    torch.nn.Dropout(0.5),               # 50% dropout

    torch.nn.Linear(256, 10)
)

Convolutional Neural Networks: 2-D Convolution

The CNN core concept from chapter_convolutional-neural-networks/conv-layer.md exploits spatial structure using Conv2d layers:

conv = torch.nn.Conv2d(in_channels=1,
                       out_channels=16,
                       kernel_size=3,
                       stride=1,
                       padding=1)
X = torch.randn(8, 1, 28, 28)  # batch of 8 grayscale images

Y = conv(X)                     # shape: (8, 16, 28, 28)

Recurrent Neural Networks: Sequence Modeling

From chapter_recurrent-neural-networks/rnn.md, RNNs handle sequential data through temporal recurrence:

rnn = torch.nn.RNN(input_size=10,
                   hidden_size=20,
                   num_layers=1,
                   batch_first=True)

X = torch.randn(32, 100, 10)   # (batch, seq_len, input_dim)

output, hn = rnn(X)           # output: (32,100,20)

Attention and Transformers: Self-Attention Mechanisms

The transformer implementation in chapter_attention-mechanisms/transformer.md introduces self-attention as an alternative to recurrence:

from d2l import torch as d2l
attention = d2l.MultiHeadAttention(num_heads=8,
                                   query_dim=64,
                                   key_dim=64,
                                   value_dim=64,
                                   dropout=0.1)

X = torch.randn(2, 10, 64)   # (batch, seq_len, d_model)

Y = attention(X, X, X)        # self-attention

Key Source Files in d2l-zh

Understanding the repository structure helps navigate these core concepts of deep learning:

Summary

  • The four fundamental components—data, model, objective function, and optimization algorithm—form the complete pipeline of any machine learning system in d2l-zh.
  • Layered architectures and back-propagation serve as the two core principles enabling deep neural networks to learn complex hierarchical representations.
  • The repository progresses logically from linear regression to transformers, maintaining the same four-component framework throughout all implementations.
  • Each concept includes runnable PyTorch code examples that demonstrate theoretical principles in practice, using the unified d2l package for consistency across MXNet, PyTorch, and TensorFlow backends.

Frequently Asked Questions

What are the four fundamental components of machine learning introduced in d2l-zh?

According to chapter_introduction/index.md, the four components are data (raw observations), model (parametric mapping), objective function (loss scalar), and optimization algorithm (parameter update method). These elements create a closed loop where data feeds into the model, the loss measures prediction error, and the optimizer adjusts parameters to minimize that error.

How does d2l-zh explain back-propagation?

The repository describes back-propagation as the application of the chain rule to compute gradients for all parameters in one backward pass. As implemented in the code examples, calling loss.backward() automatically calculates these gradients, which the optimization algorithm then uses to update weights via trainer.step(). This efficient computation makes training deep networks with millions of parameters feasible.

What optimization algorithms does d2l-zh cover?

The book covers Stochastic Gradient Descent (SGD) with various enhancements including momentum, Adam, learning-rate schedules, and weight decay. These are detailed in chapter_optimization/sgd.md and related files, demonstrating how different strategies affect convergence speed and final model performance.

How does d2l-zh structure the progression from basic to advanced deep learning concepts?

The repository follows a pedagogical path starting with linear models (least squares and softmax regression), then adding depth with multilayer perceptrons, incorporating regularization techniques (dropout, batch normalization), and finally specializing into CNNs for images, RNNs for sequences, and Transformers for attention-based modeling. Throughout this progression, the same four-component framework and back-propagation principle remain consistent.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →