How d2l-zh Explains Machine Learning Fundamentals for Practitioners: A Code-First Approach
d2l-zh teaches machine learning fundamentals through a structured seven-step workflow that interleaves conceptual theory with runnable code blocks, using a lightweight utility library in d2l/torch.py to keep examples focused on learning principles rather than boilerplate.
The Chinese edition of Dive into Deep Learning, hosted at d2l-ai/d2l-zh, introduces machine learning fundamentals to practitioners by following a pedagogical pattern that mirrors real-world ML workflows. Unlike traditional textbooks that separate theory from implementation, this open-source repository combines conceptual explanations in plain Chinese with minimal, self-contained code fragments that execute directly in Jupyter notebooks. The result is a bottom-up learning path that helps practitioners transition seamlessly from linear models to deep neural networks.
The Seven-Step Machine Learning Workflow
According to the source code in chapter_introduction/index.md, the book structures its curriculum around a seven-step workflow that reflects how practitioners actually build systems. Each step maps to specific concepts and source files within the repository.
From Problem Formulation to Generalization
The workflow follows this progression:
-
Problem formulation — Defines the input-output mapping and learning goals, covered in the introductory chapters.
-
Data — Explains collection, i.i.d. sampling assumptions, and preprocessing, emphasizing that "garbage in, garbage out."
-
Model — Introduces model families (parameterized functions) and distinguishes shallow models from deep neural networks.
-
Objective function — Quantifies performance using squared loss for regression or cross-entropy for classification.
-
Optimization algorithm — Details gradient-descent methods including SGD, momentum, and Adam, implemented in
d2l.torch. -
Training loop — Provides a concise Python loop tying together data loading, forward pass, loss computation, and parameter updates.
-
Evaluation & Generalization — Discusses training versus test error, over-fitting, and assessment protocols in
chapter_multilayer-perceptrons/underfit-overfit.md.
The d2l.torch Utility Library
To maintain focus on machine learning fundamentals rather than boilerplate code, the repository provides d2l/torch.py, a lightweight helper module that implements common utilities. This library supplies synthetic data generation, standardized loss functions, and the stochastic gradient descent optimizer, allowing practitioners to experiment with core algorithms in minimal lines of code. By abstracting repetitive operations into functions like d2l.sgd() and d2l.load_array(), the book keeps narrative attention on how components interact rather than implementation details.
Hands-On Example: End-to-End Linear Regression
The following runnable example from the 线性回归 chapter demonstrates the complete pipeline using d2l.torch utilities. This pattern appears throughout the book, reinforcing theory with immediate executable verification.
import torch
from d2l import torch as d2l
# 1️⃣ Create synthetic data (true w = [2, -3.4], b = 4.2)
true_w = torch.tensor([2.0, -3.4])
true_b = 4.2
features, labels = d2l.synthetic_data(true_w, true_b, 1000)
# 2️⃣ Build data iterator
batch_size = 10
data_iter = d2l.load_array((features, labels), batch_size)
# 3️⃣ Define model
def linreg(X, w, b):
return d2l.matmul(X, w) + b
# 4️⃣ Define loss
def squared_loss(y_hat, y):
return (y_hat - d2l.reshape(y, y_hat.shape)) ** 2 / 2
# 5️⃣ Initialize parameters
w = torch.randn((2, 1), requires_grad=True)
b = torch.zeros(1, requires_grad=True)
# 6️⃣ Training loop
num_epochs, lr = 3, 0.03
for epoch in range(num_epochs):
for X, y in data_iter:
l = squared_loss(linreg(X, w, b), y).mean()
l.backward()
d2l.sgd([w, b], lr, batch_size)
# Compute loss on the whole dataset for monitoring
train_l = squared_loss(linreg(features, w, b), labels).mean()
print(f'epoch {epoch + 1}, loss {float(train_l):.6f}')
print('learned w', w.reshape(-1).detach().numpy())
print('learned b', b.item())
Executing this code in a Jupyter notebook prints rapidly decreasing loss values and recovers parameters close to the true values [2, -3.4] and 4.2, illustrating how the core ML pipeline scales from simple linear regression to complex deep learning architectures.
Key Source Files for Practitioners
The repository organizes its machine learning fundamentals curriculum across these critical files:
-
chapter_introduction/index.md— Provides the high-level overview of ML concepts, covering data, model, loss, optimizer, and the training loop. -
chapter_linear-networks/linear-regression-scratch.md— Contains the step-by-step linear regression tutorial that the code example above mirrors. -
d2l/torch.py— Houses the core utility library providingsynthetic_data,load_array,sgd,matmul, andreshapefunctions used throughout the tutorials. -
chapter_multilayer-perceptrons/underfit-overfit.md— Discusses generalization, over-fitting, and evaluation metrics essential for production practitioners. -
chapter_optimization/sgd.mdandchapter_optimization/adam.md— Introduce stochastic gradient descent and advanced optimizers, showing the progression from basic to sophisticated training algorithms. -
chapter_deep-learning-computation/parameters.md— Explains parameter tensors and update mechanisms, connecting optimization theory to the implementation intorch.py.
Summary
-
d2l-zh follows a seven-step workflow covering problem formulation, data, model, loss, optimizer, training loop, and generalization.
-
The bottom-up pedagogical style presents concepts in plain Chinese first, then reinforces them with minimal, self-contained code fragments.
-
The
d2l/torch.pyutility library abstracts boilerplate, allowing practitioners to focus on machine learning fundamentals rather than framework specifics. -
Each chapter builds from linear-model baselines to deeper architectures, making the transition from classical ML to deep learning seamless.
-
All examples are runnable in Jupyter notebooks, bridging the gap between theoretical understanding and practical implementation.
Frequently Asked Questions
What makes d2l-zh different from other machine learning textbooks?
Unlike traditional texts that separate mathematical theory from programming exercises, d2l-zh interleaves conceptual explanations with executable code blocks on every page. The repository provides a custom d2l.torch utility module that eliminates boilerplate, enabling practitioners to experiment with gradient descent and loss functions in minimal lines of code while the narrative explains the underlying mathematics.
Do I need to know Chinese to use d2l-zh effectively?
While the text is written in Chinese, the code examples use standard Python and PyTorch APIs with English variable names, making the implementations accessible to readers worldwide. However, the conceptual explanations and mathematical intuitions are delivered in Chinese, so full comprehension requires Chinese reading proficiency.
How does the d2l.torch utility library help practitioners learn faster?
The library in d2l/torch.py provides helper functions like synthetic_data(), load_array(), and sgd() that handle repetitive data generation and optimization boilerplate. This abstraction allows learners to focus on how data flows through models and how parameters update during training, rather than debugging tensor reshaping or batch iteration logic.
Can I apply d2l-zh's teaching methodology to other deep learning frameworks?
Yes. While the code examples use PyTorch via d2l.torch, the seven-step workflow and conceptual framework are framework-agnostic. The book emphasizes fundamental principles—such as backpropagation, loss minimization, and over-fitting—that apply equally to TensorFlow, JAX, or other deep learning libraries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →