# How d2l-zh Explains Machine Learning Fundamentals for Practitioners: A Code-First Approach

> d2l-zh explains machine learning fundamentals with runnable code examples and a seven-step workflow. Learn core concepts through practical application and a focused utility library.

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: deep-dive
- Published: 2026-03-01

---

**d2l-zh** teaches machine learning fundamentals through a structured seven-step workflow that interleaves conceptual theory with runnable code blocks, using a lightweight utility library in [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py) to keep examples focused on learning principles rather than boilerplate.

The Chinese edition of *Dive into Deep Learning*, hosted at **d2l-ai/d2l-zh**, introduces machine learning fundamentals to practitioners by following a pedagogical pattern that mirrors real-world ML workflows. Unlike traditional textbooks that separate theory from implementation, this open-source repository combines conceptual explanations in plain Chinese with minimal, self-contained code fragments that execute directly in Jupyter notebooks. The result is a bottom-up learning path that helps practitioners transition seamlessly from linear models to deep neural networks.

## The Seven-Step Machine Learning Workflow

According to the source code in [`chapter_introduction/index.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_introduction/index.md), the book structures its curriculum around a seven-step workflow that reflects how practitioners actually build systems. Each step maps to specific concepts and source files within the repository.

### From Problem Formulation to Generalization

The workflow follows this progression:

1. **Problem formulation** — Defines the input-output mapping and learning goals, covered in the introductory chapters.

2. **Data** — Explains collection, i.i.d. sampling assumptions, and preprocessing, emphasizing that "garbage in, garbage out."

3. **Model** — Introduces **model families** (parameterized functions) and distinguishes shallow models from deep neural networks.

4. **Objective function** — Quantifies performance using **squared loss** for regression or **cross-entropy** for classification.

5. **Optimization algorithm** — Details gradient-descent methods including SGD, momentum, and Adam, implemented in `d2l.torch`.

6. **Training loop** — Provides a concise Python loop tying together data loading, forward pass, loss computation, and parameter updates.

7. **Evaluation & Generalization** — Discusses training versus test error, over-fitting, and assessment protocols in [`chapter_multilayer-perceptrons/underfit-overfit.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_multilayer-perceptrons/underfit-overfit.md).

## The d2l.torch Utility Library

To maintain focus on machine learning fundamentals rather than boilerplate code, the repository provides [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py), a lightweight helper module that implements common utilities. This library supplies **synthetic data generation**, standardized **loss functions**, and the **stochastic gradient descent optimizer**, allowing practitioners to experiment with core algorithms in minimal lines of code. By abstracting repetitive operations into functions like `d2l.sgd()` and `d2l.load_array()`, the book keeps narrative attention on how components interact rather than implementation details.

## Hands-On Example: End-to-End Linear Regression

The following runnable example from the *线性回归* chapter demonstrates the complete pipeline using `d2l.torch` utilities. This pattern appears throughout the book, reinforcing theory with immediate executable verification.

```python
import torch
from d2l import torch as d2l

# 1️⃣ Create synthetic data (true w = [2, -3.4], b = 4.2)

true_w = torch.tensor([2.0, -3.4])
true_b = 4.2
features, labels = d2l.synthetic_data(true_w, true_b, 1000)

# 2️⃣ Build data iterator

batch_size = 10
data_iter = d2l.load_array((features, labels), batch_size)

# 3️⃣ Define model

def linreg(X, w, b):
    return d2l.matmul(X, w) + b

# 4️⃣ Define loss

def squared_loss(y_hat, y):
    return (y_hat - d2l.reshape(y, y_hat.shape)) ** 2 / 2

# 5️⃣ Initialize parameters

w = torch.randn((2, 1), requires_grad=True)
b = torch.zeros(1, requires_grad=True)

# 6️⃣ Training loop

num_epochs, lr = 3, 0.03
for epoch in range(num_epochs):
    for X, y in data_iter:
        l = squared_loss(linreg(X, w, b), y).mean()
        l.backward()
        d2l.sgd([w, b], lr, batch_size)
    # Compute loss on the whole dataset for monitoring

    train_l = squared_loss(linreg(features, w, b), labels).mean()
    print(f'epoch {epoch + 1}, loss {float(train_l):.6f}')
print('learned w', w.reshape(-1).detach().numpy())
print('learned b', b.item())

```

Executing this code in a Jupyter notebook prints rapidly decreasing loss values and recovers parameters close to the true values `[2, -3.4]` and `4.2`, illustrating how the **core ML pipeline** scales from simple linear regression to complex deep learning architectures.

## Key Source Files for Practitioners

The repository organizes its machine learning fundamentals curriculum across these critical files:

- **[`chapter_introduction/index.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_introduction/index.md)** — Provides the high-level overview of ML concepts, covering data, model, loss, optimizer, and the training loop.

- **[`chapter_linear-networks/linear-regression-scratch.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_linear-networks/linear-regression-scratch.md)** — Contains the step-by-step linear regression tutorial that the code example above mirrors.

- **[`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py)** — Houses the core utility library providing `synthetic_data`, `load_array`, `sgd`, `matmul`, and `reshape` functions used throughout the tutorials.

- **[`chapter_multilayer-perceptrons/underfit-overfit.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_multilayer-perceptrons/underfit-overfit.md)** — Discusses generalization, over-fitting, and evaluation metrics essential for production practitioners.

- **[`chapter_optimization/sgd.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/sgd.md)** and **[`chapter_optimization/adam.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_optimization/adam.md)** — Introduce stochastic gradient descent and advanced optimizers, showing the progression from basic to sophisticated training algorithms.

- **[`chapter_deep-learning-computation/parameters.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_deep-learning-computation/parameters.md)** — Explains parameter tensors and update mechanisms, connecting optimization theory to the implementation in [`torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/torch.py).

## Summary

- **d2l-zh** follows a seven-step workflow covering problem formulation, data, model, loss, optimizer, training loop, and generalization.

- The **bottom-up pedagogical style** presents concepts in plain Chinese first, then reinforces them with minimal, self-contained code fragments.

- The **[`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py)** utility library abstracts boilerplate, allowing practitioners to focus on machine learning fundamentals rather than framework specifics.

- Each chapter builds from **linear-model baselines to deeper architectures**, making the transition from classical ML to deep learning seamless.

- All examples are **runnable in Jupyter notebooks**, bridging the gap between theoretical understanding and practical implementation.

## Frequently Asked Questions

### What makes d2l-zh different from other machine learning textbooks?

Unlike traditional texts that separate mathematical theory from programming exercises, d2l-zh interleaves conceptual explanations with executable code blocks on every page. The repository provides a custom `d2l.torch` utility module that eliminates boilerplate, enabling practitioners to experiment with gradient descent and loss functions in minimal lines of code while the narrative explains the underlying mathematics.

### Do I need to know Chinese to use d2l-zh effectively?

While the text is written in Chinese, the code examples use standard Python and PyTorch APIs with English variable names, making the implementations accessible to readers worldwide. However, the conceptual explanations and mathematical intuitions are delivered in Chinese, so full comprehension requires Chinese reading proficiency.

### How does the d2l.torch utility library help practitioners learn faster?

The library in [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py) provides helper functions like `synthetic_data()`, `load_array()`, and `sgd()` that handle repetitive data generation and optimization boilerplate. This abstraction allows learners to focus on how data flows through models and how parameters update during training, rather than debugging tensor reshaping or batch iteration logic.

### Can I apply d2l-zh's teaching methodology to other deep learning frameworks?

Yes. While the code examples use PyTorch via `d2l.torch`, the seven-step workflow and conceptual framework are framework-agnostic. The book emphasizes fundamental principles—such as backpropagation, loss minimization, and over-fitting—that apply equally to TensorFlow, JAX, or other deep learning libraries.