Mathematical Basis for Linear Networks in d2l-zh: Affine Transformations and Maximum Likelihood
Linear networks in d2l-zh are single-layer models built on affine transformations (ŷ = w^T x + b) with loss functions derived from maximum likelihood estimation—mean squared error for regression under Gaussian noise assumptions and cross-entropy for classification via multinomial logistic regression.
The d2l-zh textbook introduces linear networks as the foundational building blocks of deep learning, presenting them as the simplest neural architectures that compute affine transformations of input features. Understanding the mathematical basis for linear networks in d2l-zh requires examining how the repository derives objective functions from probabilistic principles in chapter_linear-networks/linear-regression.md and chapter_linear-networks/softmax-regression.md, then implements them through vectorized matrix operations.
Core Mathematical Formulation
Linear networks in d2l-zh reduce to computing an affine map followed by a task-specific transformation. The repository distinguishes between two primary instantiations: regression with continuous outputs and classification via probability distributions.
Linear Regression as Affine Transformation
In chapter_linear-networks/linear-regression.md, the model predicts a real-valued target by computing:
ŷ = w^T x + b
As specified in Eq. (71)-(72), where w ∈ ℝ^d represents weights and b ∈ ℝ is the bias term. For a dataset with n examples, the vectorized formulation becomes ŷ = Xw + b, where X ∈ ℝ^(n×d) denotes the design matrix.
The training objective minimizes the mean squared error (MSE):
L(w, b) = (1/n) Σᵢ (1/2)(ŷ⁽ⁱ⁾ - y⁽ⁱ⁾)²
This corresponds to Eq. (29) in the source text. The factor of 1/2 simplifies the derivative calculation during gradient descent without affecting the optimization optimum.
Softmax Regression and the Exponential Family
For classification tasks covered in chapter_linear-networks/softmax-regression.md, the network still computes an affine transformation o = Wx + b (Eq. 76), where W ∈ ℝ^(q×d) and q represents the number of classes. However, the raw scores o pass through the softmax function to produce a probability distribution:
ŷⱼ = exp(oⱼ) / Σₖ exp(oₖ)
This normalization, shown in Eq. (24), ensures outputs sum to 1 and form a valid categorical distribution. The model essentially implements multinomial logistic regression, treating the affine layer as logits for the exponential family distribution.
Probabilistic Foundations and Loss Functions
The d2l-zh repository emphasizes that loss functions emerge from maximum likelihood estimation rather than arbitrary choices, providing principled justification for the training objectives.
Gaussian Noise and Mean Squared Error
According to the "正态分布与平方损失" (Normal Distribution and Squared Loss) section in chapter_linear-networks/linear-regression.md, assuming additive Gaussian noise ε ~ N(0, σ²) leads directly to the squared error objective. The likelihood function:
P(y|x) = (1/√(2πσ²)) exp(-(1/(2σ²))(y - w^T x - b)²)
Maximizing this likelihood with respect to w and b is equivalent to minimizing the MSE loss. Furthermore, the analytical solution w* = (X^T X)^(-1) X^T y (Eq. 46) provides a closed-form optimum against which gradient-based methods can be compared, particularly revealing convergence issues when X^T X is ill-conditioned.
Multinomial Logistic Regression and Cross-Entropy
For classification, chapter_linear-networks/softmax-regression.md derives the cross-entropy loss (negative log-likelihood) from the multinomial distribution:
l(y, ŷ) = -Σⱼ yⱼ log ŷⱼ
As shown in Eq. (33), this loss measures the dissimilarity between the true one-hot distribution y and the predicted probabilities ŷ. The "对数似然" (Log Likelihood) and "softmax及其导数" sections demonstrate that minimizing this loss maximizes the probability of observing the training labels under the model's predicted distribution.
Vectorized Minibatch Implementation
Both chapters emphasize that linear networks compute forward passes through matrix multiplication for efficiency. For a minibatch X ∈ ℝ^(n×d), the operation becomes:
O = XW + b
Where O ∈ ℝ^(n×q) contains the outputs for all n examples simultaneously (Eq. 51). In the regression case, O represents the predictions directly; for classification, O feeds into the softmax function. This formulation highlights that linear networks constitute a single fully-connected layer followed by an optional non-linear activation, forming the computational backbone of all deeper architectures in the d2l-zh curriculum.
Code Examples
The following runnable implementations demonstrate how mathematical concepts translate to d2l-zh code across different backends.
Linear Regression with MXNet Backend
This example from chapter_linear-networks/linear-regression-concise.md implements the affine transformation and MSE loss using high-level APIs:
import d2l.mxnet as d2l
from mxnet import np, autograd, gluon
# Generate synthetic data with true parameters
true_w, true_b = np.array([2.0, -3.4]), 4.2
features = np.random.normal(0, 1, (1000, 2))
labels = np.dot(features, true_w) + true_b + np.random.normal(0, 0.01, (1000,))
# Define single dense layer computing Xw + b
net = gluon.nn.Dense(1)
net.initialize()
loss = gluon.loss.L2Loss() # Implements MSE from Eq. (29)
# Training loop with SGD
trainer = gluon.Trainer(net.collect_params(), 'sgd', {'learning_rate': 0.03})
for epoch in range(5):
with autograd.record():
output = net(features).reshape((-1,))
l = loss(output, labels)
l.backward()
trainer.step(batch_size=features.shape[0])
print(f'epoch {epoch+1}, loss {l.mean().asscalar():.6f}')
The gluon.nn.Dense layer automatically computes the affine transformation Xw + b, while L2Loss implements the squared error term derived from Gaussian maximum likelihood.
Softmax Regression with PyTorch Backend
This implementation from chapter_linear-networks/softmax-regression-concise.md shows the classification case:
import d2l.torch as d2l
import torch
from torch import nn, optim
# Load Fashion-MNIST dataset
batch_size = 256
train_iter, test_iter = d2l.load_data_fashion_mnist(batch_size)
# Single linear layer: affine map o = XW + b
net = nn.Linear(784, 10) # 28×28 images flattened to 784 features
loss = nn.CrossEntropyLoss() # Combines softmax + negative log-likelihood
optimizer = optim.SGD(net.parameters(), lr=0.1)
# Training loop
num_epochs = 5
for epoch in range(num_epochs):
d2l.train_ch3(net, train_iter, loss, optimizer, device=d2l.try_gpu())
acc = d2l.evaluate_accuracy(net, test_iter)
print(f'epoch {epoch+1}, test accuracy {acc:.3f}')
Here, nn.Linear implements the affine transformation XW + b (Eq. 76), while CrossEntropyLoss internally applies the softmax function and computes the cross-entropy loss derived from the multinomial distribution.
Summary
- Linear networks in d2l-zh compute affine transformations (ŷ = w^Tx + b or o = WX + b) as their fundamental operation.
- Loss functions derive from maximum likelihood estimation: MSE corresponds to Gaussian noise assumptions in regression, while cross-entropy corresponds to multinomial logistic regression in classification.
- The repository provides both analytical solutions (closed-form linear regression) and gradient-based implementations to illustrate optimization dynamics.
- Vectorized formulations (Eq. 51) enable efficient minibatch processing using standard matrix multiplication operations.
- Implementation files
chapter_linear-networks/linear-regression.mdandchapter_linear-networks/softmax-regression.mdbridge theoretical derivations with practical code.
Frequently Asked Questions
Why does d2l-zh use (1/2) times the MSE in the loss function?
The factor of 1/2 in L(w, b) = (1/n)Σᵢ(1/2)(ŷ⁽ⁱ⁾ - y⁽ⁱ⁾)² serves a computational convenience. When computing gradients with respect to the weights, the derivative of the squared term eliminates the 1/2 coefficient, yielding a cleaner ∂L/∂w expression. This convention, adopted in chapter_linear-networks/linear-regression.md Eq. (29), simplifies manual gradient calculations without affecting the location of the loss minimum.
What distinguishes softmax regression from standard linear layers?
Softmax regression adds a softmax activation to the affine output, converting raw logits into probability distributions that sum to 1. While nn.Linear or gluon.nn.Dense compute the affine map o = Wx + b, the softmax function ŷⱼ = exp(oⱼ)/Σₖexp(oₖ) and subsequent cross-entropy loss transform these scores into a multinomial logistic regression model, as derived in chapter_linear-networks/softmax-regression.md sections "对数似然" and "softmax及其导数".
When does gradient descent converge to the analytical solution in linear regression?
Gradient descent converges to the same optimum as the analytical solution w* = (X^TX)^(-1)X^Ty (Eq. 46) when the learning rate is sufficiently small and the optimization runs long enough. However, when the matrix X^TX is ill-conditioned (near-singular), gradient descent may converge slowly or numerically destabilize, whereas the analytical solution faces direct inversion challenges. The d2l-zh examples in chapter_linear-networks/linear-regression-scratch.md demonstrate both approaches for comparison.
Why does CrossEntropyLoss not require an explicit softmax layer?
In PyTorch and similar frameworks used by d2l-zh, CrossEntropyLoss combines LogSoftmax and NLLLoss (negative log-likelihood) in a single, numerically stable operation. Mathematically, this computes log(softmax(o)) before applying the negative likelihood, avoiding separate exponentiation and division steps that could cause numerical overflow. This implementation directly corresponds to the cross-entropy loss l(y, ŷ) = -Σⱼyⱼlogŷⱼ from Eq. (33) in chapter_linear-networks/softmax-regression.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →