# Key Architectural Components of Convolutional Neural Networks (CNNs) According to d2l-zh

> Discover the 8 key architectural components of CNNs explained by d2l-zh including convolutional layers pooling layers batch normalization and more Learn how CNNs achieve spatial invariance and efficiency

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: architecture
- Published: 2026-03-01

---

**According to the d2l-zh educational repository, CNNs achieve spatial invariance and parameter efficiency through a stack of eight specialized components: convolutional layers, padding/stride controls, pooling layers, multi-channel feature maps, batch normalization, activation functions, fully-connected layers, and hierarchical architectural patterns.**

The **d2l-zh** repository (the Chinese edition of *Dive into Deep Learning*) provides an open-source reference implementation and theoretical explanation of deep learning architectures. The primary keyword architectural components of Convolutional Neural Networks are systematically defined across eight markdown chapters that explain how weight sharing, local receptive fields, and hierarchical feature extraction enable efficient visual processing.

## The Eight Core Architectural Components

### Convolutional Layers with Cross-Correlation

In [`chapter_convolutional-neural-networks/conv-layer.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_convolutional-neural-networks/conv-layer.md), the foundational operation is defined as **2-D cross-correlation** between a learnable kernel and the input tensor. The kernel slides across all spatial locations while sharing parameters, producing feature maps that detect local patterns like edges and textures.

The `corr2d` function in the source demonstrates the mathematical operation without framework abstractions:

```python
def corr2d(X, K):          # X: (H, W), K: (kH, kW)

    h, w = K.shape
    Y = d2l.zeros((X.shape[0]-h+1, X.shape[1]-w+1))
    for i in range(Y.shape[0]):
        for j in range(Y.shape[1]):
            Y[i, j] = d2l.reduce_sum(X[i:i+h, j:j+w] * K)
    return Y

```

This weight-sharing mechanism drastically reduces parameter counts compared to fully-connected layers while maintaining spatial locality.

### Padding and Stride Controls

As detailed in [`chapter_convolutional-neural-networks/padding-and-strides.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_convolutional-neural-networks/padding-and-strides.md), **padding** adds zeros around the input border to preserve spatial dimensions during convolution, while **stride** controls the step size of the kernel slide, enabling controlled downsampling. These hyperparameters determine the output tensor shape and receptive field growth through the network.

### Pooling Layers for Translation Invariance

The [`pooling.md`](https://github.com/d2l-ai/d2l-zh/blob/main/pooling.md) chapter introduces **max-pooling** and **average-pooling** layers that aggregate statistics over small windows (typically 2×2). These layers provide translation invariance by making the network less sensitive to exact pixel positions while reducing spatial resolution and computational load.

The reference implementation `pool2d` shows the aggregation logic:

```python
def pool2d(X, pool_size, mode='max'):
    p_h, p_w = pool_size
    Y = d2l.zeros((X.shape[0]-p_h+1, X.shape[1]-p_w+1))
    for i in range(Y.shape[0]):
        for j in range(Y.shape[1]):
            window = X[i:i+p_h, j:j+p_w]
            Y[i, j] = window.max() if mode == 'max' else window.mean()
    return Y

```

### Multi-Channel Feature Maps

According to [`channels.md`](https://github.com/d2l-ai/d2l-zh/blob/main/channels.md), modern CNNs operate on **multiple input and output channels**, where each channel represents a distinct feature map. A convolutional layer transforms input channels into output channels by learning separate kernels for each input-output pair, then summing the results. This multi-channel structure enables the network to learn complex, compositional features across depth.

### Batch Normalization for Training Stability

The [`chapter_convolutional-modern/batch-norm.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_convolutional-modern/batch-norm.md) file adds **batch normalization** as a critical architectural component for deep CNNs. This layer normalizes activations across the mini-batch dimension, stabilizing gradient flow during training and acting as a regularizer that reduces the need for careful weight initialization.

### Activation Functions for Non-Linearity

While not a separate layer type in the file structure, **activation functions** (typically ReLU) are applied immediately after convolution operations to introduce non-linearity. As implemented in the d2l-zh examples, these point-wise transformations allow the network to model complex, non-linear mappings from pixel inputs to semantic outputs.

### Fully-Connected Classification Heads

After stacking multiple convolution-pooling blocks, the high-level spatial representations are flattened and fed into **fully-connected (dense) layers** for final classification or regression. The [`lenet.md`](https://github.com/d2l-ai/d2l-zh/blob/main/lenet.md) chapter demonstrates this transition from spatial feature extraction to vector-based prediction in the LeNet architecture.

### Hierarchical Architectural Patterns

The components are organized into **repeated blocks** that form complete network architectures. [`lenet.md`](https://github.com/d2l-ai/d2l-zh/blob/main/lenet.md) presents the classic LeNet-5 pattern (Conv → ReLU → Pool → Conv → ReLU → Pool → Dense), while [`resnet.md`](https://github.com/d2l-ai/d2l-zh/blob/main/resnet.md) extends these components with skip connections for very deep networks.

## Implementing the Canonical CNN Stack

The d2l-zh repository demonstrates these components through framework-agnostic abstractions. Below is the canonical **Conv-BN-ReLU-Pool** block using MXNet Gluon syntax (identical logic applies to PyTorch and TensorFlow variants):

```python
from d2l import mxnet as d2l
from mxnet import np, npx
from mxnet.gluon import nn
npx.set_np()

class ConvBlock(nn.Block):
    def __init__(self, kernel_size, num_channels, pool_size=(2,2), **kwargs):
        super().__init__(**kwargs)
        self.conv = nn.Conv2D(num_channels, kernel_size=kernel_size,
                             padding=kernel_size//2)  # same padding

        self.bn   = nn.BatchNorm()
        self.pool = nn.MaxPool2D(pool_size)

    def forward(self, X):
        return self.pool(d2l.relu(self.bn(self.conv(X))))

```

A complete LeNet-style architecture combining all eight components appears as:

```python
class LeNet(nn.Block):
    def __init__(self, **kwargs):
        super().__init__(**kwargs)
        self.conv1 = nn.Conv2D(6, kernel_size=5, padding=2)
        self.bn1   = nn.BatchNorm()
        self.pool1 = nn.MaxPool2D(2)
        self.conv2 = nn.Conv2D(16, kernel_size=5)
        self.bn2   = nn.BatchNorm()
        self.pool2 = nn.MaxPool2D(2)
        self.fc1   = nn.Dense(120)
        self.fc2   = nn.Dense(84)
        self.fc3   = nn.Dense(10)

    def forward(self, X):
        X = self.pool1(d2l.relu(self.bn1(self.conv1(X))))
        X = self.pool2(d2l.relu(self.bn2(self.conv2(X))))
        X = X.reshape((0, -1))          # flatten

        X = d2l.relu(self.fc1(X))
        X = d2l.relu(self.fc2(X))
        return self.fc3(X)

```

## Summary

- **Convolutional layers** with shared kernels provide parameter efficiency and local receptive fields as defined in [`conv-layer.md`](https://github.com/d2l-ai/d2l-zh/blob/main/conv-layer.md).
- **Padding and stride** control spatial dimensions and downsampling rates according to [`padding-and-strides.md`](https://github.com/d2l-ai/d2l-zh/blob/main/padding-and-strides.md).
- **Pooling layers** deliver translation invariance through max or average aggregation detailed in [`pooling.md`](https://github.com/d2l-ai/d2l-zh/blob/main/pooling.md).
- **Multi-channel convolutions** enable hierarchical feature extraction across depth as explained in [`channels.md`](https://github.com/d2l-ai/d2l-zh/blob/main/channels.md).
- **Batch normalization** stabilizes deep network training per [`batch-norm.md`](https://github.com/d2l-ai/d2l-zh/blob/main/batch-norm.md).
- **Fully-connected layers** convert spatial features into final predictions, demonstrated in [`lenet.md`](https://github.com/d2l-ai/d2l-zh/blob/main/lenet.md).
- **Architectural patterns** like LeNet and ResNet show how to stack these components into production-ready CNNs.

## Frequently Asked Questions

### What is the role of weight sharing in CNN architectures?

Weight sharing, implemented through the sliding kernel mechanism in convolutional layers, drastically reduces the number of learnable parameters compared to fully-connected networks. According to [`why-conv.md`](https://github.com/d2l-ai/d2l-zh/blob/main/why-conv.md), this sharing also provides **spatial invariance**, allowing the network to detect patterns regardless of their position in the input image.

### How does batch normalization improve CNN training?

As described in [`batch-norm.md`](https://github.com/d2l-ai/d2l-zh/blob/main/batch-norm.md), batch normalization normalizes layer outputs across the mini-batch dimension, which stabilizes the distribution of activations during training. This reduces internal covariate shift, allows higher learning rates, and acts as a regularizer that reduces overfitting in deep convolutional stacks.

### Why do CNNs use pooling layers instead of just strided convolutions?

While strided convolutions can achieve downsampling, **pooling layers** provide explicit translation invariance through max or average aggregation, making the network robust to small input shifts. The [`pooling.md`](https://github.com/d2l-ai/d2l-zh/blob/main/pooling.md) chapter notes that pooling also introduces a form of structural regularization by discarding exact spatial coordinates, which often improves generalization.

### What distinguishes the LeNet architecture in d2l-zh?

The LeNet implementation in [`lenet.md`](https://github.com/d2l-ai/d2l-zh/blob/main/lenet.md) represents the canonical CNN template that combines all eight architectural components: two convolution-pooling blocks for feature extraction followed by three fully-connected layers for classification. This pattern established the Conv-BN-ReLU-Pool block structure that modern architectures like ResNet still follow, though with added skip connections and greater depth.