1x1 vs 3x3 vs Dilated Convolutions in CNNs: Key Differences Explained

1x1 convolutions perform channel-wise dimensionality reduction, 3x3 convolutions extract local spatial features with a 3×3 receptive field, and dilated convolutions expand the receptive field exponentially without adding parameters or sacrificing resolution.

The scutan90/DeepLearning-500-questions repository categorizes these kernel types in ch05_卷积神经网络(CNN)/第五章_卷积神经网络(CNN).md, distinguishing standard convolutions (encompassing both 1×1 and 3×3 kernels) from dilated (空洞) convolutions based on their sampling strategies and effective receptive field calculations.

1×1 Convolutions: Channel Mixing and Dimensionality Reduction

A 1×1 convolution—despite its small spatial footprint—operates across all input channels at each spatial location. According to the repository’s classification of standard convolutions in lines 30–35, this kernel size functions as a fully-connected layer applied per-pixel, making it ideal for bottleneck designs.

  • Parameters: 1 × in_channels × out_channels (minimal spatial parameters)
  • Receptive field: 1×1 (single pixel, no spatial context)
  • Primary use: Reducing channel depth before expensive 3×3 operations (ResNet bottleneck blocks) or cross-channel feature mixing (Network-in-Network architectures)

3×3 Convolutions: Standard Local Feature Extraction

The 3×3 standard convolution represents the default spatial feature extractor in modern CNNs. As documented in the repository’s CNN chapter, this kernel size strikes a balance between representational capacity and parameter efficiency.

  • Parameters: 9 × in_channels × out_channels
  • Receptive field: Exactly 3×3 (no dilation)
  • Computation: O(9 × in × out × output_area)
  • Typical application: Early convolutional layers, stacked depthwise-separable blocks, and general local pattern detection (edges, textures)

Dilated (Atrous) Convolutions: Expanding Context Without Pooling

Dilated convolutions insert zeros between kernel elements to "space out" sampling points while maintaining the same physical kernel size. The repository explicitly defines dilated (空洞) convolution at line 1005 as distinct from standard convolution due to its modified receptive field calculation.

Effective Receptive Field Formula

For a 3×3 kernel with dilation rate d:


Effective size = kernel_size + (kernel_size - 1) × (dilation - 1)

  • Dilation = 1: Standard 3×3 (3×3 effective field)
  • Dilation = 2: 5×5 effective field
  • Dilation = 4: 9×9 effective field

Key Characteristics

Aspect Standard 3×3 Dilated 3×3
Physical parameters 9 per channel 9 per channel (identical)
Multiply-add operations Standard count Same as standard (only memory access pattern changes)
Spatial resolution Preserved with stride=1 and padding=1 Preserved with padding = dilation × (kernel-1)/2
Gridding risk None Potential "gridding" artifacts if dilation rates are mismatched between layers

Practical Implementation in PyTorch

The following examples demonstrate the three convolution types using the same input tensor. Note that 1×1 and 3×3 standard convolutions differ only in kernel_size, while dilated convolution adds the dilation parameter.

import torch
import torch.nn as nn

# Input: batch=1, channels=64, 64×64 spatial

x = torch.randn(1, 64, 64, 64)

# 1. 1×1 Convolution: Channel reduction (64 → 128)

conv_1x1 = nn.Conv2d(
    in_channels=64, 
    out_channels=128, 
    kernel_size=1,      # 1×1 spatial coverage

    stride=1,
    padding=0           # No padding needed for 1×1

)

# 2. Standard 3×3 Convolution: Local feature extraction

conv_3x3 = nn.Conv2d(
    in_channels=64, 
    out_channels=128, 
    kernel_size=3,
    stride=1,
    padding=1,          # padding=(3-1)/2 = 1

    dilation=1          # Default, no dilation

)

# 3. Dilated 3×3 Convolution: 5×5 effective receptive field

conv_dilated = nn.Conv2d(
    in_channels=64,
    out_channels=128,
    kernel_size=3,
    stride=1,
    padding=2,          # padding=dilation*(kernel-1)/2 = 2

    dilation=2          # 1 zero between kernel elements

)

# Output shapes (all preserve 64×64 spatial resolution)

y1 = conv_1x1(x)   # Shape: (1, 128, 64, 64)

y2 = conv_3x3(x)   # Shape: (1, 128, 64, 64)

y3 = conv_dilated(x)  # Shape: (1, 128, 64, 64)

print(f"1×1 params: {sum(p.numel() for p in conv_1x1.parameters())}")
print(f"3×3 params: {sum(p.numel() for p in conv_3x3.parameters())}")
print(f"Dilated params: {sum(p.numel() for p in conv_dilated.parameters())}")

Parameter count verification: All three layers above use 128 output channels. The 1×1 layer has 64×128×1 = 8,192 weights plus 128 biases; both 3×3 layers (standard and dilated) have 64×128×9 = 73,728 weights plus 128 biases, confirming that dilation adds zero parameters.

Architectural Usage Patterns

Modern CNNs combine these three convolution types strategically:

  1. 1×1 bottleneck: Precede 3×3 layers to reduce computational load (ResNet, MobileNetV2 inverted residuals)
  2. Stacked 3×3 standard: Build local hierarchies in early networks (VGG-style stacks)
  3. Dilated cascades: Replace pooling in semantic segmentation backbones (DeepLab, PSPNet) to maintain high-resolution feature maps while capturing global context

According to the repository source, combining standard 3×3 stacks with dilated layers allows networks to learn local patterns first, then aggregate broader context without the information loss associated with max-pooling or strided convolutions.

Summary

  • 1×1 convolutions minimize parameters and perform cross-channel mixing with no spatial context, ideal for bottlenecks
  • 3×3 standard convolutions capture local 3×3 spatial patterns and serve as the workhorse for feature extraction
  • Dilated 3×3 convolutions match standard 3×3 in parameter count and computation but achieve larger receptive fields (5×5, 9×9, etc.) through spaced sampling
  • All three preserve spatial resolution when configured with appropriate padding, but only dilated convolutions avoid the need for deeper stacking or pooling to increase context

Frequently Asked Questions

What is the effective receptive field of a 3×3 convolution with dilation rate 2?

The effective receptive field expands to 5×5. Using the formula kernel_size + (kernel_size - 1) × (dilation - 1), a 3×3 kernel with dilation=2 covers 3 + (2 × 1) = 5 pixels in each dimension, matching the context area of a 5×5 standard kernel while using only 9 parameters instead of 25.

Does a 1×1 convolution have any spatial receptive field?

No. A 1×1 convolution operates on single spatial locations only, processing across the channel dimension. It cannot detect spatial patterns like edges or corners; it exclusively performs linear combinations of input channels to produce output channels at each pixel position.

Why use dilated convolution instead of just increasing the kernel size to 5×5 or 7×7?

Dilated convolution maintains the same parameter count and computational cost as a 3×3 kernel (9 weights) while achieving the receptive field of larger kernels. A standard 5×5 kernel requires 25 parameters and significantly more multiply-add operations. Additionally, dilated convolutions preserve spatial resolution without requiring pooling layers that discard fine-grained detail critical for dense prediction tasks like semantic segmentation.

Can I stack dilated convolutions with different rates to avoid gridding artifacts?

Yes. The repository notes that careless use of dilation can cause "gridding" where input pixels are sampled unevenly. To mitigate this, practitioners often apply hybrid dilated convolution (HDC)—stacking layers with carefully chosen dilation rates (e.g., 1, 2, 4) that collectively cover the receptive field without leaving gaps, or combining dilated layers with standard 3×3 convolutions to fill in between sampled points.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →