Using CNNs for Computer Vision Tasks: From LeNet to ResNet

Convolutional Neural Networks (CNNs) provide translation-equivariant feature extraction through weight-sharing convolutional layers and residual shortcuts, enabling efficient image classification, detection, and segmentation pipelines as implemented in the ai-engineering-from-scratch curriculum.

The rohitg00/ai-engineering-from-scratch repository contains a comprehensive Vision phase (Phase 4) that demonstrates how to build modern computer vision systems using CNNs. This educational resource walks through the architectural evolution from classic LeNet-5 to residual networks, providing lightweight, production-ready implementations that illustrate how each building block contributes to hierarchical feature learning.

Core CNN Architecture Concepts

The lesson "CNNs — LeNet to ResNet" located at phases/04-computer-vision/03-cnns-lenet-to-resnet/docs/en.md explains the theoretical foundations that make CNNs effective for image processing.

Convolution and Pooling

Convolutional layers enforce spatial locality and weight sharing, allowing the network to learn edge detectors that generalize across the image plane. When combined with pooling operations, these layers create hierarchical receptive fields that capture increasingly complex patterns—from simple edges to object parts—while reducing spatial dimensions.

Depth, Bottlenecks, and Residual Connections

Modern CNNs increase representational power through depth while maintaining computational efficiency using bottleneck designs. Residual shortcuts (skip connections) allow gradients to flow directly through very deep networks, solving the vanishing gradient problem that plagued earlier architectures. The repository implements these concepts alongside batch normalization and ReLU activations to standardize activations and accelerate convergence.

Reference Implementations in the Repository

The file phases/04-computer-vision/03-cnns-lenet-to-resnet/code/main.py contains three deliberately lightweight implementations (approximately 100 lines each) that demonstrate the architectural progression:

LeNet5 for MNIST-Style Grayscale Images

The LeNet5 class (defined at line 6) provides a classic two-layer convolutional architecture suitable for 28×28 grayscale inputs:

import torch
from phases.04_computer_vision.03_cnns_lenet_to_resnet.code.main import LeNet5

# Initialize for 10-class classification

net = LeNet5(num_classes=10)
x = torch.randn(8, 1, 28, 28)  # batch of 8 grayscale images

logits = net(x)                # → (8, 10)

MiniVGG for Color Image Classification

The MiniVGG class (defined at line 40) stacks VGG-style blocks using 3×3 convolutions with batch normalization and ReLU activations, designed for 32×32 RGB inputs:

net = MiniVGG(num_classes=10)
x = torch.randn(8, 3, 32, 32)  # batch of 8 color images

logits = net(x)                # → (8, 10)

TinyResNet as a Modern Backbone

The TinyResNet class (defined at line 80) implements a residual network with four groups of BasicBlock modules, illustrating shortcut connections for 64×64 inputs:

net = TinyResNet(num_classes=100)
x = torch.randn(8, 3, 64, 64)  # higher resolution inputs

logits = net(x)                # → (8, 100)

Building End-to-End Vision Pipelines

The repository demonstrates how these CNN backbones integrate into complete computer vision systems:

Running the Examples

Execute the reference implementations using the repository's entry point script. The main() function (line 27) prints a concise model summary with per-group parameter counts:

python phases/04-computer-vision/03-cnns-lenet-to-resnet/code/main.py

For production deployment decisions, consult the prompt utility at phases/04-computer-vision/03-cnns-lenet-to-resnet/outputs/prompt-backbone-selector.md, which selects appropriate CNN architectures based on task requirements and resource constraints.

Summary

  • Architectural Evolution: The repository provides progressive implementations from LeNet5 (simple grayscale classifier) to TinyResNet (deep residual network), illustrating how convolutional layers, pooling, and skip connections improve feature extraction.

  • Translation Equivariance: CNNs leverage weight sharing and local receptive fields to detect patterns regardless of spatial position, making them inherently robust to object translation within images.

  • Production Integration: These backbones serve as the foundation for downstream tasks including classification, YOLO-based detection, Mask R-CNN segmentation, and DeepSORT tracking.

  • Educational Clarity: Each implementation in main.py remains under 100 lines, allowing direct inspection of forward passes, parameter counting, and architectural experimentation.

Frequently Asked Questions

What is the difference between LeNet5 and ResNet architectures?

LeNet5 uses a shallow stack of convolutional and pooling layers connected to fully-connected classifiers, suitable for simple grayscale datasets like MNIST. ResNet (Residual Network) introduces skip connections that bypass convolutional blocks, allowing the network to learn residual mappings rather than direct transformations and enabling training of much deeper architectures (50+ layers) without vanishing gradients, as implemented in the TinyResNet class.

When should I choose MiniVGG versus TinyResNet for a computer vision task?

Select MiniVGG for small-scale color image classification (32×32 inputs) where simplicity and fast training are priorities; its stacked 3×3 convolution blocks provide sufficient representational power for CIFAR-10 style datasets. Choose TinyResNet when processing higher resolution inputs (64×64+) or when you need deeper feature hierarchies for complex patterns, as the residual shortcuts allow stable training of deeper architectures with four distinct feature groups.

How do residual connections prevent vanishing gradients in deep CNNs?

Residual connections create shortcut paths that allow gradients to flow directly from deeper layers back to earlier ones during backpropagation. By reformulating the learning objective to fit a residual mapping (F(x) = H(x) - x) rather than the direct underlying mapping, the network ensures that gradients cannot completely vanish—even if the residual block learns nothing, the identity shortcut preserves gradient flow, which is critical for training networks with 50+ layers.

Can these CNN implementations be extended for object detection tasks?

Yes, the TinyResNet backbone can serve as a feature extractor for detection frameworks like YOLO or Mask R-CNN by removing the final classification head and attaching detection-specific heads that predict bounding box coordinates and class probabilities. The repository demonstrates this pattern in Phase 4 Lessons 06 and 08, where CNN backbones feed into region proposal networks and feature pyramid networks for object detection and instance segmentation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →