# Using CNNs for Computer Vision Tasks: From LeNet to ResNet

> Learn how to use CNNs for computer vision tasks like image classification and detection. Explore the evolution from LeNet to ResNet with efficient implementations. Discover the ai-engineering-from-scratch curriculum.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-19

---

**Convolutional Neural Networks (CNNs) provide translation-equivariant feature extraction through weight-sharing convolutional layers and residual shortcuts, enabling efficient image classification, detection, and segmentation pipelines as implemented in the `ai-engineering-from-scratch` curriculum.**

The `rohitg00/ai-engineering-from-scratch` repository contains a comprehensive Vision phase (Phase 4) that demonstrates how to build modern computer vision systems using CNNs. This educational resource walks through the architectural evolution from classic LeNet-5 to residual networks, providing lightweight, production-ready implementations that illustrate how each building block contributes to hierarchical feature learning.

## Core CNN Architecture Concepts

The lesson **"CNNs — LeNet to ResNet"** located at [`phases/04-computer-vision/03-cnns-lenet-to-resnet/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/04-computer-vision/03-cnns-lenet-to-resnet/docs/en.md) explains the theoretical foundations that make CNNs effective for image processing.

### Convolution and Pooling

**Convolutional layers** enforce spatial locality and weight sharing, allowing the network to learn edge detectors that generalize across the image plane. When combined with **pooling operations**, these layers create hierarchical receptive fields that capture increasingly complex patterns—from simple edges to object parts—while reducing spatial dimensions.

### Depth, Bottlenecks, and Residual Connections

Modern CNNs increase representational power through depth while maintaining computational efficiency using bottleneck designs. **Residual shortcuts** (skip connections) allow gradients to flow directly through very deep networks, solving the vanishing gradient problem that plagued earlier architectures. The repository implements these concepts alongside **batch normalization** and **ReLU activations** to standardize activations and accelerate convergence.

## Reference Implementations in the Repository

The file [`phases/04-computer-vision/03-cnns-lenet-to-resnet/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/04-computer-vision/03-cnns-lenet-to-resnet/code/main.py) contains three deliberately lightweight implementations (approximately 100 lines each) that demonstrate the architectural progression:

### LeNet5 for MNIST-Style Grayscale Images

The `LeNet5` class (defined at line 6) provides a classic two-layer convolutional architecture suitable for 28×28 grayscale inputs:

```python
import torch
from phases.04_computer_vision.03_cnns_lenet_to_resnet.code.main import LeNet5

# Initialize for 10-class classification

net = LeNet5(num_classes=10)
x = torch.randn(8, 1, 28, 28)  # batch of 8 grayscale images

logits = net(x)                # → (8, 10)

```

### MiniVGG for Color Image Classification

The `MiniVGG` class (defined at line 40) stacks VGG-style blocks using 3×3 convolutions with batch normalization and ReLU activations, designed for 32×32 RGB inputs:

```python
net = MiniVGG(num_classes=10)
x = torch.randn(8, 3, 32, 32)  # batch of 8 color images

logits = net(x)                # → (8, 10)

```

### TinyResNet as a Modern Backbone

The `TinyResNet` class (defined at line 80) implements a residual network with four groups of `BasicBlock` modules, illustrating shortcut connections for 64×64 inputs:

```python
net = TinyResNet(num_classes=100)
x = torch.randn(8, 3, 64, 64)  # higher resolution inputs

logits = net(x)                # → (8, 100)

```

## Building End-to-End Vision Pipelines

The repository demonstrates how these CNN backbones integrate into complete computer vision systems:

- **Image Classification**: CNN feature maps feed directly into classifier heads that map spatial features to class logits, as detailed in [`phases/04-computer-vision/04-image-classification/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/04-computer-vision/04-image-classification/docs/en.md).

- **Object Detection**: The CNN backbone serves as a feature extractor for detectors like YOLO, with task-specific heads added for bounding box regression and classification, covered in [`phases/04-computer-vision/06-object-detection-yolo/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/04-computer-vision/06-object-detection-yolo/docs/en.md).

- **Instance Segmentation and Tracking**: CNN features power mask prediction in Mask R-CNN and appearance embeddings in DeepSORT multi-object tracking systems, documented in [`phases/04-computer-vision/08-instance-segmentation-mask-rcnn/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/04-computer-vision/08-instance-segmentation-mask-rcnn/docs/en.md) and [`phases/04-computer-vision/27-multi-object-tracking/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/04-computer-vision/27-multi-object-tracking/docs/en.md).

## Running the Examples

Execute the reference implementations using the repository's entry point script. The `main()` function (line 27) prints a concise model summary with per-group parameter counts:

```bash
python phases/04-computer-vision/03-cnns-lenet-to-resnet/code/main.py

```

For production deployment decisions, consult the prompt utility at [`phases/04-computer-vision/03-cnns-lenet-to-resnet/outputs/prompt-backbone-selector.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/04-computer-vision/03-cnns-lenet-to-resnet/outputs/prompt-backbone-selector.md), which selects appropriate CNN architectures based on task requirements and resource constraints.

## Summary

- **Architectural Evolution**: The repository provides progressive implementations from LeNet5 (simple grayscale classifier) to TinyResNet (deep residual network), illustrating how convolutional layers, pooling, and skip connections improve feature extraction.

- **Translation Equivariance**: CNNs leverage weight sharing and local receptive fields to detect patterns regardless of spatial position, making them inherently robust to object translation within images.

- **Production Integration**: These backbones serve as the foundation for downstream tasks including classification, YOLO-based detection, Mask R-CNN segmentation, and DeepSORT tracking.

- **Educational Clarity**: Each implementation in [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py) remains under 100 lines, allowing direct inspection of forward passes, parameter counting, and architectural experimentation.

## Frequently Asked Questions

### What is the difference between LeNet5 and ResNet architectures?

**LeNet5** uses a shallow stack of convolutional and pooling layers connected to fully-connected classifiers, suitable for simple grayscale datasets like MNIST. **ResNet** (Residual Network) introduces skip connections that bypass convolutional blocks, allowing the network to learn residual mappings rather than direct transformations and enabling training of much deeper architectures (50+ layers) without vanishing gradients, as implemented in the `TinyResNet` class.

### When should I choose MiniVGG versus TinyResNet for a computer vision task?

Select **MiniVGG** for small-scale color image classification (32×32 inputs) where simplicity and fast training are priorities; its stacked 3×3 convolution blocks provide sufficient representational power for CIFAR-10 style datasets. Choose **TinyResNet** when processing higher resolution inputs (64×64+) or when you need deeper feature hierarchies for complex patterns, as the residual shortcuts allow stable training of deeper architectures with four distinct feature groups.

### How do residual connections prevent vanishing gradients in deep CNNs?

Residual connections create shortcut paths that allow gradients to flow directly from deeper layers back to earlier ones during backpropagation. By reformulating the learning objective to fit a residual mapping (F(x) = H(x) - x) rather than the direct underlying mapping, the network ensures that gradients cannot completely vanish—even if the residual block learns nothing, the identity shortcut preserves gradient flow, which is critical for training networks with 50+ layers.

### Can these CNN implementations be extended for object detection tasks?

Yes, the `TinyResNet` backbone can serve as a feature extractor for detection frameworks like YOLO or Mask R-CNN by removing the final classification head and attaching detection-specific heads that predict bounding box coordinates and class probabilities. The repository demonstrates this pattern in Phase 4 Lessons 06 and 08, where CNN backbones feed into region proposal networks and feature pyramid networks for object detection and instance segmentation.