# How Residual Connections Enable Training of Very Deep Neural Networks

> Discover how residual connections enable training of very deep neural networks by preserving gradient flow and preventing the vanishing gradient problem. Learn about identity shortcut paths and residual learning.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**TLDR:** Residual connections allow neural networks to learn residual functions (the difference between desired output and input) rather than unreferenced mappings, creating identity shortcut paths that preserve gradient flow during backpropagation and prevent the vanishing gradient problem in deep architectures.

Training very deep neural networks was notoriously difficult before the introduction of residual connections, which fundamentally changed how gradient descent optimizes layered representations. In the `karpathy/nn-zero-to-hero` educational repository, Andrej Karpathy identifies residual connections as critical building blocks for modern deep learning, noting them as essential topics for future lectures. These skip connections solve the degradation problem by allowing layers to learn incremental improvements rather than complete transformations from scratch.

## The Vanishing Gradient Problem in Deep Architectures

Without residual connections, each layer in a deep network must learn a direct mapping from input to output. As networks grow beyond 20 to 30 layers, gradients computed during backpropagation shrink exponentially as they propagate backward through successive non-linear transformations. This **vanishing gradient phenomenon** makes it nearly impossible for stochastic gradient descent to update weights in early layers effectively, causing deep networks to perform worse than their shallower counterparts despite having greater representational capacity.

## How Residual Connections Reformulate Learning

Rather than forcing a stack of layers to approximate a desired mapping **H(x)**, residual connections reformulate the problem so the layers learn a residual function **F(x) := H(x) - x**. This subtle shift dramatically alters the optimization landscape.

### The Mathematical Framework

The core operation appears deceptively simple. For an input tensor **x** and a composite function **F** representing one or more layers (typically convolutions, batch normalization, and activations), the forward pass becomes:

```

y = F(x) + x

```

This additive formulation appears in the ResNet architecture and is emphasized in deep learning curricula as the foundation for building networks with hundreds of layers. By learning the residual rather than the absolute mapping, the network can easily push the residual to zero if the optimal function is close to the identity, avoiding the difficulties of learning identity transformations through multiple non-linear layers.

### Gradient Flow During Backpropagation

During backpropagation, the additive identity path creates a direct highway for gradients. The derivative of the output with respect to the input includes a term equal to **1** (from the identity connection **∂(x)/∂x**) alongside the gradient flowing through **F**. According to the `nn-zero-to-hero` source code analysis, this ensures that even if the Jacobian of **F** becomes ill-conditioned or vanishing, the gradient **∂L/∂x** can still propagate effectively to earlier layers, maintaining strong learning signals throughout the entire depth of the network.

## Implementing Residual Blocks in PyTorch

The following implementation follows patterns established in the `nn-zero-to-hero` repository, particularly building on concepts from `lectures/makemore/makemore_part5_cnn1.ipynb`. This example demonstrates how to construct deep, stable networks using residual connections:

```python
import torch
import torch.nn as nn
import torch.nn.functional as F

class ResidualBlock(nn.Module):
    """A simple residual block with two 3×3 conv layers."""
    def __init__(self, channels: int):
        super().__init__()
        self.conv1 = nn.Conv2d(channels, channels, kernel_size=3, padding=1)
        self.bn1   = nn.BatchNorm2d(channels)
        self.conv2 = nn.Conv2d(channels, channels, kernel_size=3, padding=1)
        self.bn2   = nn.BatchNorm2d(channels)

    def forward(self, x):
        out = F.relu(self.bn1(self.conv1(x)))
        out = self.bn2(self.conv2(out))
        # Skip connection: add the original input

        out = out + x
        return F.relu(out)

# Stack many residual blocks – this model can easily be 50+ layers deep.

class DeepResNet(nn.Module):
    def __init__(self, num_blocks: int = 20, channels: int = 64):
        super().__init__()
        self.input_conv = nn.Conv2d(3, channels, kernel_size=3, padding=1)
        self.blocks = nn.Sequential(*[ResidualBlock(channels) for _ in range(num_blocks)])
        self.pool = nn.AdaptiveAvgPool2d(1)
        self.fc   = nn.Linear(channels, 10)   # e.g. CIFAR-10

    def forward(self, x):
        x = F.relu(self.input_conv(x))
        x = self.blocks(x)
        x = self.pool(x).view(x.size(0), -1)
        return self.fc(x)

# Example usage

model = DeepResNet(num_blocks=30)   # 30 residual blocks → >60 convolutional layers

dummy_input = torch.randn(8, 3, 32, 32)   # batch of 8 CIFAR-10 images

logits = model(dummy_input)
print(logits.shape)   # → torch.Size([8, 10])

```

Key implementation details illustrated in this code:

- **Skip connection**: The critical line `out = out + x` creates the additive identity path that allows gradient flow
- **Depth scaling**: The `DeepResNet` class demonstrates how to stack 30 residual blocks to create networks with over 60 convolutional layers without training degradation
- **Activation placement**: The final ReLU appears after the addition, preserving the identity path's linearity and ensuring gradients can flow unimpeded through the shortcut

## Synergy with Batch Normalization

While residual connections solve gradient flow, they achieve maximum effectiveness when combined with **Batch Normalization**, a technique covered extensively in `lectures/makemore/makemore_part3_bn.ipynb`. Batch normalization normalizes layer inputs to reduce internal covariate shift, preventing exploding activations that could destabilize deep residual networks. The [`README.md`](https://github.com/karpathy/nn-zero-to-hero/blob/main/README.md) in `karpathy/nn-zero-to-hero` notes that both techniques together enable the training of "very deep networks" that would otherwise be impossible to optimize.

## Summary

- Residual connections reformulate layer learning from direct mapping **H(x)** to residual learning **F(x) + x**, making optimization easier for deep architectures
- The additive identity path creates a **gradient highway** that prevents vanishing gradients during backpropagation, allowing training of networks with 100+ layers
- Implementation requires only the single addition operation `out = out + x`, but this simple change enables stable training of arbitrarily deep models
- Residual blocks act as implicit regularizers, allowing the network to fall back to identity mappings if additional depth provides no benefit
- When combined with Batch Normalization (as taught in `makemore_part3_bn.ipynb`), residual connections enable practical training of very deep convolutional networks on standard hardware

## Frequently Asked Questions

### What is the difference between residual connections and skip connections?

In practice, the terms are often used interchangeably, though "residual connection" specifically refers to the additive skip where the input is added to the transformed output (**F(x) + x**). Skip connections in general can refer to any pathway that bypasses layers, including concatenative skips (as in DenseNet) where tensors are joined along the channel dimension rather than added.

### Do residual connections prevent overfitting?

Residual connections act as an **implicit regularizer** by allowing the network to easily learn identity mappings. If a residual block adds no value, the network can set its weights to near zero and pass the signal through unchanged, effectively reducing the model's complexity during training. However, they do not replace explicit regularization techniques like dropout or weight decay.

### How many residual blocks can I stack?

Modern architectures commonly use 50 to 152 residual blocks in computer vision tasks, and some implementations exceed 1,000 layers. The key constraint is not gradient flow (which residual connections preserve) but computational cost and memory usage. The `nn-zero-to-hero` repository demonstrates that even 30 blocks (over 60 convolutional layers) train stably on standard hardware.

### Are residual connections only for convolutional networks?

No. While the code example in `makemore_part5_cnn1.ipynb` focuses on convolutional architectures, residual connections are equally effective in fully-connected networks, transformers, and recurrent architectures. The Transformer architecture relies heavily on residual connections around its self-attention and feed-forward layers, demonstrating that this technique generalizes across all deep learning domains.