# How Backpropagation Works in Neural Networks from Scratch: A Complete Implementation Guide

> Understand how backpropagation works in neural networks from scratch. Learn to implement this core algorithm step by step and enable gradient descent updates.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Backpropagation is the algorithm that computes the gradient of a loss function with respect to every learnable parameter by applying the chain rule in reverse, starting from the output layer and propagating error signals backward to enable gradient descent updates.**

Backpropagation forms the mathematical foundation of modern deep learning. In the `rohitg00/ai-engineering-from-scratch` repository, the algorithm is implemented from first principles without automatic differentiation libraries, revealing exactly how gradients flow through each layer. This guide examines the concrete implementation across two projects: a classic two-layer XOR network and a hand-coded Variational Auto-Encoder.

## The Four Stages of Backpropagation

Every training iteration follows a strict sequence. First, the **forward pass** feeds input through the network, computing linear transformations and activation outputs. Second, **loss evaluation** measures the discrepancy between prediction and target. Third, the **backward pass** applies the chain rule to propagate error signals from output to input. Finally, **parameter updates** adjust weights and biases using the computed gradients.

This pattern remains identical whether training a simple perceptron or a complex generative model.

## Backpropagation from Scratch: Two Reference Implementations

### Two-Layer XOR Network

The repository's perceptron implementation in [`phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py) demonstrates backpropagation on a minimal fully-connected network learning the XOR truth table. The `TwoLayerNetwork` class stores intermediate activations during the forward pass in `self.hidden_outputs` and `self.output` for later reuse.

During backpropagation, the algorithm first computes the output error `d_output` by combining the prediction error with the derivative of the sigmoid activation: `error * output * (1 - output)`. This error signal propagates backward to the hidden layer as `hidden_deltas`, which are then used to calculate gradients for both weight matrices and bias vectors.

```python

# From phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py

net = TwoLayerNetwork(learning_rate=2.0)
net.train(xor_data, epochs=10000)

```

After approximately 10,000 epochs, the network correctly learns the non-linear XOR function, proving that even a raw Python implementation of backpropagation converges effectively.

### Variational Auto-Encoder with Manual Gradients

The VAE implementation in [`phases/08-generative-ai/02-autoencoders-vae/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/02-autoencoders-vae/code/main.py) extends backpropagation to probabilistic models. The `backward` function computes gradients for both the reconstruction loss and the KL-divergence term, handling the re-parameterization trick that introduces additional gradient paths through `d_mu_recon` and `d_lv_recon`.

```python

# From phases/08-generative-ai/02-autoencoders-vae/code/main.py

params = init_vae(in_dim, hidden, z_dim, rng)
for x in data:
    eps = [rng.gauss(0, 1) for _ in range(z_dim)]
    fwd = forward(x, params, eps)
    loss, _, _ = loss_value(x, fwd, beta)
    grads = backward(x, fwd, params, beta)
    apply_update(params, grads, lr)

```

The function returns a matching-shaped gradient dictionary that mirrors the parameter structure, allowing the `apply_update` function to perform `param -= lr * grad` for every weight and bias.

## The Mathematical Pattern of Gradient Flow

Both implementations follow an identical mathematical sequence, revealing the mechanical nature of backpropagation:

**Forward Computation.** Each layer computes `z = W·x + b` followed by an activation function like `sigmoid` or `tanh`. The code explicitly stores these intermediate values (`self.hidden_outputs.append(self.sigmoid(z))` in the perceptron, or `h_enc = tanh(add(matmul(enc["W1"], x), enc["b1"]))` in the VAE) because they are required for the backward pass.

**Loss Gradient with Respect to Output.** The derivative of the loss relative to the network output provides the initial error signal. For mean-squared error with sigmoid outputs, this becomes `d_output = (target - output) * output * (1 - output)` as implemented in the XOR network.

**Backpropagation Through Activations.** The error signal gets multiplied by the derivative of the activation function. For sigmoid, this is `output * (1 - output)`; for tanh, it is `1 - h^2`. The VAE code demonstrates this explicitly: `d_pre_enc = [dg * g for dg, g in zip(d_h_enc, tanh_grad(h_enc))]`.

**Backpropagation Through Linear Layers.** To propagate error to the previous layer, the implementation uses the transposed weight matrix. In the VAE, this appears as `d_h_dec = [sum(dec["W_out"][i][j] * d_x_hat[i] for i ...) for j ...]`, mapping output errors back to hidden layer errors.

**Parameter Gradients.** Gradients for weights are computed as the outer product of the error signal and the incoming activation. The VAE constructs weight gradients with `grads["dec"]["W_out"] = [[d * h for h in h_dec] for d in d_x_hat]`.

**Parameter Updates.** Finally, weights are adjusted in the opposite direction of the gradient: `row[j] -= lr * g[i][j]`, where `lr` represents the learning rate.

## Summary

- **Backpropagation** applies the chain rule in reverse to compute gradients efficiently across computational graphs.
- The `rohitg00/ai-engineering-from-scratch` repository implements this in [`phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py) using a two-layer XOR solver and in [`phases/08-generative-ai/02-autoencoders-vae/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/02-autoencoders-vae/code/main.py) using a manual VAE.
- **Forward pass** caching is essential for gradient computation during the backward pass.
- **Gradient flow** follows a strict sequence: loss derivative → activation derivative → linear transformation → parameter gradients.
- **Manual implementations** explicitly show the matrix transposes and outer products that autograd systems hide, making the learning process transparent.

## Frequently Asked Questions

### How does backpropagation calculate gradients without automatic differentiation?

Backpropagation manually applies the chain rule from calculus. Each layer stores its inputs and outputs during the forward pass. During the backward pass, the algorithm receives an error signal from the subsequent layer, multiplies it by the derivative of the local activation function, and uses the transposed weight matrix to compute the error for the previous layer. This explicit computation of derivatives replaces the need for autograd systems.

### Why does the XOR network need backpropagation to learn?

The XOR function is not linearly separable, meaning a single-layer perceptron cannot learn it. The two-layer network introduces hidden units that create non-linear decision boundaries. Backpropagation computes how each weight in both layers contributes to the prediction error, allowing the network to adjust the hidden layer representations until the non-linear function is approximated correctly, typically converging after around 10,000 epochs with a learning rate of 2.0.

### What is the re-parameterization trick in the VAE backpropagation code?

The re-parameterization trick allows gradients to flow through random sampling operations. Instead of sampling directly from the latent distribution `N(μ, σ)`, the code samples from a standard normal `ε` and computes `z = μ + σ * ε`. During backpropagation, gradients flow through `d_mu_recon` and `d_lv_recon` (log-variance), making it possible to optimize the stochastic nodes using standard gradient descent while maintaining the probabilistic nature of the model.

### How are bias gradients computed differently from weight gradients?

Bias gradients are simpler because they do not depend on the input activations—they depend only on the error signal flowing into the layer. While weight gradients require the outer product of the error signal and the input activation (as seen in `grads["dec"]["W_out"]` calculations), bias gradients are simply the sum of the error signals across the batch dimension, computed as `d_b = sum(d_z)` for layer `z`.