How Backpropagation Works in Neural Networks from Scratch: A Complete Implementation Guide

Backpropagation is the algorithm that computes the gradient of a loss function with respect to every learnable parameter by applying the chain rule in reverse, starting from the output layer and propagating error signals backward to enable gradient descent updates.

Backpropagation forms the mathematical foundation of modern deep learning. In the rohitg00/ai-engineering-from-scratch repository, the algorithm is implemented from first principles without automatic differentiation libraries, revealing exactly how gradients flow through each layer. This guide examines the concrete implementation across two projects: a classic two-layer XOR network and a hand-coded Variational Auto-Encoder.

The Four Stages of Backpropagation

Every training iteration follows a strict sequence. First, the forward pass feeds input through the network, computing linear transformations and activation outputs. Second, loss evaluation measures the discrepancy between prediction and target. Third, the backward pass applies the chain rule to propagate error signals from output to input. Finally, parameter updates adjust weights and biases using the computed gradients.

This pattern remains identical whether training a simple perceptron or a complex generative model.

Backpropagation from Scratch: Two Reference Implementations

Two-Layer XOR Network

The repository's perceptron implementation in phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py demonstrates backpropagation on a minimal fully-connected network learning the XOR truth table. The TwoLayerNetwork class stores intermediate activations during the forward pass in self.hidden_outputs and self.output for later reuse.

During backpropagation, the algorithm first computes the output error d_output by combining the prediction error with the derivative of the sigmoid activation: error * output * (1 - output). This error signal propagates backward to the hidden layer as hidden_deltas, which are then used to calculate gradients for both weight matrices and bias vectors.


# From phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py

net = TwoLayerNetwork(learning_rate=2.0)
net.train(xor_data, epochs=10000)

After approximately 10,000 epochs, the network correctly learns the non-linear XOR function, proving that even a raw Python implementation of backpropagation converges effectively.

Variational Auto-Encoder with Manual Gradients

The VAE implementation in phases/08-generative-ai/02-autoencoders-vae/code/main.py extends backpropagation to probabilistic models. The backward function computes gradients for both the reconstruction loss and the KL-divergence term, handling the re-parameterization trick that introduces additional gradient paths through d_mu_recon and d_lv_recon.


# From phases/08-generative-ai/02-autoencoders-vae/code/main.py

params = init_vae(in_dim, hidden, z_dim, rng)
for x in data:
    eps = [rng.gauss(0, 1) for _ in range(z_dim)]
    fwd = forward(x, params, eps)
    loss, _, _ = loss_value(x, fwd, beta)
    grads = backward(x, fwd, params, beta)
    apply_update(params, grads, lr)

The function returns a matching-shaped gradient dictionary that mirrors the parameter structure, allowing the apply_update function to perform param -= lr * grad for every weight and bias.

The Mathematical Pattern of Gradient Flow

Both implementations follow an identical mathematical sequence, revealing the mechanical nature of backpropagation:

Forward Computation. Each layer computes z = W·x + b followed by an activation function like sigmoid or tanh. The code explicitly stores these intermediate values (self.hidden_outputs.append(self.sigmoid(z)) in the perceptron, or h_enc = tanh(add(matmul(enc["W1"], x), enc["b1"])) in the VAE) because they are required for the backward pass.

Loss Gradient with Respect to Output. The derivative of the loss relative to the network output provides the initial error signal. For mean-squared error with sigmoid outputs, this becomes d_output = (target - output) * output * (1 - output) as implemented in the XOR network.

Backpropagation Through Activations. The error signal gets multiplied by the derivative of the activation function. For sigmoid, this is output * (1 - output); for tanh, it is 1 - h^2. The VAE code demonstrates this explicitly: d_pre_enc = [dg * g for dg, g in zip(d_h_enc, tanh_grad(h_enc))].

Backpropagation Through Linear Layers. To propagate error to the previous layer, the implementation uses the transposed weight matrix. In the VAE, this appears as d_h_dec = [sum(dec["W_out"][i][j] * d_x_hat[i] for i ...) for j ...], mapping output errors back to hidden layer errors.

Parameter Gradients. Gradients for weights are computed as the outer product of the error signal and the incoming activation. The VAE constructs weight gradients with grads["dec"]["W_out"] = [[d * h for h in h_dec] for d in d_x_hat].

Parameter Updates. Finally, weights are adjusted in the opposite direction of the gradient: row[j] -= lr * g[i][j], where lr represents the learning rate.

Summary

  • Backpropagation applies the chain rule in reverse to compute gradients efficiently across computational graphs.
  • The rohitg00/ai-engineering-from-scratch repository implements this in phases/03-deep-learning-core/01-the-perceptron/code/perceptron.py using a two-layer XOR solver and in phases/08-generative-ai/02-autoencoders-vae/code/main.py using a manual VAE.
  • Forward pass caching is essential for gradient computation during the backward pass.
  • Gradient flow follows a strict sequence: loss derivative → activation derivative → linear transformation → parameter gradients.
  • Manual implementations explicitly show the matrix transposes and outer products that autograd systems hide, making the learning process transparent.

Frequently Asked Questions

How does backpropagation calculate gradients without automatic differentiation?

Backpropagation manually applies the chain rule from calculus. Each layer stores its inputs and outputs during the forward pass. During the backward pass, the algorithm receives an error signal from the subsequent layer, multiplies it by the derivative of the local activation function, and uses the transposed weight matrix to compute the error for the previous layer. This explicit computation of derivatives replaces the need for autograd systems.

Why does the XOR network need backpropagation to learn?

The XOR function is not linearly separable, meaning a single-layer perceptron cannot learn it. The two-layer network introduces hidden units that create non-linear decision boundaries. Backpropagation computes how each weight in both layers contributes to the prediction error, allowing the network to adjust the hidden layer representations until the non-linear function is approximated correctly, typically converging after around 10,000 epochs with a learning rate of 2.0.

What is the re-parameterization trick in the VAE backpropagation code?

The re-parameterization trick allows gradients to flow through random sampling operations. Instead of sampling directly from the latent distribution N(μ, σ), the code samples from a standard normal ε and computes z = μ + σ * ε. During backpropagation, gradients flow through d_mu_recon and d_lv_recon (log-variance), making it possible to optimize the stochastic nodes using standard gradient descent while maintaining the probabilistic nature of the model.

How are bias gradients computed differently from weight gradients?

Bias gradients are simpler because they do not depend on the input activations—they depend only on the error signal flowing into the layer. While weight gradients require the outer product of the error signal and the input activation (as seen in grads["dec"]["W_out"] calculations), bias gradients are simply the sum of the error signals across the batch dimension, computed as d_b = sum(d_z) for layer z.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →