# How the Custom Neural Network Framework in OwnFramework.ipynb Implements Backpropagation

> Learn how OwnFramework.ipynb implements backpropagation using a computational graph and chain rule. Discover gradient accumulation and parameter updates for AI beginners.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: internals
- Published: 2026-08-27

---

**The OwnFramework.ipynb notebook from microsoft/AI-For-Beginners implements backpropagation through a computational graph where each layer class defines `forward()` and `backward()` methods to apply the chain rule, accumulating gradients in `dW` and `db` attributes before updating parameters via stochastic gradient descent.**

The educational neural network framework inside `OwnFramework.ipynb` demonstrates backpropagation from scratch using pure NumPy. Located in the `microsoft/AI-For-Beginners` repository at `lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb`, this implementation reveals how modern deep learning libraries compute gradients by traversing a computational graph in reverse.

## Computational Graph Architecture

The framework treats each neural network component as a node in a computational graph. Every layer class exposes two critical methods:

- `forward(x)` – computes the output and stores intermediates needed for differentiation
- `backward(δ)` – receives the upstream gradient δ and returns the downstream gradient using the chain rule

Trainable layers additionally accumulate parameter gradients in temporary attributes (`dW`, `db`) during the backward pass, then apply updates through an `update(lr)` method that performs stochastic gradient descent: θ ← θ – η·∂ℒ/∂θ.

## Layer-Level Backpropagation Implementation

The backpropagation mechanism delegates derivative computation to individual layers. Each layer analytically derives its local Jacobian and multiplies it by the incoming gradient δ.

### Linear Layer Gradients

In `OwnFramework.ipynb`, the `Linear` class implements a fully-connected layer where forward computation follows z = x·Wᵀ + b. During the backward pass defined in the same class, the layer computes three quantities:

- **dx = δ·W** – propagates the gradient to the previous layer
- **dW = δᵀ·x** – computes the gradient with respect to weights
- **db = δ.sum(axis=0)** – computes the gradient with respect to biases

These matrix operations automatically average gradients across the minibatch dimension.

### Softmax Jacobian Computation

The `Softmax` layer handles the non-trivial Jacobian of the softmax function. Rather than explicitly constructing the full Jacobian matrix, the implementation uses the efficient simplified form:

1. Re-computes probabilities: `p = forward(z)`
2. Element-wise product: `dp = p * δ`
3. Output gradient: `dx = dp - p * sum(dp, axis=1, keepdims=True)`

This formulation computes the vector-Jacobian product without materializing the full n×n Jacobian matrix.

### Cross-Entropy Loss Gradient

The `CrossEntropyLoss` class computes the gradient of the negative log-likelihood. For a batch of predictions p and true labels y, the backward method returns a sparse gradient where the entry corresponding to the true class receives −1/(N·p_y). The implementation returns `dlog_softmax / p` to be consumed directly by the preceding Softmax layer's backward method.

## Network Orchestration and Parameter Updates

A `Net` container class orchestrates the backward pass by iterating through layers in reverse order. According to the source code in `OwnFramework.ipynb`, the container:

1. Sequentially calls each layer's `forward` during the forward pass
2. Iterates layers in reverse during `backward(dp)`, feeding each layer's output gradient as the next layer's input δ
3. Invokes `update(lr)` on each trainable layer to apply the accumulated gradients

This design strictly follows the chain rule: the gradient flowing into a layer is multiplied by that layer's local derivatives to produce the gradient for preceding layers.

## Practical Training Example

The following pattern from `OwnFramework.ipynb` illustrates the complete forward-backward-update cycle:

```python

# Initialize layers

lin = Linear(nin=2, nout=2)
softmax = Softmax()
loss_fn = CrossEntropyLoss()

# Forward pass

z = lin.forward(x_batch)
p = softmax.forward(z)
loss = loss_fn.forward(p, y_batch)

# Backward pass (chain rule application)

dp = loss_fn.backward(loss)  # ∂ℒ/∂p

dz = softmax.backward(dp)     # ∂ℒ/∂z

dx = lin.backward(dz)         # ∂ℒ/∂x (for previous layer)

# Parameter update

lin.update(lr=0.01)           # SGD: W -= 0.01 * dW, b -= 0.01 * db

```

For a complete network using the `Net` container:

```python
net = Net()
net.add(Linear(784, 128))
net.add(Softmax())
loss_fn = CrossEntropyLoss()

# Training loop

for epoch in range(10):
    for i in range(0, len(train_x), batch_size):
        xb = train_x[i:i+batch_size]
        yb = train_labels[i:i+batch_size]
        
        # Forward

        p = net.forward(xb)
        loss = loss_fn.forward(p, yb)
        
        # Backward

        dp = loss_fn.backward(loss)
        net.backward(dp)
        
        # Update

        net.update(lr=0.1)

```

## Summary

- The custom neural network framework in OwnFramework.ipynb implements backpropagation through explicit `forward()` and `backward()` methods in each layer class.
- **Linear** layers compute gradients via matrix multiplication: `dW = δᵀ·x` and `db = δ.sum(axis=0)`.
- **Softmax** layers efficiently compute the vector-Jacobian product without constructing the full Jacobian matrix.
- The **Net** container orchestrates reverse-mode automatic differentiation by iterating layers backward and invoking `update(lr)` for stochastic gradient descent.
- All operations use NumPy matrix arithmetic to process entire minibatches simultaneously, automatically averaging gradients across the batch dimension.

## Frequently Asked Questions

### Where is the backward pass implemented in OwnFramework.ipynb?

The backward pass is implemented within individual layer classes (such as `Linear`, `Softmax`, and `CrossEntropyLoss`) located in `lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb`. Each layer's `backward(δ)` method computes the local gradient using the chain rule and returns the downstream gradient to the previous layer. The `Net` container class handles the reverse iteration through these layers.

### How does the Linear layer compute weight gradients?

According to the source code in `OwnFramework.ipynb`, the `Linear` layer computes weight gradients during the `backward` method using the formula `dW = δᵀ·x`, where δ is the upstream gradient and x is the cached input from the forward pass. Biases receive gradients via `db = δ.sum(axis=0)`, automatically summing across the batch dimension.

### Does OwnFramework.ipynb use automatic differentiation libraries?

No, OwnFramework.ipynb implements backpropagation manually using pure NumPy without automatic differentiation frameworks like PyTorch or TensorFlow. Each layer explicitly defines its forward and backward computations, making it an educational tool for understanding how gradients flow through a computational graph before relying on high-level deep learning libraries.

### How are parameters updated after the backward pass?

After the backward pass completes, each trainable layer calls its `update(lr)` method to perform stochastic gradient descent. The method updates parameters according to θ ← θ – η·∂ℒ/∂θ, where η is the learning rate passed as `lr` and ∂ℒ/∂θ represents the accumulated gradients stored in `dW` and `db` during backpropagation.