# How to Build a Custom Neural Network Framework from Scratch (Lesson 04)

> Learn to build a custom neural network framework from scratch. Implement modular layers, form a computational graph, and apply SGD optimization in this AI-For-Beginners lesson.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: tutorial
- Published: 2026-08-23

---

**To build a custom neural network framework from scratch, you implement modular layer classes with `forward` and `backward` methods, stack them in a `Net` container to form a computational graph, then apply SGD optimization to minimize loss.**

The microsoft/AI-For-Beginners repository demonstrates this architecture in **Lesson 04-OwnFramework**, where a minimal yet fully functional deep learning library emerges in approximately 100 lines of pure Python. This walkthrough references the source code in `OwnFramework.ipynb` to explain how modern frameworks like PyTorch implement automatic differentiation and layer composition under the hood.

## Core Architecture: Layers and the Computational Graph

### The Layer Interface

Every component in the framework follows a strict contract defined by two essential methods. The `forward` method computes the layer's output tensor, while the `backward` method receives upstream gradients and computes gradients with respect to the layer's inputs and parameters. This design mirrors the **autograd** systems found in production frameworks.

In `OwnFramework.ipynb`, concrete implementations include `Linear`, `Softmax`, `Tanh`, and `CrossEntropyLoss`. Each class stores intermediate values during the forward pass (such as inputs `x` or activations `out`) to enable gradient computation during backpropagation.

### The Net Container

The framework uses a `Net` class to manage the **computational graph**. This container maintains an ordered list of layer objects and orchestrates the flow of data:

```python
class Net:
    def __init__(self):
        self.layers = []
    
    def add(self, layer):
        self.layers.append(layer)
    
    def forward(self, x):
        for layer in self.layers:
            x = layer.forward(x)
        return x
    
    def backward(self, loss_grad):
        for layer in reversed(self.layers):
            loss_grad = layer.backward(loss_grad)

```

By iterating forward through the list for predictions and backward for gradients, `Net` implements reverse-mode automatic differentiation without external dependencies.

## Implementing Trainable Components

### The Linear Layer

The `Linear` layer demonstrates **parameter storage** and gradient computation. It initializes weights `W` with small random values and biases `b` with zeros, then accumulates gradients `dW` and `db` during the backward pass:

```python
class Linear:
    def __init__(self, in_dim, out_dim):
        self.W = np.random.randn(in_dim, out_dim) * 0.01
        self.b = np.zeros(out_dim)
    
    def forward(self, x):
        self.x = x                     # store for backward pass

        return np.dot(x, self.W) + self.b
    
    def backward(self, grad_output):
        # grad_output = dL/dz where z = Wx + b

        self.dW = np.dot(self.x.T, grad_output) / self.x.shape[0]
        self.db = grad_output.mean(axis=0)
        return np.dot(grad_output, self.W.T)   # propagate to previous layer

```

The backward method computes the Jacobian-vector product and returns the gradient with respect to inputs, enabling the chain rule to propagate errors through arbitrary depths.

### Activation Functions

Non-linearities like `Tanh` and `Softmax` follow the same interface but contain no trainable parameters. The `Softmax` implementation in the lesson includes numerical stabilization and a simplified backward pass when paired with cross-entropy loss:

```python
class Softmax:
    def forward(self, z):
        exp_z = np.exp(z - np.max(z, axis=1, keepdims=True))
        self.out = exp_z / exp_z.sum(axis=1, keepdims=True)
        return self.out
    
    def backward(self, grad_output):
        # grad_output is dL/dp where p = softmax(z)

        # Using Jacobian of softmax -> simplified form for cross-entropy

        return grad_output - self.out

```

### Loss Computation

The `CrossEntropyLoss` layer bridges the network outputs and training labels. It computes the scalar loss and returns the initial gradient to seed the backward pass:

```python
class CrossEntropyLoss:
    def forward(self, probs, targets):
        self.probs = probs
        self.targets = targets
        n = probs.shape[0]
        loss = -np.log(probs[range(n), targets]).sum() / n
        return loss
    
    def backward(self):
        n = self.probs.shape[0]
        grad = self.probs.copy()
        grad[range(n), self.targets] -= 1
        return grad / n

```

## Training Mechanism

### Forward and Backward Passes

Training begins by calling `net.forward(train_x)` to compute predictions, then `criterion.forward(logits, train_labels)` to calculate loss. The `backward` method on the loss layer produces the initial gradient, which the `Net.backward` method routes through each layer in reverse order to compute parameter gradients.

### SGD Optimization

After gradients are computed, **stochastic gradient descent** updates the parameters in-place. The framework iterates through `net.layers` and applies the learning rate `lr` to weights and biases that expose the `W` and `b` attributes:

```python
lr = 0.1
for layer in net.layers:
    if hasattr(layer, "W"):
        layer.W -= lr * layer.dW
        layer.b -= lr * layer.db

```

Mini-batch processing is supported naturally by passing sliced arrays (`train_x[batch_start:batch_end]`) to the forward method, following the same interface as full-batch training.

### Complete Training Loop

The `OwnFramework.ipynb` notebook demonstrates convergence by combining these components into a training loop that visualizes decision boundaries in real-time:

```python
net = Net()
net.add(Linear(2, 8))
net.add(Tanh())
net.add(Linear(8, 2))
net.add(Softmax())
criterion = CrossEntropyLoss()

lr = 0.1
for epoch in range(200):
    # forward pass

    logits = net.forward(train_x)
    loss = criterion.forward(logits, train_labels)
    
    # backward pass

    grad = criterion.backward()
    net.backward(grad)
    
    # SGD update

    for layer in net.layers:
        if hasattr(layer, "W"):
            layer.W -= lr * layer.dW
            layer.b -= lr * layer.db
    
    if epoch % 20 == 0:
        print(f"epoch {epoch}, loss {loss:.4f}")

```

## Extending the Framework

Because every layer adheres to the same `forward`/`backward` interface, you can add new functionality without modifying the training loop. Implementing `ReLU` or `Sigmoid` requires only defining the two core methods and inserting the layer into the `Net` stack. This **plug-and-play architecture** enables rapid experimentation, as demonstrated in `lab/MyFW_MNIST.ipynb` where the framework trains on real image data.

## Summary

- **Layer abstraction**: Every component implements `forward()` for computation and `backward()` for gradient propagation.
- **Computational graph**: The `Net` class in `OwnFramework.ipynb` sequences layers and automatically handles reverse-mode differentiation.
- **Parameter management**: Trainable layers store `W`, `b`, `dW`, and `db`, initialized with small random weights and zero biases.
- **Training pipeline**: Forward pass computes loss, backward pass accumulates gradients, and SGD updates parameters via simple attribute arithmetic.
- **Extensibility**: New layers and loss functions integrate seamlessly by following the established interface contract.

## Frequently Asked Questions

### How does the Net class handle backpropagation automatically?

The `Net.backward` method iterates through `self.layers` in reverse order, passing the gradient from the current layer as input to the previous layer's `backward` method. This manual implementation of the chain rule mimics PyTorch's autograd engine but in pure Python, as shown in the `OwnFramework.ipynb` source code.

### What files contain the complete framework implementation?

The core implementation resides in `lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb`, which contains the `Linear`, `Softmax`, `Net`, and loss class definitions. Extended examples for MNIST classification appear in `lessons/3-NeuralNetworks/04-OwnFramework/lab/MyFW_MNIST.ipynb`.

### Can this framework train on mini-batches instead of full datasets?

Yes. The `forward` and `backward` methods operate on the first dimension of the input array, allowing you to pass slices like `train_x[i:i+batch_size]` without changing any layer code. The gradient calculations in `Linear.backward` explicitly divide by batch size (`self.x.shape[0]`) to average gradients correctly.

### Why store intermediate values like `self.x` during the forward pass?

Layers cache inputs and outputs (e.g., `self.x` in `Linear`, `self.out` in `Softmax`) because computing gradients during `backward` requires knowing the forward-pass values. For example, `dW` depends on the input `x`, which is no longer available unless stored during the forward computation. This trade-off of memory for computation is standard in deep learning frameworks.