# How to Create Custom Neural Network Frameworks in AI for Beginners: A Step-by-Step Guide

> Learn to build custom neural network frameworks in AI. This beginner-friendly guide provides step-by-step instructions for implementing modular Python classes for layers, loss, and network containers.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: how-to-guide
- Published: 2026-08-29

---

**You can create a custom neural network framework by implementing modular Python classes for layers, loss functions, and network containers that handle forward propagation, backpropagation, and parameter updates using SGD.**

The Microsoft `AI-For-Beginners` repository provides a complete educational implementation in the *OwnFramework* lesson, located at `lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb`. This tutorial walks you through building a tiny deep-learning library from scratch to understand the mechanics behind PyTorch and TensorFlow.

## Understanding the Core Architecture

Every modern deep-learning library rests on three fundamental abstractions: **layers** that transform data, **loss functions** that measure error, and **optimizers** that update parameters. In the `OwnFramework` implementation, each component is a Python class with explicit `forward` and `backward` methods.

The architecture follows a clean separation of concerns. Layers store their parameters (`W`, `b`) and gradients (`dW`, `db`) internally. The `Net` container orchestrates the forward and backward passes by iterating through a stack of layers. This design mirrors production frameworks but in pure NumPy, making the mathematics transparent.

## Building the Layer Abstraction

### The Linear Layer

The `Linear` class implements a fully-connected transformation with Xavier initialization. According to the source code in `OwnFramework.ipynb`, it caches the input during the forward pass to compute gradients during backpropagation.

```python
class Linear:
    def __init__(self, nin, nout):
        self.W = np.random.normal(0, 1.0/np.sqrt(nin), (nout, nin))
        self.b = np.zeros((1, nout))
        self.dW = np.zeros_like(self.W)
        self.db = np.zeros_like(self.b)

    def forward(self, x):
        self.x = x                     # cache for backward

        return np.dot(x, self.W.T) + self.b

    def backward(self, dz):
        dx = np.dot(dz, self.W)        # ∂L/∂x

        self.dW = np.dot(dz.T, self.x)   # ∂L/∂W

        self.db = dz.sum(axis=0)       # ∂L/∂b

        return dx

    def update(self, lr):
        self.W -= lr * self.dW
        self.b -= lr * self.db

```

The `update` method applies vanilla SGD using the accumulated gradients stored in `self.dW` and `self.db`.

### The Softmax Layer

The `Softmax` layer converts logits into probability distributions. It stores the input `z` during the forward pass and recomputes the activation during the backward pass to calculate the Jacobian.

```python
class Softmax:
    def forward(self, z):
        self.z = z
        zmax = z.max(axis=1, keepdims=True)
        expz = np.exp(z - zmax)
        Z = expz.sum(axis=1, keepdims=True)
        return expz / Z

    def backward(self, dp):
        p = self.forward(self.z)               # recompute softmax

        pdp = p * dp
        return pdp - p * pdp.sum(axis=1, keepdims=True)

```

This implementation uses the log-sum-exp trick for numerical stability by subtracting `zmax` before exponentiation.

## Implementing Loss and Network Containers

### Cross-Entropy Loss Layer

The `CrossEntropyLoss` class separates the loss computation from the network layers. It calculates the negative log-likelihood and returns the gradient with respect to the softmax probabilities.

```python
class CrossEntropyLoss:
    def forward(self, p, y):
        self.p = p
        self.y = y
        p_of_y = p[np.arange(len(y)), y]
        return -np.log(p_of_y).mean()

    def backward(self, loss):
        dlog_softmax = np.zeros_like(self.p)
        dlog_softmax[np.arange(len(self.y)), self.y] = -1.0 / len(self.y)
        return dlog_softmax / self.p

```

The `backward` method returns the initial gradient that starts the backpropagation chain through the `Net` container.

### The Net Container

The `Net` class abstracts the plumbing of stacking layers. It implements the composite pattern, treating individual layers and the network as a unified interface.

```python
class Net:
    def __init__(self):
        self.layers = []

    def add(self, l):
        self.layers.append(l)

    def forward(self, x):
        for l in self.layers:
            x = l.forward(x)
        return x

    def backward(self, dz):
        for l in reversed(self.layers):
            dz = l.backward(dz)
        return dz

    def update(self, lr):
        for l in self.layers:
            if hasattr(l, "update"):
                l.update(lr)

```

The `backward` method iterates in reverse order, passing gradients from the loss layer back to the input layer.

## Training with Mini-Batch SGD

The training loop in `OwnFramework.ipynb` implements mini-batch stochastic gradient descent. The `train_epoch` function processes the dataset in batches, performing forward propagation, loss evaluation, backpropagation, and parameter updates.

```python
def train_epoch(net, X, y, loss, batch_size=4, lr=0.1):
    for i in range(0, len(X), batch_size):
        xb = X[i:i+batch_size]
        yb = y[i:i+batch_size]

        p = net.forward(xb)               # forward

        l = loss.forward(p, yb)            # loss

        dp = loss.backward(l)              # ∂L/∂p

        net.backward(dp)                   # back-prop

        net.update(lr)                     # SGD step

```

This loop demonstrates the four-step training paradigm used in all deep learning frameworks: forward, loss, backward, update.

## Extending to Real Datasets

To create custom neural network frameworks that handle real-world data, you can extend this architecture to the MNIST dataset using the lab file `lessons/3-NeuralNetworks/04-OwnFramework/lab/MyFW_MNIST.ipynb`.

```python

# Build a 2-layer classifier for MNIST (784 input pixels, 128 hidden units, 10 classes)

net = Net()
net.add(Linear(784, 128))   # Input → Hidden

net.add(Linear(128, 10))    # Hidden → Logits

net.add(Softmax())           # Logits → Probabilities

loss = CrossEntropyLoss()

# Train using the same loop structure

train_epoch(net, train_x, train_labels, loss, batch_size=64, lr=0.01)

```

The `Net` abstraction handles any layer depth automatically, allowing you to experiment with deeper architectures without modifying the training loop.

## Summary

- **Layer abstraction**: Implement `forward` for computation and `backward` for gradients, caching inputs as needed for derivative calculations.
- **Parameter management**: Store weights `W`, biases `b`, and their gradients `dW`, `db` inside layer classes, exposing an `update(lr)` method for SGD.
- **Network composition**: Use a `Net` container to stack layers and orchestrate forward/backward passes through sequential iteration.
- **Loss separation**: Create dedicated loss classes like `CrossEntropyLoss` that compute both the scalar loss value and the initial gradient for backpropagation.
- **Training loop**: Implement mini-batch SGD by iterating over `range(0, len(X), batch_size)`, calling `forward`, `loss.backward`, `net.backward`, and `net.update` in sequence.

## Frequently Asked Questions

### What is the difference between the Linear and Softmax layers in this custom framework?

The `Linear` layer contains learnable parameters (`W` and `b`) and implements an affine transformation, while the `Softmax` layer is a parameter-free activation function that normalizes logits into probabilities. Only layers with parameters require an `update` method; the `Softmax` layer simply transforms data during the forward pass and computes the Jacobian during the backward pass.

### How does the Net container handle backpropagation through multiple layers?

The `Net.backward` method iterates through `reversed(self.layers)`, passing the gradient from the current layer to the previous one. Each layer returns `dx` (the gradient with respect to its input), which becomes the `dz` input for the preceding layer. This chain rule application continues until the input layer is reached.

### Can I add dropout or convolutional layers to this framework?

Yes. You create custom neural network frameworks by following the same interface pattern: implement `forward` to compute the output and store any necessary context, implement `backward` to return the gradient with respect to the input, and include an `update` method if the layer has learnable parameters. Dropout would mask activations during `forward` and scale gradients during `backward`, while convolutional layers would implement im2col or direct convolution in their `forward` methods.

### Where can I find the complete working example of this framework?

The full implementation resides in `lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb` within the `microsoft/AI-For-Beginners` repository. This notebook includes the `Linear`, `Softmax`, and `CrossEntropyLoss` classes, the `Net` container, visualization utilities like `plot_decision_boundary`, and training loops for both synthetic 2D classification and MNIST digit recognition.