How to Create Custom Neural Network Frameworks in AI for Beginners: A Step-by-Step Guide

You can create a custom neural network framework by implementing modular Python classes for layers, loss functions, and network containers that handle forward propagation, backpropagation, and parameter updates using SGD.

The Microsoft AI-For-Beginners repository provides a complete educational implementation in the OwnFramework lesson, located at lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb. This tutorial walks you through building a tiny deep-learning library from scratch to understand the mechanics behind PyTorch and TensorFlow.

Understanding the Core Architecture

Every modern deep-learning library rests on three fundamental abstractions: layers that transform data, loss functions that measure error, and optimizers that update parameters. In the OwnFramework implementation, each component is a Python class with explicit forward and backward methods.

The architecture follows a clean separation of concerns. Layers store their parameters (W, b) and gradients (dW, db) internally. The Net container orchestrates the forward and backward passes by iterating through a stack of layers. This design mirrors production frameworks but in pure NumPy, making the mathematics transparent.

Building the Layer Abstraction

The Linear Layer

The Linear class implements a fully-connected transformation with Xavier initialization. According to the source code in OwnFramework.ipynb, it caches the input during the forward pass to compute gradients during backpropagation.

class Linear:
    def __init__(self, nin, nout):
        self.W = np.random.normal(0, 1.0/np.sqrt(nin), (nout, nin))
        self.b = np.zeros((1, nout))
        self.dW = np.zeros_like(self.W)
        self.db = np.zeros_like(self.b)

    def forward(self, x):
        self.x = x                     # cache for backward

        return np.dot(x, self.W.T) + self.b

    def backward(self, dz):
        dx = np.dot(dz, self.W)        # ∂L/∂x

        self.dW = np.dot(dz.T, self.x)   # ∂L/∂W

        self.db = dz.sum(axis=0)       # ∂L/∂b

        return dx

    def update(self, lr):
        self.W -= lr * self.dW
        self.b -= lr * self.db

The update method applies vanilla SGD using the accumulated gradients stored in self.dW and self.db.

The Softmax Layer

The Softmax layer converts logits into probability distributions. It stores the input z during the forward pass and recomputes the activation during the backward pass to calculate the Jacobian.

class Softmax:
    def forward(self, z):
        self.z = z
        zmax = z.max(axis=1, keepdims=True)
        expz = np.exp(z - zmax)
        Z = expz.sum(axis=1, keepdims=True)
        return expz / Z

    def backward(self, dp):
        p = self.forward(self.z)               # recompute softmax

        pdp = p * dp
        return pdp - p * pdp.sum(axis=1, keepdims=True)

This implementation uses the log-sum-exp trick for numerical stability by subtracting zmax before exponentiation.

Implementing Loss and Network Containers

Cross-Entropy Loss Layer

The CrossEntropyLoss class separates the loss computation from the network layers. It calculates the negative log-likelihood and returns the gradient with respect to the softmax probabilities.

class CrossEntropyLoss:
    def forward(self, p, y):
        self.p = p
        self.y = y
        p_of_y = p[np.arange(len(y)), y]
        return -np.log(p_of_y).mean()

    def backward(self, loss):
        dlog_softmax = np.zeros_like(self.p)
        dlog_softmax[np.arange(len(self.y)), self.y] = -1.0 / len(self.y)
        return dlog_softmax / self.p

The backward method returns the initial gradient that starts the backpropagation chain through the Net container.

The Net Container

The Net class abstracts the plumbing of stacking layers. It implements the composite pattern, treating individual layers and the network as a unified interface.

class Net:
    def __init__(self):
        self.layers = []

    def add(self, l):
        self.layers.append(l)

    def forward(self, x):
        for l in self.layers:
            x = l.forward(x)
        return x

    def backward(self, dz):
        for l in reversed(self.layers):
            dz = l.backward(dz)
        return dz

    def update(self, lr):
        for l in self.layers:
            if hasattr(l, "update"):
                l.update(lr)

The backward method iterates in reverse order, passing gradients from the loss layer back to the input layer.

Training with Mini-Batch SGD

The training loop in OwnFramework.ipynb implements mini-batch stochastic gradient descent. The train_epoch function processes the dataset in batches, performing forward propagation, loss evaluation, backpropagation, and parameter updates.

def train_epoch(net, X, y, loss, batch_size=4, lr=0.1):
    for i in range(0, len(X), batch_size):
        xb = X[i:i+batch_size]
        yb = y[i:i+batch_size]

        p = net.forward(xb)               # forward

        l = loss.forward(p, yb)            # loss

        dp = loss.backward(l)              # ∂L/∂p

        net.backward(dp)                   # back-prop

        net.update(lr)                     # SGD step

This loop demonstrates the four-step training paradigm used in all deep learning frameworks: forward, loss, backward, update.

Extending to Real Datasets

To create custom neural network frameworks that handle real-world data, you can extend this architecture to the MNIST dataset using the lab file lessons/3-NeuralNetworks/04-OwnFramework/lab/MyFW_MNIST.ipynb.


# Build a 2-layer classifier for MNIST (784 input pixels, 128 hidden units, 10 classes)

net = Net()
net.add(Linear(784, 128))   # Input → Hidden

net.add(Linear(128, 10))    # Hidden → Logits

net.add(Softmax())           # Logits → Probabilities

loss = CrossEntropyLoss()

# Train using the same loop structure

train_epoch(net, train_x, train_labels, loss, batch_size=64, lr=0.01)

The Net abstraction handles any layer depth automatically, allowing you to experiment with deeper architectures without modifying the training loop.

Summary

  • Layer abstraction: Implement forward for computation and backward for gradients, caching inputs as needed for derivative calculations.
  • Parameter management: Store weights W, biases b, and their gradients dW, db inside layer classes, exposing an update(lr) method for SGD.
  • Network composition: Use a Net container to stack layers and orchestrate forward/backward passes through sequential iteration.
  • Loss separation: Create dedicated loss classes like CrossEntropyLoss that compute both the scalar loss value and the initial gradient for backpropagation.
  • Training loop: Implement mini-batch SGD by iterating over range(0, len(X), batch_size), calling forward, loss.backward, net.backward, and net.update in sequence.

Frequently Asked Questions

What is the difference between the Linear and Softmax layers in this custom framework?

The Linear layer contains learnable parameters (W and b) and implements an affine transformation, while the Softmax layer is a parameter-free activation function that normalizes logits into probabilities. Only layers with parameters require an update method; the Softmax layer simply transforms data during the forward pass and computes the Jacobian during the backward pass.

How does the Net container handle backpropagation through multiple layers?

The Net.backward method iterates through reversed(self.layers), passing the gradient from the current layer to the previous one. Each layer returns dx (the gradient with respect to its input), which becomes the dz input for the preceding layer. This chain rule application continues until the input layer is reached.

Can I add dropout or convolutional layers to this framework?

Yes. You create custom neural network frameworks by following the same interface pattern: implement forward to compute the output and store any necessary context, implement backward to return the gradient with respect to the input, and include an update method if the layer has learnable parameters. Dropout would mask activations during forward and scale gradients during backward, while convolutional layers would implement im2col or direct convolution in their forward methods.

Where can I find the complete working example of this framework?

The full implementation resides in lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb within the microsoft/AI-For-Beginners repository. This notebook includes the Linear, Softmax, and CrossEntropyLoss classes, the Net container, visualization utilities like plot_decision_boundary, and training loops for both synthetic 2D classification and MNIST digit recognition.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →