How to Create Custom Neural Network Frameworks in AI for Beginners: A Step-by-Step Guide
You can create a custom neural network framework by implementing modular Python classes for layers, loss functions, and network containers that handle forward propagation, backpropagation, and parameter updates using SGD.
The Microsoft AI-For-Beginners repository provides a complete educational implementation in the OwnFramework lesson, located at lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb. This tutorial walks you through building a tiny deep-learning library from scratch to understand the mechanics behind PyTorch and TensorFlow.
Understanding the Core Architecture
Every modern deep-learning library rests on three fundamental abstractions: layers that transform data, loss functions that measure error, and optimizers that update parameters. In the OwnFramework implementation, each component is a Python class with explicit forward and backward methods.
The architecture follows a clean separation of concerns. Layers store their parameters (W, b) and gradients (dW, db) internally. The Net container orchestrates the forward and backward passes by iterating through a stack of layers. This design mirrors production frameworks but in pure NumPy, making the mathematics transparent.
Building the Layer Abstraction
The Linear Layer
The Linear class implements a fully-connected transformation with Xavier initialization. According to the source code in OwnFramework.ipynb, it caches the input during the forward pass to compute gradients during backpropagation.
class Linear:
def __init__(self, nin, nout):
self.W = np.random.normal(0, 1.0/np.sqrt(nin), (nout, nin))
self.b = np.zeros((1, nout))
self.dW = np.zeros_like(self.W)
self.db = np.zeros_like(self.b)
def forward(self, x):
self.x = x # cache for backward
return np.dot(x, self.W.T) + self.b
def backward(self, dz):
dx = np.dot(dz, self.W) # ∂L/∂x
self.dW = np.dot(dz.T, self.x) # ∂L/∂W
self.db = dz.sum(axis=0) # ∂L/∂b
return dx
def update(self, lr):
self.W -= lr * self.dW
self.b -= lr * self.db
The update method applies vanilla SGD using the accumulated gradients stored in self.dW and self.db.
The Softmax Layer
The Softmax layer converts logits into probability distributions. It stores the input z during the forward pass and recomputes the activation during the backward pass to calculate the Jacobian.
class Softmax:
def forward(self, z):
self.z = z
zmax = z.max(axis=1, keepdims=True)
expz = np.exp(z - zmax)
Z = expz.sum(axis=1, keepdims=True)
return expz / Z
def backward(self, dp):
p = self.forward(self.z) # recompute softmax
pdp = p * dp
return pdp - p * pdp.sum(axis=1, keepdims=True)
This implementation uses the log-sum-exp trick for numerical stability by subtracting zmax before exponentiation.
Implementing Loss and Network Containers
Cross-Entropy Loss Layer
The CrossEntropyLoss class separates the loss computation from the network layers. It calculates the negative log-likelihood and returns the gradient with respect to the softmax probabilities.
class CrossEntropyLoss:
def forward(self, p, y):
self.p = p
self.y = y
p_of_y = p[np.arange(len(y)), y]
return -np.log(p_of_y).mean()
def backward(self, loss):
dlog_softmax = np.zeros_like(self.p)
dlog_softmax[np.arange(len(self.y)), self.y] = -1.0 / len(self.y)
return dlog_softmax / self.p
The backward method returns the initial gradient that starts the backpropagation chain through the Net container.
The Net Container
The Net class abstracts the plumbing of stacking layers. It implements the composite pattern, treating individual layers and the network as a unified interface.
class Net:
def __init__(self):
self.layers = []
def add(self, l):
self.layers.append(l)
def forward(self, x):
for l in self.layers:
x = l.forward(x)
return x
def backward(self, dz):
for l in reversed(self.layers):
dz = l.backward(dz)
return dz
def update(self, lr):
for l in self.layers:
if hasattr(l, "update"):
l.update(lr)
The backward method iterates in reverse order, passing gradients from the loss layer back to the input layer.
Training with Mini-Batch SGD
The training loop in OwnFramework.ipynb implements mini-batch stochastic gradient descent. The train_epoch function processes the dataset in batches, performing forward propagation, loss evaluation, backpropagation, and parameter updates.
def train_epoch(net, X, y, loss, batch_size=4, lr=0.1):
for i in range(0, len(X), batch_size):
xb = X[i:i+batch_size]
yb = y[i:i+batch_size]
p = net.forward(xb) # forward
l = loss.forward(p, yb) # loss
dp = loss.backward(l) # ∂L/∂p
net.backward(dp) # back-prop
net.update(lr) # SGD step
This loop demonstrates the four-step training paradigm used in all deep learning frameworks: forward, loss, backward, update.
Extending to Real Datasets
To create custom neural network frameworks that handle real-world data, you can extend this architecture to the MNIST dataset using the lab file lessons/3-NeuralNetworks/04-OwnFramework/lab/MyFW_MNIST.ipynb.
# Build a 2-layer classifier for MNIST (784 input pixels, 128 hidden units, 10 classes)
net = Net()
net.add(Linear(784, 128)) # Input → Hidden
net.add(Linear(128, 10)) # Hidden → Logits
net.add(Softmax()) # Logits → Probabilities
loss = CrossEntropyLoss()
# Train using the same loop structure
train_epoch(net, train_x, train_labels, loss, batch_size=64, lr=0.01)
The Net abstraction handles any layer depth automatically, allowing you to experiment with deeper architectures without modifying the training loop.
Summary
- Layer abstraction: Implement
forwardfor computation andbackwardfor gradients, caching inputs as needed for derivative calculations. - Parameter management: Store weights
W, biasesb, and their gradientsdW,dbinside layer classes, exposing anupdate(lr)method for SGD. - Network composition: Use a
Netcontainer to stack layers and orchestrate forward/backward passes through sequential iteration. - Loss separation: Create dedicated loss classes like
CrossEntropyLossthat compute both the scalar loss value and the initial gradient for backpropagation. - Training loop: Implement mini-batch SGD by iterating over
range(0, len(X), batch_size), callingforward,loss.backward,net.backward, andnet.updatein sequence.
Frequently Asked Questions
What is the difference between the Linear and Softmax layers in this custom framework?
The Linear layer contains learnable parameters (W and b) and implements an affine transformation, while the Softmax layer is a parameter-free activation function that normalizes logits into probabilities. Only layers with parameters require an update method; the Softmax layer simply transforms data during the forward pass and computes the Jacobian during the backward pass.
How does the Net container handle backpropagation through multiple layers?
The Net.backward method iterates through reversed(self.layers), passing the gradient from the current layer to the previous one. Each layer returns dx (the gradient with respect to its input), which becomes the dz input for the preceding layer. This chain rule application continues until the input layer is reached.
Can I add dropout or convolutional layers to this framework?
Yes. You create custom neural network frameworks by following the same interface pattern: implement forward to compute the output and store any necessary context, implement backward to return the gradient with respect to the input, and include an update method if the layer has learnable parameters. Dropout would mask activations during forward and scale gradients during backward, while convolutional layers would implement im2col or direct convolution in their forward methods.
Where can I find the complete working example of this framework?
The full implementation resides in lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb within the microsoft/AI-For-Beginners repository. This notebook includes the Linear, Softmax, and CrossEntropyLoss classes, the Net container, visualization utilities like plot_decision_boundary, and training loops for both synthetic 2D classification and MNIST digit recognition.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →