How to Build a Custom Neural Network Framework from Scratch (Lesson 04)

To build a custom neural network framework from scratch, you implement modular layer classes with forward and backward methods, stack them in a Net container to form a computational graph, then apply SGD optimization to minimize loss.

The microsoft/AI-For-Beginners repository demonstrates this architecture in Lesson 04-OwnFramework, where a minimal yet fully functional deep learning library emerges in approximately 100 lines of pure Python. This walkthrough references the source code in OwnFramework.ipynb to explain how modern frameworks like PyTorch implement automatic differentiation and layer composition under the hood.

Core Architecture: Layers and the Computational Graph

The Layer Interface

Every component in the framework follows a strict contract defined by two essential methods. The forward method computes the layer's output tensor, while the backward method receives upstream gradients and computes gradients with respect to the layer's inputs and parameters. This design mirrors the autograd systems found in production frameworks.

In OwnFramework.ipynb, concrete implementations include Linear, Softmax, Tanh, and CrossEntropyLoss. Each class stores intermediate values during the forward pass (such as inputs x or activations out) to enable gradient computation during backpropagation.

The Net Container

The framework uses a Net class to manage the computational graph. This container maintains an ordered list of layer objects and orchestrates the flow of data:

class Net:
    def __init__(self):
        self.layers = []
    
    def add(self, layer):
        self.layers.append(layer)
    
    def forward(self, x):
        for layer in self.layers:
            x = layer.forward(x)
        return x
    
    def backward(self, loss_grad):
        for layer in reversed(self.layers):
            loss_grad = layer.backward(loss_grad)

By iterating forward through the list for predictions and backward for gradients, Net implements reverse-mode automatic differentiation without external dependencies.

Implementing Trainable Components

The Linear Layer

The Linear layer demonstrates parameter storage and gradient computation. It initializes weights W with small random values and biases b with zeros, then accumulates gradients dW and db during the backward pass:

class Linear:
    def __init__(self, in_dim, out_dim):
        self.W = np.random.randn(in_dim, out_dim) * 0.01
        self.b = np.zeros(out_dim)
    
    def forward(self, x):
        self.x = x                     # store for backward pass

        return np.dot(x, self.W) + self.b
    
    def backward(self, grad_output):
        # grad_output = dL/dz where z = Wx + b

        self.dW = np.dot(self.x.T, grad_output) / self.x.shape[0]
        self.db = grad_output.mean(axis=0)
        return np.dot(grad_output, self.W.T)   # propagate to previous layer

The backward method computes the Jacobian-vector product and returns the gradient with respect to inputs, enabling the chain rule to propagate errors through arbitrary depths.

Activation Functions

Non-linearities like Tanh and Softmax follow the same interface but contain no trainable parameters. The Softmax implementation in the lesson includes numerical stabilization and a simplified backward pass when paired with cross-entropy loss:

class Softmax:
    def forward(self, z):
        exp_z = np.exp(z - np.max(z, axis=1, keepdims=True))
        self.out = exp_z / exp_z.sum(axis=1, keepdims=True)
        return self.out
    
    def backward(self, grad_output):
        # grad_output is dL/dp where p = softmax(z)

        # Using Jacobian of softmax -> simplified form for cross-entropy

        return grad_output - self.out

Loss Computation

The CrossEntropyLoss layer bridges the network outputs and training labels. It computes the scalar loss and returns the initial gradient to seed the backward pass:

class CrossEntropyLoss:
    def forward(self, probs, targets):
        self.probs = probs
        self.targets = targets
        n = probs.shape[0]
        loss = -np.log(probs[range(n), targets]).sum() / n
        return loss
    
    def backward(self):
        n = self.probs.shape[0]
        grad = self.probs.copy()
        grad[range(n), self.targets] -= 1
        return grad / n

Training Mechanism

Forward and Backward Passes

Training begins by calling net.forward(train_x) to compute predictions, then criterion.forward(logits, train_labels) to calculate loss. The backward method on the loss layer produces the initial gradient, which the Net.backward method routes through each layer in reverse order to compute parameter gradients.

SGD Optimization

After gradients are computed, stochastic gradient descent updates the parameters in-place. The framework iterates through net.layers and applies the learning rate lr to weights and biases that expose the W and b attributes:

lr = 0.1
for layer in net.layers:
    if hasattr(layer, "W"):
        layer.W -= lr * layer.dW
        layer.b -= lr * layer.db

Mini-batch processing is supported naturally by passing sliced arrays (train_x[batch_start:batch_end]) to the forward method, following the same interface as full-batch training.

Complete Training Loop

The OwnFramework.ipynb notebook demonstrates convergence by combining these components into a training loop that visualizes decision boundaries in real-time:

net = Net()
net.add(Linear(2, 8))
net.add(Tanh())
net.add(Linear(8, 2))
net.add(Softmax())
criterion = CrossEntropyLoss()

lr = 0.1
for epoch in range(200):
    # forward pass

    logits = net.forward(train_x)
    loss = criterion.forward(logits, train_labels)
    
    # backward pass

    grad = criterion.backward()
    net.backward(grad)
    
    # SGD update

    for layer in net.layers:
        if hasattr(layer, "W"):
            layer.W -= lr * layer.dW
            layer.b -= lr * layer.db
    
    if epoch % 20 == 0:
        print(f"epoch {epoch}, loss {loss:.4f}")

Extending the Framework

Because every layer adheres to the same forward/backward interface, you can add new functionality without modifying the training loop. Implementing ReLU or Sigmoid requires only defining the two core methods and inserting the layer into the Net stack. This plug-and-play architecture enables rapid experimentation, as demonstrated in lab/MyFW_MNIST.ipynb where the framework trains on real image data.

Summary

  • Layer abstraction: Every component implements forward() for computation and backward() for gradient propagation.
  • Computational graph: The Net class in OwnFramework.ipynb sequences layers and automatically handles reverse-mode differentiation.
  • Parameter management: Trainable layers store W, b, dW, and db, initialized with small random weights and zero biases.
  • Training pipeline: Forward pass computes loss, backward pass accumulates gradients, and SGD updates parameters via simple attribute arithmetic.
  • Extensibility: New layers and loss functions integrate seamlessly by following the established interface contract.

Frequently Asked Questions

How does the Net class handle backpropagation automatically?

The Net.backward method iterates through self.layers in reverse order, passing the gradient from the current layer as input to the previous layer's backward method. This manual implementation of the chain rule mimics PyTorch's autograd engine but in pure Python, as shown in the OwnFramework.ipynb source code.

What files contain the complete framework implementation?

The core implementation resides in lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb, which contains the Linear, Softmax, Net, and loss class definitions. Extended examples for MNIST classification appear in lessons/3-NeuralNetworks/04-OwnFramework/lab/MyFW_MNIST.ipynb.

Can this framework train on mini-batches instead of full datasets?

Yes. The forward and backward methods operate on the first dimension of the input array, allowing you to pass slices like train_x[i:i+batch_size] without changing any layer code. The gradient calculations in Linear.backward explicitly divide by batch size (self.x.shape[0]) to average gradients correctly.

Why store intermediate values like self.x during the forward pass?

Layers cache inputs and outputs (e.g., self.x in Linear, self.out in Softmax) because computing gradients during backward requires knowing the forward-pass values. For example, dW depends on the input x, which is no longer available unless stored during the forward computation. This trade-off of memory for computation is standard in deep learning frameworks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →