# How to Build a Single-Layer Neural Network Using the Perceptron in AI-For-Beginners

> Learn to build a single-layer neural network with the Perceptron. This AI for Beginners guide explains the linear classifier and perceptron learning rule for binary classification.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: how-to-guide
- Published: 2026-08-29

---

**The Perceptron is a linear classifier that learns a weight vector to separate binary classes via the perceptron learning rule, implemented in the microsoft/AI-For-Beginners repository as a foundational single-layer neural network.**

The Perceptron represents the simplest form of a neural network—a single-layer linear classifier that maps input vectors to binary outputs. In the microsoft/AI-For-Beginners curriculum, this algorithm serves as the entry point for understanding how to build a single-layer neural network using the Perceptron before advancing to multi-layer architectures. The implementation in `lessons/3-NeuralNetworks/03-Perceptron/Perceptron.ipynb` demonstrates core concepts including weight updates, bias absorption, and decision boundaries through hands-on Python examples.

## Understanding the Perceptron Algorithm

The Perceptron implements a linear decision function that classifies inputs based on a learned weight vector. According to the source code in `Perceptron.ipynb`, the model maps an input vector **x** to an output **y** using a weight vector **w** and an optional bias term.

### The Decision Rule

The classification follows a step function applied to the weighted sum of inputs. The model computes **z = w^T x** (the dot product of weights and inputs), then applies the activation:

- Output **+1** if **z ≥ 0**
- Output **-1** if **z < 0**

This binary threshold function makes the Perceptron a linear classifier that separates data into two regions using a hyperplane defined by **w**.

### The Learning Rule

Training minimizes classification error through iterative weight adjustments. When the model misclassifies an example **x** with label **t** (where **t** ∈ {+1, -1}), the weights update according to:

**w_new = w_old + η * x * t**

Here, **η** represents the learning rate, **x** is the misclassified input vector, and **t** is the true label. Positive examples that score negative trigger weight increases, while negative examples that score positive trigger weight decreases.

## Implementing the Perceptron in Python

The reference implementation provides a complete training pipeline that handles data preparation, weight initialization, and stochastic updates.

### Absorbing the Bias Term

To simplify the model architecture, the implementation absorbs the bias by appending a constant feature `1` to every input vector. This allows the weight vector **w** to include the bias component naturally, eliminating the need for separate bias variable tracking. For a 2D input **[x₁, x₂]**, the augmented vector becomes **[x₁, x₂, 1]`.

### The Training Function

The `train` function implements stochastic weight updates by repeatedly sampling one positive and one negative example per iteration. Located in the lesson notebook, this function initializes weights to zero, then iteratively corrects misclassifications based on the sign of the dot product.

## Complete Implementation: Binary Classification on Synthetic Data

This example demonstrates how to build a single-layer neural network using the Perceptron on a 2D toy dataset from scikit-learn. The code follows the exact implementation shown in the repository's lesson notebook.

```python
import numpy as np
from sklearn.datasets import make_classification
import random

# Create a small synthetic dataset

X, Y = make_classification(n_samples=50, n_features=2,
                           n_redundant=0, n_informative=2, flip_y=0)
Y = Y*2 - 1                     # Convert 0/1 to -1/+1

X = X.astype(np.float32)
Y = Y.astype(np.int32)

# Split into train / test

train_x, test_x = np.split(X, [40])
train_y, test_y = np.split(Y, [40])

# Add bias term (extra column of 1s)

train_pos = np.array([np.append(x, 1) for x, t in zip(train_x, train_y) if t > 0])
train_neg = np.array([np.append(x, 1) for x, t in zip(train_x, train_y) if t < 0])

# Perceptron training (same code as in the notebook)

def train(positive_examples, negative_examples,
          num_iterations=100, learning_rate=0.01):
    num_dims = positive_examples.shape[1]
    w = np.zeros((num_dims, 1))          # initialise weights

    for _ in range(num_iterations):
        pos = random.choice(positive_examples)
        neg = random.choice(negative_examples)

        if np.dot(pos, w) < 0:            # mis‑classified positive

            w += learning_rate * pos.reshape(w.shape)
        if np.dot(neg, w) >= 0:           # mis‑classified negative

            w -= learning_rate * neg.reshape(w.shape)
    return w

weights = train(train_pos, train_neg)

# Simple accuracy on the test set

def accuracy(w, X, Y):
    Xb = np.c_[X, np.ones(len(X))]
    preds = np.sign(np.dot(Xb, w)).flatten()
    return np.mean(preds == Y)

print("Test accuracy:", accuracy(weights, test_x, test_y))

```

## Scaling to Image Data: MNIST Digit Classification

The same `train` function generalizes to high-dimensional data without architectural changes. The repository includes `data/mnist.pkl.gz`, which provides flattened 28×28 grayscale images. The following implementation extracts two digits (1 and 0) and trains the perceptron to distinguish between them:

```python
import gzip, pickle
import numpy as np

# Load the MNIST pickle supplied with the repo

with gzip.open('data/mnist.pkl.gz', 'rb') as f:
    mnist = pickle.load(f, encoding='latin1')

train_imgs = mnist['Train']['Features'].astype(np.float32) / 255.0
train_lbls = mnist['Train']['Labels']

# Helper to pull out two digits as positive / negative samples

def get_mnist_pos_neg(pos_digit, neg_digit):
    pos_idx = np.where(train_lbls == pos_digit)[0]
    neg_idx = np.where(train_lbls == neg_digit)[0]
    pos = train_imgs[pos_idx]
    neg = train_imgs[neg_idx]
    # flatten images and add bias term

    pos = np.c_[pos.reshape(pos.shape[0], -1), np.ones(pos.shape[0])]
    neg = np.c_[neg.reshape(neg.shape[0], -1), np.ones(neg.shape[0])]
    return pos, neg

pos, neg = get_mnist_pos_neg(1, 0)
w = train(pos, neg, num_iterations=2000, learning_rate=0.001)

# Evaluate on a held‑out subset

def mnist_accuracy(w, imgs, labels, pos_digit, neg_digit):
    Xb = np.c_[imgs.reshape(imgs.shape[0], -1), np.ones(imgs.shape[0])]
    pred = np.sign(np.dot(Xb, w)).flatten()
    true = np.where(labels == pos_digit, 1, -1)
    return np.mean(pred == true)

print("MNIST 1 vs 0 accuracy:",
      mnist_accuracy(w, train_imgs[2000:3000], train_lbls[2000:3000], 1, 0))

```

## Evaluating Model Performance

After training, the implementation measures accuracy by comparing predicted signs against ground-truth labels. The `accuracy` function augments test inputs with the bias column, computes the dot product with learned weights, and applies the sign function to generate predictions. For the MNIST example, the `mnist_accuracy` function reshapes image tensors and converts digit labels into the +1/-1 format expected by the binary classifier.

## Summary

- The Perceptron in `microsoft/AI-For-Beginners` implements a single-layer linear classifier using the perceptron learning rule to iteratively update weights on misclassified examples.
- The algorithm absorbs bias terms by appending a constant `1` to input vectors, simplifying the weight vector structure while maintaining mathematical equivalence to explicit bias notation.
- The `train` function uses stochastic sampling of positive and negative examples to converge toward a separating hyperplane that minimizes classification error.
- The implementation scales from 2D synthetic data to high-dimensional MNIST digit classification (784 features) without requiring architectural modifications.

## Frequently Asked Questions

### What is the primary purpose of the Perceptron lesson in AI-For-Beginners?

The lesson serves as the foundational building block for neural network education. It teaches linear classification, the perceptron learning rule, and weight update mechanics before students advance to multi-layer architectures. The `Perceptron.ipynb` notebook provides interactive visualizations of decision boundaries to reinforce these concepts.

### Why does the implementation append a constant 1 to input vectors instead of using a separate bias term?

This technique absorbs the bias into the weight vector to simplify the model implementation. By treating the bias as an additional weight connected to a constant input feature, the mathematics remains equivalent while eliminating the need to track separate bias variables during the weight update loop. This augmentation appears in both the synthetic data and MNIST examples.

### How does the perceptron learning rule handle misclassified examples during training?

The `train` function checks the sign of the dot product for sampled positive and negative examples. For misclassified positives (where **w^T x < 0**), it adds **η * x** to the weights. For misclassified negatives (where **w^T x ≥ 0**), it subtracts **η * x**. This iterative correction pushes the decision boundary toward correctly separating the two classes.

### Can the single-layer Perceptron solve multi-class classification problems like full MNIST?

No, the single-layer Perceptron demonstrated in the repository is strictly a binary classifier designed to distinguish between two classes (e.g., digits 1 vs 0). Solving multi-class problems like full 10-digit MNIST classification requires multiple perceptrons (one per class) in a one-vs-rest configuration or advancing to multi-layer neural networks covered in subsequent lessons of the curriculum.