# Bigram vs MLP Language Models: From Counting to Neural Networks

> Explore the difference between Bigram and MLP language models. Understand how Bigram uses counts while MLP leverages neural networks and context for character prediction.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**Bigram language models predict the next character using only frequency counts of character pairs, while MLP language models use learnable neural network weights to process a context window of multiple preceding characters through embeddings and non-linear transformations.**

The `karpathy/nn-zero-to-hero` repository walks through this evolution in the *makemore* tutorial series, demonstrating how character-level language modeling progresses from simple statistical counting to sophisticated neural architectures. These implementations illustrate the fundamental trade-off between computational simplicity and predictive power.

## Context Window and Information Usage

### Bigram Models: Single Character Dependencies

A **Bigram model** restricts its prediction to only the immediately preceding character. For a given character sequence, the model looks up the probability of the next character based solely on the current one, ignoring any broader context from earlier in the sequence.

In `lectures/makemore/makemore_part1_bigrams.ipynb`, this is implemented as a simple counting loop:

```python
b = {}
for w in words:
    chs = ['<S>'] + list(w) + ['<E>']
    for ch1, ch2 in zip(chs, chs[1:]):
        bigram = (ch1, ch2)
        b[bigram] = b.get(bigram, 0) + 1

```

### MLP Models: Multi-Character Context Windows

The **MLP model** expands the receptive field to a configurable **context window** (here `block_size = 3`), allowing the network to learn patterns from multiple preceding characters. This architecture first embeds each context token into a continuous space, concatenates these embeddings, and processes them through hidden layers.

As implemented in `lectures/makemore/makemore_part2_mlp.ipynb`:

```python

# Build dataset with context window of 3

block_size = 3
X, Y = [], []
for w in words:
    context = [0] * block_size          # 0 = <S>

    for ch in w + '.':                  # '.' = <E>

        ix = stoi[ch]
        X.append(context)
        Y.append(ix)
        context = context[1:] + [ix]

X = torch.tensor(X)   # shape [N, 3]

Y = torch.tensor(Y)   # shape [N]

```

## Parameterization and Learnable Weights

### Statistical Tables vs. Neural Weights

The Bigram model stores probability distributions in a **count table** with dimensions `|V|²` (where `|V|` is vocabulary size). No gradient-based learning occurs; probabilities derive from normalized frequencies:

```python
prob = {pair: cnt / sum(cnt for (p, _), cnt in b.items() if p == pair[0])
        for pair, cnt in b.items()}

```

Conversely, the MLP introduces **learnable parameters** including:
- **Embedding matrix** `C` with shape `(vocab_size, embed_dim)`
- **First layer weights** `W1` with shape `(block_size*embed_dim, hidden_dim)`
- **Hidden bias** `b1` and **output bias** `b2`
- **Second layer weights** `W2` with shape `(hidden_dim, vocab_size)`

The forward pass in the MLP combines these matrices with non-linear activations:

```python
C = torch.randn((vocab_size, embed_dim))
W1 = torch.randn((block_size*embed_dim, hidden_dim))
b1 = torch.randn(hidden_dim)
W2 = torch.randn((hidden_dim, vocab_size))
b2 = torch.randn(vocab_size)

# Forward pass

emb = C[X]                                 # (N, 3, embed_dim)

h = torch.tanh(emb.view(-1, 6) @ W1 + b1)  # flatten and apply hidden layer

logits = h @ W2 + b2                       # (N, vocab_size)

```

## Training Objectives and Optimization

Both models optimize the same **cross-entropy loss** against the next-character target, but differ fundamentally in how they learn:

- **Bigram**: No optimization loop required. The model maximizes likelihood by counting observed frequencies, which represents the closed-form maximum likelihood solution for the bigram distribution.

- **MLP**: Requires **iterative gradient descent** via backpropagation. The loss flows backward through `W2`, `W1`, and `C`, updating continuous representations that capture semantic relationships between characters and positions.

```python
loss = F.cross_entropy(logits, Y)
loss.backward()

# optimizer step updates C, W1, W2, b1, b2

```

## Expressive Capacity and Computational Trade-offs

### Bigram Limitations and Efficiency

The Bigram model captures only **first-order dependencies** (immediate co-occurrence statistics). While limited in expressiveness, it offers:
- **Minimal memory**: Only `|V|²` counts to store
- **Instant training**: Single pass counting requires no GPU
- **Fast inference**: Simple table lookup

### MLP Advantages and Costs

The MLP can approximate **arbitrary functions** of its input context up to the capacity of its hidden layer. According to the `karpathy/nn-zero-to-hero` source code, this enables learning of phonetic patterns, syllable structures, and multi-character dependencies. However, this comes with:
- **Higher memory usage**: Scales with `embed_dim × vocab_size` plus hidden layer parameters
- **Computational requirements**: Matrix operations and gradient computation demand GPU/CPU acceleration
- **Training time**: Requires many epochs of gradient descent

## Summary

- **Bigram models** use simple frequency counting of character pairs in `makemore_part1_bigrams.ipynb`, offering fast training but limited to single-character context.
- **MLP models** in `makemore_part2_mlp.ipynb` employ learnable embeddings and feed-forward layers to process context windows of three or more characters.
- The transition from counting to backpropagation enables the model to learn continuous representations and non-linear interactions, dramatically improving predictive accuracy at the cost of computational resources.
- Bigram models serve as essential baselines and educational tools, while MLPs represent the foundation for modern neural language architectures.

## Frequently Asked Questions

### What is the fundamental architectural difference between Bigram and MLP language models?

The Bigram model is a **tabular probability estimator** that stores normalized counts of character pairs, containing no learnable parameters. The MLP is a **feed-forward neural network** with weight matrices (`C`, `W1`, `W2`) and bias vectors (`b1`, `b2`) that learns distributed representations through gradient descent.

### How does context window size affect model performance?

Bigram models are constrained to a context window of exactly one character, making them unable to capture word roots or syllable patterns. The MLP model uses `block_size = 3` (configurable) to attend to multiple preceding characters, enabling it to learn that "q" is almost always followed by "u" based on earlier context, even when processing character by character.

### Which model trains faster and requires less memory?

The **Bigram model** trains orders of magnitude faster because it requires only a single counting pass through the corpus and stores simply a `|V|²` probability table. The **MLP model** requires iterative optimization over many epochs, stores embedding matrices and hidden layers, and performs expensive matrix multiplications during both forward and backward passes.

### When should I choose a Bigram model over an MLP?

Use a **Bigram model** for rapid prototyping, educational demonstrations, sanity checks on data quality, or when working with extremely small corpora where neural networks would overfit. Choose an **MLP model** when you need to capture linguistic patterns beyond immediate character adjacency, or when building toward larger architectures like Transformers that require the foundation of learned embeddings and backpropagation.