Bigram vs MLP Language Models: From Counting to Neural Networks

Bigram language models predict the next character using only frequency counts of character pairs, while MLP language models use learnable neural network weights to process a context window of multiple preceding characters through embeddings and non-linear transformations.

The karpathy/nn-zero-to-hero repository walks through this evolution in the makemore tutorial series, demonstrating how character-level language modeling progresses from simple statistical counting to sophisticated neural architectures. These implementations illustrate the fundamental trade-off between computational simplicity and predictive power.

Context Window and Information Usage

Bigram Models: Single Character Dependencies

A Bigram model restricts its prediction to only the immediately preceding character. For a given character sequence, the model looks up the probability of the next character based solely on the current one, ignoring any broader context from earlier in the sequence.

In lectures/makemore/makemore_part1_bigrams.ipynb, this is implemented as a simple counting loop:

b = {}
for w in words:
    chs = ['<S>'] + list(w) + ['<E>']
    for ch1, ch2 in zip(chs, chs[1:]):
        bigram = (ch1, ch2)
        b[bigram] = b.get(bigram, 0) + 1

MLP Models: Multi-Character Context Windows

The MLP model expands the receptive field to a configurable context window (here block_size = 3), allowing the network to learn patterns from multiple preceding characters. This architecture first embeds each context token into a continuous space, concatenates these embeddings, and processes them through hidden layers.

As implemented in lectures/makemore/makemore_part2_mlp.ipynb:


# Build dataset with context window of 3

block_size = 3
X, Y = [], []
for w in words:
    context = [0] * block_size          # 0 = <S>

    for ch in w + '.':                  # '.' = <E>

        ix = stoi[ch]
        X.append(context)
        Y.append(ix)
        context = context[1:] + [ix]

X = torch.tensor(X)   # shape [N, 3]

Y = torch.tensor(Y)   # shape [N]

Parameterization and Learnable Weights

Statistical Tables vs. Neural Weights

The Bigram model stores probability distributions in a count table with dimensions |V|² (where |V| is vocabulary size). No gradient-based learning occurs; probabilities derive from normalized frequencies:

prob = {pair: cnt / sum(cnt for (p, _), cnt in b.items() if p == pair[0])
        for pair, cnt in b.items()}

Conversely, the MLP introduces learnable parameters including:

  • Embedding matrix C with shape (vocab_size, embed_dim)
  • First layer weights W1 with shape (block_size*embed_dim, hidden_dim)
  • Hidden bias b1 and output bias b2
  • Second layer weights W2 with shape (hidden_dim, vocab_size)

The forward pass in the MLP combines these matrices with non-linear activations:

C = torch.randn((vocab_size, embed_dim))
W1 = torch.randn((block_size*embed_dim, hidden_dim))
b1 = torch.randn(hidden_dim)
W2 = torch.randn((hidden_dim, vocab_size))
b2 = torch.randn(vocab_size)

# Forward pass

emb = C[X]                                 # (N, 3, embed_dim)

h = torch.tanh(emb.view(-1, 6) @ W1 + b1)  # flatten and apply hidden layer

logits = h @ W2 + b2                       # (N, vocab_size)

Training Objectives and Optimization

Both models optimize the same cross-entropy loss against the next-character target, but differ fundamentally in how they learn:

  • Bigram: No optimization loop required. The model maximizes likelihood by counting observed frequencies, which represents the closed-form maximum likelihood solution for the bigram distribution.

  • MLP: Requires iterative gradient descent via backpropagation. The loss flows backward through W2, W1, and C, updating continuous representations that capture semantic relationships between characters and positions.

loss = F.cross_entropy(logits, Y)
loss.backward()

# optimizer step updates C, W1, W2, b1, b2

Expressive Capacity and Computational Trade-offs

Bigram Limitations and Efficiency

The Bigram model captures only first-order dependencies (immediate co-occurrence statistics). While limited in expressiveness, it offers:

  • Minimal memory: Only |V|² counts to store
  • Instant training: Single pass counting requires no GPU
  • Fast inference: Simple table lookup

MLP Advantages and Costs

The MLP can approximate arbitrary functions of its input context up to the capacity of its hidden layer. According to the karpathy/nn-zero-to-hero source code, this enables learning of phonetic patterns, syllable structures, and multi-character dependencies. However, this comes with:

  • Higher memory usage: Scales with embed_dim × vocab_size plus hidden layer parameters
  • Computational requirements: Matrix operations and gradient computation demand GPU/CPU acceleration
  • Training time: Requires many epochs of gradient descent

Summary

  • Bigram models use simple frequency counting of character pairs in makemore_part1_bigrams.ipynb, offering fast training but limited to single-character context.
  • MLP models in makemore_part2_mlp.ipynb employ learnable embeddings and feed-forward layers to process context windows of three or more characters.
  • The transition from counting to backpropagation enables the model to learn continuous representations and non-linear interactions, dramatically improving predictive accuracy at the cost of computational resources.
  • Bigram models serve as essential baselines and educational tools, while MLPs represent the foundation for modern neural language architectures.

Frequently Asked Questions

What is the fundamental architectural difference between Bigram and MLP language models?

The Bigram model is a tabular probability estimator that stores normalized counts of character pairs, containing no learnable parameters. The MLP is a feed-forward neural network with weight matrices (C, W1, W2) and bias vectors (b1, b2) that learns distributed representations through gradient descent.

How does context window size affect model performance?

Bigram models are constrained to a context window of exactly one character, making them unable to capture word roots or syllable patterns. The MLP model uses block_size = 3 (configurable) to attend to multiple preceding characters, enabling it to learn that "q" is almost always followed by "u" based on earlier context, even when processing character by character.

Which model trains faster and requires less memory?

The Bigram model trains orders of magnitude faster because it requires only a single counting pass through the corpus and stores simply a |V|² probability table. The MLP model requires iterative optimization over many epochs, stores embedding matrices and hidden layers, and performs expensive matrix multiplications during both forward and backward passes.

When should I choose a Bigram model over an MLP?

Use a Bigram model for rapid prototyping, educational demonstrations, sanity checks on data quality, or when working with extremely small corpora where neural networks would overfit. Choose an MLP model when you need to capture linguistic patterns beyond immediate character adjacency, or when building toward larger architectures like Transformers that require the foundation of learned embeddings and backpropagation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →