Bigram vs MLP Language Models: From Counting to Neural Networks
Bigram language models predict the next character using only frequency counts of character pairs, while MLP language models use learnable neural network weights to process a context window of multiple preceding characters through embeddings and non-linear transformations.
The karpathy/nn-zero-to-hero repository walks through this evolution in the makemore tutorial series, demonstrating how character-level language modeling progresses from simple statistical counting to sophisticated neural architectures. These implementations illustrate the fundamental trade-off between computational simplicity and predictive power.
Context Window and Information Usage
Bigram Models: Single Character Dependencies
A Bigram model restricts its prediction to only the immediately preceding character. For a given character sequence, the model looks up the probability of the next character based solely on the current one, ignoring any broader context from earlier in the sequence.
In lectures/makemore/makemore_part1_bigrams.ipynb, this is implemented as a simple counting loop:
b = {}
for w in words:
chs = ['<S>'] + list(w) + ['<E>']
for ch1, ch2 in zip(chs, chs[1:]):
bigram = (ch1, ch2)
b[bigram] = b.get(bigram, 0) + 1
MLP Models: Multi-Character Context Windows
The MLP model expands the receptive field to a configurable context window (here block_size = 3), allowing the network to learn patterns from multiple preceding characters. This architecture first embeds each context token into a continuous space, concatenates these embeddings, and processes them through hidden layers.
As implemented in lectures/makemore/makemore_part2_mlp.ipynb:
# Build dataset with context window of 3
block_size = 3
X, Y = [], []
for w in words:
context = [0] * block_size # 0 = <S>
for ch in w + '.': # '.' = <E>
ix = stoi[ch]
X.append(context)
Y.append(ix)
context = context[1:] + [ix]
X = torch.tensor(X) # shape [N, 3]
Y = torch.tensor(Y) # shape [N]
Parameterization and Learnable Weights
Statistical Tables vs. Neural Weights
The Bigram model stores probability distributions in a count table with dimensions |V|² (where |V| is vocabulary size). No gradient-based learning occurs; probabilities derive from normalized frequencies:
prob = {pair: cnt / sum(cnt for (p, _), cnt in b.items() if p == pair[0])
for pair, cnt in b.items()}
Conversely, the MLP introduces learnable parameters including:
- Embedding matrix
Cwith shape(vocab_size, embed_dim) - First layer weights
W1with shape(block_size*embed_dim, hidden_dim) - Hidden bias
b1and output biasb2 - Second layer weights
W2with shape(hidden_dim, vocab_size)
The forward pass in the MLP combines these matrices with non-linear activations:
C = torch.randn((vocab_size, embed_dim))
W1 = torch.randn((block_size*embed_dim, hidden_dim))
b1 = torch.randn(hidden_dim)
W2 = torch.randn((hidden_dim, vocab_size))
b2 = torch.randn(vocab_size)
# Forward pass
emb = C[X] # (N, 3, embed_dim)
h = torch.tanh(emb.view(-1, 6) @ W1 + b1) # flatten and apply hidden layer
logits = h @ W2 + b2 # (N, vocab_size)
Training Objectives and Optimization
Both models optimize the same cross-entropy loss against the next-character target, but differ fundamentally in how they learn:
-
Bigram: No optimization loop required. The model maximizes likelihood by counting observed frequencies, which represents the closed-form maximum likelihood solution for the bigram distribution.
-
MLP: Requires iterative gradient descent via backpropagation. The loss flows backward through
W2,W1, andC, updating continuous representations that capture semantic relationships between characters and positions.
loss = F.cross_entropy(logits, Y)
loss.backward()
# optimizer step updates C, W1, W2, b1, b2
Expressive Capacity and Computational Trade-offs
Bigram Limitations and Efficiency
The Bigram model captures only first-order dependencies (immediate co-occurrence statistics). While limited in expressiveness, it offers:
- Minimal memory: Only
|V|²counts to store - Instant training: Single pass counting requires no GPU
- Fast inference: Simple table lookup
MLP Advantages and Costs
The MLP can approximate arbitrary functions of its input context up to the capacity of its hidden layer. According to the karpathy/nn-zero-to-hero source code, this enables learning of phonetic patterns, syllable structures, and multi-character dependencies. However, this comes with:
- Higher memory usage: Scales with
embed_dim × vocab_sizeplus hidden layer parameters - Computational requirements: Matrix operations and gradient computation demand GPU/CPU acceleration
- Training time: Requires many epochs of gradient descent
Summary
- Bigram models use simple frequency counting of character pairs in
makemore_part1_bigrams.ipynb, offering fast training but limited to single-character context. - MLP models in
makemore_part2_mlp.ipynbemploy learnable embeddings and feed-forward layers to process context windows of three or more characters. - The transition from counting to backpropagation enables the model to learn continuous representations and non-linear interactions, dramatically improving predictive accuracy at the cost of computational resources.
- Bigram models serve as essential baselines and educational tools, while MLPs represent the foundation for modern neural language architectures.
Frequently Asked Questions
What is the fundamental architectural difference between Bigram and MLP language models?
The Bigram model is a tabular probability estimator that stores normalized counts of character pairs, containing no learnable parameters. The MLP is a feed-forward neural network with weight matrices (C, W1, W2) and bias vectors (b1, b2) that learns distributed representations through gradient descent.
How does context window size affect model performance?
Bigram models are constrained to a context window of exactly one character, making them unable to capture word roots or syllable patterns. The MLP model uses block_size = 3 (configurable) to attend to multiple preceding characters, enabling it to learn that "q" is almost always followed by "u" based on earlier context, even when processing character by character.
Which model trains faster and requires less memory?
The Bigram model trains orders of magnitude faster because it requires only a single counting pass through the corpus and stores simply a |V|² probability table. The MLP model requires iterative optimization over many epochs, stores embedding matrices and hidden layers, and performs expensive matrix multiplications during both forward and backward passes.
When should I choose a Bigram model over an MLP?
Use a Bigram model for rapid prototyping, educational demonstrations, sanity checks on data quality, or when working with extremely small corpora where neural networks would overfit. Choose an MLP model when you need to capture linguistic patterns beyond immediate character adjacency, or when building toward larger architectures like Transformers that require the foundation of learned embeddings and backpropagation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →