How to Implement a Character-Level Language Model from Scratch: A Complete Guide
A character-level language model predicts the next token in a sequence based on previous characters; the nn-zero-to-hero repository implements this by starting with a statistical bigram table and progressively upgrading to trainable neural networks, MLPs, and transformers.
This tutorial walks through building a character-level language model from the ground up using the educational PyTorch code in karpathy/nn-zero-to-hero. We begin with a simple count-based bigram approach, then transition to gradient-based learning with neural architectures that generalize beyond observed character pairs.
Data Preparation and Tokenization
Every character-level language model starts with a corpus and a vocabulary mapping. In lectures/makemore/makemore_part1_bigrams.ipynb, the dataset consists of names loaded from a text file:
words = open('names.txt').read().splitlines()
Tokenization requires converting characters to integer indices. The code creates two dictionaries: stoi (string-to-int) and itos (int-to-string). Special . tokens mark word boundaries (index 0), while the 26 letters occupy indices 1-27:
chars = sorted(set(''.join(words)))
stoi = {s: i+1 for i, s in enumerate(chars)} # reserve 0 for '.'
stoi['.'] = 0
itos = {i: s for s, i in stoi.items()}
Each name is wrapped with start and end tokens so the model learns where words begin and terminate: ['.'] + list(name) + ['.'].
Building the Bigram Count Matrix
The simplest character-level language model is a bigram model that counts how often character i is followed by character j. The repository initializes a 27×27 integer matrix N and populates it by iterating through the dataset:
N = torch.zeros((27, 27), dtype=torch.int32)
for w in words:
chs = ['.'] + list(w) + ['.']
for ch1, ch2 in zip(chs, chs[1:]):
ix1 = stoi[ch1]
ix2 = stoi[ch2]
N[ix1, ix2] += 1
After this loop, N[ix1, ix2] contains the raw frequency of the bigram pair found in the training data.
Converting Counts to Probabilities with Laplace Smoothing
Raw counts must become valid probability distributions. To avoid division by zero on unseen bigrams, the code applies Laplace smoothing (add-1 smoothing) and normalizes each row:
P = (N + 1).float() # add 1 to every count
P /= P.sum(1, keepdims=True) # normalize rows to sum to 1
P[i, j] now represents the probability of character j appearing immediately after character i, according to the learned statistics.
Sampling from the Bigram Model
Generating text requires sampling from these probability distributions sequentially. The sampling loop starts at the . token (index 0), draws from P[ix] using torch.multinomial, and continues until it generates another .:
g = torch.Generator().manual_seed(2147483647)
ix = 0
out = []
while True:
p = P[ix] # probability distribution for current char
ix = torch.multinomial(p, 1, generator=g).item()
out.append(itos[ix])
if ix == 0: # end token reached
break
print(''.join(out)) # e.g., "mor.", "axx."
This produces novel character sequences that follow the statistical patterns of the training corpus.
Upgrading to a Trainable Neural Network
The repository's makemore_part2_mlp.ipynb replaces the static count table with a trainable weight matrix W. Instead of counting frequencies, the model learns parameters via gradient descent:
# One-hot encode the input character
xenc = F.one_hot(xs, num_classes=27).float()
# Forward pass: linear layer + softmax
logits = xenc @ W # W shape: (27, 27)
probs = logits.exp()
probs = probs / probs.sum(1, keepdim=True)
# Negative log-likelihood loss
loss = -probs[torch.arange(N), ys].log().mean()
This neural approach can generalize to unseen bigrams and serves as the foundation for deeper architectures.
Extending to Deep Architectures and Transformers
The nn-zero-to-hero series scales the same character-level pipeline to modern depths:
- Part 3 (
makemore_part3_bn.ipynb) adds batch normalization and deeper MLP layers - Part 5 (
makemore_part5_cnn1.ipynb) implements a WaveNet-style convolutional architecture - Part 7 culminates in a full GPT-style transformer, reusing the same
stoi/itostokenization and.boundary markers
This progression demonstrates how a simple character prediction task scales from statistical tables to state-of-the-art language models.
Complete Working Example
Below is a self-contained script implementing the full bigram pipeline. Save this as bigram_model.py after cloning the repository (ensure names.txt is in the same directory):
import torch
# 1. Load corpus
with open('names.txt') as f:
words = f.read().splitlines()
# 2. Build vocabulary
chars = sorted(set(''.join(words)))
stoi = {c: i+1 for i, c in enumerate(chars)}
stoi['.'] = 0
itos = {i: c for c, i in stoi.items()}
# 3. Count bigrams
N = torch.zeros((27, 27), dtype=torch.int32)
for w in words:
seq = ['.'] + list(w) + ['.']
for a, b in zip(seq, seq[1:]):
N[stoi[a], stoi[b]] += 1
# 4. Create probability table (with smoothing)
P = (N + 1).float()
P /= P.sum(1, keepdim=True)
# 5. Generate a name
g = torch.Generator().manual_seed(12345)
idx = 0
generated = []
while True:
probs = P[idx]
idx = torch.multinomial(probs, 1, generator=g).item()
generated.append(itos[idx])
if idx == 0:
break
print('Generated:', ''.join(generated))
Summary
- Character-level language models predict the next token given previous characters, starting from simple bigram statistics stored in a count matrix
N - Laplace smoothing (
N + 1) prevents zero probabilities on unseen character pairs before normalization - Sampling uses
torch.multinomialto draw from the learned probability distribution row by row - The neural network version replaces the static table with a trainable weight matrix
Wand optimizes via negative log-likelihood loss - The same tokenization pipeline (
stoi,itos,.delimiters) scales from bigram tables through MLPs, CNNs, and finally transformer architectures
Frequently Asked Questions
What is the difference between a bigram table and a neural character-level language model?
A bigram table stores explicit counts of character pairs in a 27×27 matrix N, while the neural version in makemore_part2_mlp.ipynb learns a weight matrix W that produces probability distributions through a softmax. The neural approach generalizes better to unseen sequences and provides a foundation for deeper architectures.
Why is Laplace smoothing (+1) necessary before normalizing the count matrix?
Without smoothing, any bigram absent from the training data would have zero probability, causing the model to break when encountering that sequence during sampling. Adding 1 to every count in N ensures every character transition has non-zero probability while minimally distorting the observed frequencies.
How does this character-level implementation scale to modern transformers?
The repository reuses the same vocabulary mappings (stoi/itos) and . boundary tokens throughout all lectures, gradually replacing the prediction mechanism—from count tables to MLPs to multi-head self-attention—while keeping the character-level prediction objective constant.
Which dataset does the nn-zero-to-hero repository use for training?
The examples use names.txt, a list of English names, but the architecture works with any text corpus. The code treats each line as a training example wrapped with . start/end tokens, making it adaptable to any character sequence dataset.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →