How to Implement a Character-Level Language Model from Scratch: A Complete Guide

A character-level language model predicts the next token in a sequence based on previous characters; the nn-zero-to-hero repository implements this by starting with a statistical bigram table and progressively upgrading to trainable neural networks, MLPs, and transformers.

This tutorial walks through building a character-level language model from the ground up using the educational PyTorch code in karpathy/nn-zero-to-hero. We begin with a simple count-based bigram approach, then transition to gradient-based learning with neural architectures that generalize beyond observed character pairs.

Data Preparation and Tokenization

Every character-level language model starts with a corpus and a vocabulary mapping. In lectures/makemore/makemore_part1_bigrams.ipynb, the dataset consists of names loaded from a text file:

words = open('names.txt').read().splitlines()

Tokenization requires converting characters to integer indices. The code creates two dictionaries: stoi (string-to-int) and itos (int-to-string). Special . tokens mark word boundaries (index 0), while the 26 letters occupy indices 1-27:

chars = sorted(set(''.join(words)))
stoi = {s: i+1 for i, s in enumerate(chars)}  # reserve 0 for '.'

stoi['.'] = 0
itos = {i: s for s, i in stoi.items()}

Each name is wrapped with start and end tokens so the model learns where words begin and terminate: ['.'] + list(name) + ['.'].

Building the Bigram Count Matrix

The simplest character-level language model is a bigram model that counts how often character i is followed by character j. The repository initializes a 27×27 integer matrix N and populates it by iterating through the dataset:

N = torch.zeros((27, 27), dtype=torch.int32)

for w in words:
    chs = ['.'] + list(w) + ['.']
    for ch1, ch2 in zip(chs, chs[1:]):
        ix1 = stoi[ch1]
        ix2 = stoi[ch2]
        N[ix1, ix2] += 1

After this loop, N[ix1, ix2] contains the raw frequency of the bigram pair found in the training data.

Converting Counts to Probabilities with Laplace Smoothing

Raw counts must become valid probability distributions. To avoid division by zero on unseen bigrams, the code applies Laplace smoothing (add-1 smoothing) and normalizes each row:

P = (N + 1).float()              # add 1 to every count

P /= P.sum(1, keepdims=True)     # normalize rows to sum to 1

P[i, j] now represents the probability of character j appearing immediately after character i, according to the learned statistics.

Sampling from the Bigram Model

Generating text requires sampling from these probability distributions sequentially. The sampling loop starts at the . token (index 0), draws from P[ix] using torch.multinomial, and continues until it generates another .:

g = torch.Generator().manual_seed(2147483647)
ix = 0
out = []

while True:
    p = P[ix]                    # probability distribution for current char

    ix = torch.multinomial(p, 1, generator=g).item()
    out.append(itos[ix])
    if ix == 0:                  # end token reached

        break

print(''.join(out))              # e.g., "mor.", "axx."

This produces novel character sequences that follow the statistical patterns of the training corpus.

Upgrading to a Trainable Neural Network

The repository's makemore_part2_mlp.ipynb replaces the static count table with a trainable weight matrix W. Instead of counting frequencies, the model learns parameters via gradient descent:


# One-hot encode the input character

xenc = F.one_hot(xs, num_classes=27).float()

# Forward pass: linear layer + softmax

logits = xenc @ W                # W shape: (27, 27)

probs = logits.exp()
probs = probs / probs.sum(1, keepdim=True)

# Negative log-likelihood loss

loss = -probs[torch.arange(N), ys].log().mean()

This neural approach can generalize to unseen bigrams and serves as the foundation for deeper architectures.

Extending to Deep Architectures and Transformers

The nn-zero-to-hero series scales the same character-level pipeline to modern depths:

  • Part 3 (makemore_part3_bn.ipynb) adds batch normalization and deeper MLP layers
  • Part 5 (makemore_part5_cnn1.ipynb) implements a WaveNet-style convolutional architecture
  • Part 7 culminates in a full GPT-style transformer, reusing the same stoi/itos tokenization and . boundary markers

This progression demonstrates how a simple character prediction task scales from statistical tables to state-of-the-art language models.

Complete Working Example

Below is a self-contained script implementing the full bigram pipeline. Save this as bigram_model.py after cloning the repository (ensure names.txt is in the same directory):

import torch

# 1. Load corpus

with open('names.txt') as f:
    words = f.read().splitlines()

# 2. Build vocabulary

chars = sorted(set(''.join(words)))
stoi = {c: i+1 for i, c in enumerate(chars)}
stoi['.'] = 0
itos = {i: c for c, i in stoi.items()}

# 3. Count bigrams

N = torch.zeros((27, 27), dtype=torch.int32)
for w in words:
    seq = ['.'] + list(w) + ['.']
    for a, b in zip(seq, seq[1:]):
        N[stoi[a], stoi[b]] += 1

# 4. Create probability table (with smoothing)

P = (N + 1).float()
P /= P.sum(1, keepdim=True)

# 5. Generate a name

g = torch.Generator().manual_seed(12345)
idx = 0
generated = []
while True:
    probs = P[idx]
    idx = torch.multinomial(probs, 1, generator=g).item()
    generated.append(itos[idx])
    if idx == 0:
        break

print('Generated:', ''.join(generated))

Summary

  • Character-level language models predict the next token given previous characters, starting from simple bigram statistics stored in a count matrix N
  • Laplace smoothing (N + 1) prevents zero probabilities on unseen character pairs before normalization
  • Sampling uses torch.multinomial to draw from the learned probability distribution row by row
  • The neural network version replaces the static table with a trainable weight matrix W and optimizes via negative log-likelihood loss
  • The same tokenization pipeline (stoi, itos, . delimiters) scales from bigram tables through MLPs, CNNs, and finally transformer architectures

Frequently Asked Questions

What is the difference between a bigram table and a neural character-level language model?

A bigram table stores explicit counts of character pairs in a 27×27 matrix N, while the neural version in makemore_part2_mlp.ipynb learns a weight matrix W that produces probability distributions through a softmax. The neural approach generalizes better to unseen sequences and provides a foundation for deeper architectures.

Why is Laplace smoothing (+1) necessary before normalizing the count matrix?

Without smoothing, any bigram absent from the training data would have zero probability, causing the model to break when encountering that sequence during sampling. Adding 1 to every count in N ensures every character transition has non-zero probability while minimally distorting the observed frequencies.

How does this character-level implementation scale to modern transformers?

The repository reuses the same vocabulary mappings (stoi/itos) and . boundary tokens throughout all lectures, gradually replacing the prediction mechanism—from count tables to MLPs to multi-head self-attention—while keeping the character-level prediction objective constant.

Which dataset does the nn-zero-to-hero repository use for training?

The examples use names.txt, a list of English names, but the architecture works with any text corpus. The code treats each line as a training example wrapped with . start/end tokens, making it adaptable to any character sequence dataset.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →