Byte Pair Encoding (BPE) Tokenization: How GPT Tokenizers Work
Byte Pair Encoding (BPE) tokenization is a sub-word segmentation algorithm that converts raw text into integer tokens by iteratively merging the most frequent adjacent character pairs, creating a vocabulary that balances compression with coverage.
In the karpathy/nn-zero-to-hero educational series, Byte Pair Encoding (BPE) tokenization is introduced in Lecture 8 – "Let's build the GPT Tokenizer" – where the author explains that tokenizers are a separate, critical stage of large language model pipelines. While the lecture notebooks demonstrate the concepts, the production implementation resides in the external minBPE library referenced at line 83 of the repository's README.md.
What Is Byte Pair Encoding (BPE) Tokenization?
Byte Pair Encoding (BPE) tokenization operates on the principle that text can be compressed by representing common character sequences as single tokens. Unlike word-based tokenization, which struggles with out-of-vocabulary terms, or character-level tokenization, which creates overly long sequences, BPE constructs a sub-word vocabulary that handles rare words by decomposing them into smaller units. Because the algorithm processes raw bytes, it naturally supports multilingual text without language-specific preprocessing.
The BPE Algorithm Step-by-Step
The implementation in minbpe.py follows a six-stage training and inference pipeline:
- Data Preparation – Raw text is split into an initial list of individual characters or bytes.
- Pair Counting – All adjacent pairs in the corpus are counted, and the most frequent pair is identified.
- Merge Operation – The selected pair is merged into a new token, and the underlying token list is updated.
- Vocabulary Expansion – The new token is added to the BPE vocabulary; steps 2–4 repeat until reaching the target
vocab_size. - Encoding – New strings are processed by repeatedly replacing substrings with the longest matching tokens from the learned vocabulary, outputting a list of token IDs.
- Decoding – Token IDs are mapped back to their string representations and concatenated to reconstruct the original text.
Implementation in the minBPE Library
According to the nn-zero-to-hero source code, the actual BPE implementation is provided by the minBPE library rather than the lecture repository itself. In minbpe.py, the algorithm exposes a minimal API with three core functions:
# Train on a corpus to generate vocabulary
tokens, vocab = minbpe.train(corpus, vocab_size=512)
# Convert string to token IDs
ids = minbpe.encode("Hello, world!", vocab)
# Convert token IDs back to string
text = minbpe.decode(ids, vocab)
The train() function implements the iterative merge loop, while encode() and decode() handle the inference-time conversion between text and integer sequences. This separation allows the same vocabulary to be reused across different text samples without retraining.
Practical Code Examples
The following examples demonstrate the BPE workflow using the minBPE library, mirroring the style used in the Lecture 8 notebook.
Training a BPE Tokenizer
from minbpe import train, encode, decode
# Sample training corpus (typically a large text file)
corpus = [
"the quick brown fox jumps over the lazy dog",
"the quick brown fox was very quick"
]
# Train vocabulary with 256 tokens
tokens, vocab = train(corpus, vocab_size=256)
print("Number of tokens learned:", len(vocab))
# => Number of tokens learned: 256
Encoding and Decoding Text
sentence = "the quick brown fox"
ids = encode(sentence, vocab) # Convert to token IDs
print("Token IDs:", ids)
# => Token IDs: [12, 45, 78, 34]
recovered = decode(ids, vocab) # Reconstruct original string
print("Decoded text:", recovered)
# => Decoded text: the quick brown fox
Integration with PyTorch Models
import torch
from torch.nn import Embedding
# Convert token IDs to tensor
input_ids = torch.tensor(ids, dtype=torch.long).unsqueeze(0) # shape (1, seq_len)
# Create embedding layer matching vocab size
embedding = Embedding(num_embeddings=len(vocab), embedding_dim=64)
# Generate token embeddings for neural network input
embeds = embedding(input_ids) # shape (1, seq_len, 64)
print(embeds.shape)
# => torch.Size([1, 4, 64])
Summary
- Byte Pair Encoding (BPE) tokenization compresses text by merging frequent character pairs into sub-word tokens, solving the trade-off between vocabulary size and coverage.
- The algorithm consists of six stages: data preparation, pair counting, merge operations, vocabulary expansion, encoding, and decoding.
- In the nn-zero-to-hero series, the implementation is provided by the minBPE library (
minbpe.py), referenced in theREADME.mdat line 83. - The minBPE API provides three essential functions:
train()for vocabulary construction,encode()for text-to-IDs conversion, anddecode()for reconstruction. - Because BPE operates on byte-level data, it handles any Unicode text without language-specific preprocessing, making it ideal for multilingual GPT-style models.
Frequently Asked Questions
How does BPE tokenization handle rare or misspelled words?
BPE tokenization decomposes rare or misspelled words into smaller sub-word units that exist in the vocabulary. For example, a rare word like "tokenization" might be split into ["token", "ization"] or even smaller fragments if necessary, ensuring the model can process any input text without encountering unknown tokens.
What is the difference between the nn-zero-to-hero repo and the minBPE library?
The nn-zero-to-hero repository contains educational materials and Lecture 8's "GPT Tokenizer" notebook explaining BPE concepts, while the minBPE library (karpathy/minbpe) provides the actual production implementation in minbpe.py. The README at line 83 explicitly links to minBPE as the reference implementation for the code demonstrated in the lectures.
Why does BPE use byte-level processing instead of character-level?
Byte-level processing allows BPE to handle any Unicode character, including emojis and non-Latin scripts, without requiring language-specific preprocessing or character set limitations. Since bytes are the fundamental unit of digital text, the tokenizer can process 256 possible byte values, ensuring universal text coverage across all languages and symbols.
How do I choose the right vocabulary size for BPE training?
The optimal vocab_size depends on your corpus size and downstream model requirements. Smaller vocabularies (256–1,000 tokens) require the model to learn more composition patterns but reduce embedding memory, while larger vocabularies (10,000–50,000 tokens) capture more complete words but increase memory usage and risk overfitting on small datasets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →