# Byte Pair Encoding (BPE) Tokenization: How GPT Tokenizers Work

> Learn how Byte Pair Encoding BPE tokenization works. Discover this sub-word segmentation algorithm for efficient text to integer token conversion in GPT models.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**Byte Pair Encoding (BPE) tokenization is a sub-word segmentation algorithm that converts raw text into integer tokens by iteratively merging the most frequent adjacent character pairs, creating a vocabulary that balances compression with coverage.**

In the *karpathy/nn-zero-to-hero* educational series, Byte Pair Encoding (BPE) tokenization is introduced in **Lecture 8** – "Let's build the GPT Tokenizer" – where the author explains that tokenizers are a separate, critical stage of large language model pipelines. While the lecture notebooks demonstrate the concepts, the production implementation resides in the external **minBPE** library referenced at line 83 of the repository's [`README.md`](https://github.com/karpathy/nn-zero-to-hero/blob/main/README.md).

## What Is Byte Pair Encoding (BPE) Tokenization?

Byte Pair Encoding (BPE) tokenization operates on the principle that text can be compressed by representing common character sequences as single tokens. Unlike word-based tokenization, which struggles with out-of-vocabulary terms, or character-level tokenization, which creates overly long sequences, BPE constructs a sub-word vocabulary that handles rare words by decomposing them into smaller units. Because the algorithm processes raw bytes, it naturally supports multilingual text without language-specific preprocessing.

## The BPE Algorithm Step-by-Step

The implementation in [`minbpe.py`](https://github.com/karpathy/nn-zero-to-hero/blob/main/minbpe.py) follows a six-stage training and inference pipeline:

1. **Data Preparation** – Raw text is split into an initial list of individual characters or bytes.
2. **Pair Counting** – All adjacent pairs in the corpus are counted, and the most frequent pair is identified.
3. **Merge Operation** – The selected pair is merged into a new token, and the underlying token list is updated.
4. **Vocabulary Expansion** – The new token is added to the BPE vocabulary; steps 2–4 repeat until reaching the target `vocab_size`.
5. **Encoding** – New strings are processed by repeatedly replacing substrings with the longest matching tokens from the learned vocabulary, outputting a list of token IDs.
6. **Decoding** – Token IDs are mapped back to their string representations and concatenated to reconstruct the original text.

## Implementation in the minBPE Library

According to the *nn-zero-to-hero* source code, the actual BPE implementation is provided by the **minBPE** library rather than the lecture repository itself. In [`minbpe.py`](https://github.com/karpathy/nn-zero-to-hero/blob/main/minbpe.py), the algorithm exposes a minimal API with three core functions:

```python

# Train on a corpus to generate vocabulary

tokens, vocab = minbpe.train(corpus, vocab_size=512)

# Convert string to token IDs

ids = minbpe.encode("Hello, world!", vocab)

# Convert token IDs back to string

text = minbpe.decode(ids, vocab)

```

The `train()` function implements the iterative merge loop, while `encode()` and `decode()` handle the inference-time conversion between text and integer sequences. This separation allows the same vocabulary to be reused across different text samples without retraining.

## Practical Code Examples

The following examples demonstrate the BPE workflow using the minBPE library, mirroring the style used in the Lecture 8 notebook.

### Training a BPE Tokenizer

```python
from minbpe import train, encode, decode

# Sample training corpus (typically a large text file)

corpus = [
    "the quick brown fox jumps over the lazy dog",
    "the quick brown fox was very quick"
]

# Train vocabulary with 256 tokens

tokens, vocab = train(corpus, vocab_size=256)

print("Number of tokens learned:", len(vocab))

# => Number of tokens learned: 256

```

### Encoding and Decoding Text

```python
sentence = "the quick brown fox"
ids = encode(sentence, vocab)      # Convert to token IDs

print("Token IDs:", ids)

# => Token IDs: [12, 45, 78, 34]

recovered = decode(ids, vocab)    # Reconstruct original string

print("Decoded text:", recovered)

# => Decoded text: the quick brown fox

```

### Integration with PyTorch Models

```python
import torch
from torch.nn import Embedding

# Convert token IDs to tensor

input_ids = torch.tensor(ids, dtype=torch.long).unsqueeze(0)  # shape (1, seq_len)

# Create embedding layer matching vocab size

embedding = Embedding(num_embeddings=len(vocab), embedding_dim=64)

# Generate token embeddings for neural network input

embeds = embedding(input_ids)   # shape (1, seq_len, 64)

print(embeds.shape)

# => torch.Size([1, 4, 64])

```

## Summary

- **Byte Pair Encoding (BPE) tokenization** compresses text by merging frequent character pairs into sub-word tokens, solving the trade-off between vocabulary size and coverage.
- The algorithm consists of six stages: data preparation, pair counting, merge operations, vocabulary expansion, encoding, and decoding.
- In the *nn-zero-to-hero* series, the implementation is provided by the **minBPE** library ([`minbpe.py`](https://github.com/karpathy/nn-zero-to-hero/blob/main/minbpe.py)), referenced in the [`README.md`](https://github.com/karpathy/nn-zero-to-hero/blob/main/README.md) at line 83.
- The minBPE API provides three essential functions: `train()` for vocabulary construction, `encode()` for text-to-IDs conversion, and `decode()` for reconstruction.
- Because BPE operates on byte-level data, it handles any Unicode text without language-specific preprocessing, making it ideal for multilingual GPT-style models.

## Frequently Asked Questions

### How does BPE tokenization handle rare or misspelled words?

BPE tokenization decomposes rare or misspelled words into smaller sub-word units that exist in the vocabulary. For example, a rare word like "tokenization" might be split into ["token", "ization"] or even smaller fragments if necessary, ensuring the model can process any input text without encountering unknown tokens.

### What is the difference between the nn-zero-to-hero repo and the minBPE library?

The *nn-zero-to-hero* repository contains educational materials and Lecture 8's "GPT Tokenizer" notebook explaining BPE concepts, while the **minBPE** library (`karpathy/minbpe`) provides the actual production implementation in [`minbpe.py`](https://github.com/karpathy/nn-zero-to-hero/blob/main/minbpe.py). The README at line 83 explicitly links to minBPE as the reference implementation for the code demonstrated in the lectures.

### Why does BPE use byte-level processing instead of character-level?

Byte-level processing allows BPE to handle any Unicode character, including emojis and non-Latin scripts, without requiring language-specific preprocessing or character set limitations. Since bytes are the fundamental unit of digital text, the tokenizer can process 256 possible byte values, ensuring universal text coverage across all languages and symbols.

### How do I choose the right vocabulary size for BPE training?

The optimal `vocab_size` depends on your corpus size and downstream model requirements. Smaller vocabularies (256–1,000 tokens) require the model to learn more composition patterns but reduce embedding memory, while larger vocabularies (10,000–50,000 tokens) capture more complete words but increase memory usage and risk overfitting on small datasets.