# Understanding Tokenization Algorithms: BPE, WordPiece, and SentencePiece

> Master BPE, WordPiece, and SentencePiece tokenization algorithms used in LLMs. Learn their differences in merge selection, frequency, likelihood, and Unicode handling for multilingual NLP.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-26

---

**Byte‑Pair Encoding (BPE), WordPiece, and SentencePiece are the three dominant sub‑word tokenization algorithms used in modern LLMs, differing primarily in how they select symbol merges—BPE uses raw frequency, WordPiece uses likelihood ratios, and SentencePiece treats text as a raw Unicode stream enabling multilingual support.**

Tokenization bridges raw text and neural model vocabularies, making the choice of algorithm a critical design decision in AI engineering. This article explores the mechanics of BPE, WordPiece, and SentencePiece as implemented in the `rohitg00/ai-engineering-from-scratch` repository, providing reference implementations and practical guidance for selecting the right approach.

## How BPE Tokenization Works

**Byte‑Pair Encoding (BPE)** constructs vocabularies through iterative merging of the most frequent adjacent symbol pairs. According to the curriculum in [`phases/10-llms-from-scratch/01-tokenizers/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/01-tokenizers/docs/en.md), the algorithm initializes with a base vocabulary of individual characters (or bytes), then greedily merges the pair with the highest `count(pair)` until reaching a target vocabulary size.

The repository provides a **byte‑level BPE** implementation in [`phases/10-llms-from-scratch/01-tokenizers/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/01-tokenizers/code/main.py) that operates on raw bytes (0‑255) rather than characters. This guarantees **no out‑of‑vocabulary (OOV) tokens**, as any Unicode character can be decomposed into byte sequences.

Key characteristics of BPE include:
- Deterministic merges based purely on frequency statistics
- Optional pre‑tokenization that splits on whitespace before merging
- Cross‑word boundary merges unless explicitly prevented by pre‑processing

## WordPiece vs BPE: Likelihood‑Based Merging

**WordPiece** follows a similar bottom‑up approach to BPE but changes the merge criterion from raw frequency to **likelihood maximization**. As detailed in the subword tokenization lesson at [`phases/05-nlp-foundations-to-advanced/19-subword-tokenization/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/05-nlp-foundations-to-advanced/19-subword-tokenization/docs/en.md), WordPiece selects merges that maximize the ratio `count(AB)/(count(A)·count(B))`.

This probabilistic approach produces a more linguistically meaningful vocabulary by preferring **surprising co‑occurrences** over common adjacent pairs. WordPiece also introduces explicit word‑boundary markers by prefixing continuation sub‑words with `"##"`, which proves beneficial for token‑level classification tasks in BERT‑family models.

## SentencePiece: Language‑Agnostic Tokenization

**SentencePiece** takes a fundamentally different approach by treating the input as a raw Unicode stream, including whitespace characters. This design, documented in [`phases/10-llms-from-scratch/02-building-a-tokenizer/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/02-building-a-tokenizer/docs/en.md), enables language‑agnostic training on corpora without spaces (such as Chinese, Thai, or Japanese).

Key differentiators of SentencePiece include:
- Whitespace encoding as a special token (`▁`) rather than pre‑tokenization
- Support for both **BPE** and **Unigram** training modes within the same framework
- Optional `byte_fallback=True` for handling rare scripts by falling back to UTF‑8 byte sequences

Because SentencePiece requires no language‑specific pre‑tokenizer, it serves as the standard for multilingual models like LLaMA and T5.

## Implementation Examples from the Repository

### Byte‑Level BPE from Scratch

The curriculum includes a minimal‑dependency `BPETokenizer` class in [`phases/10-llms-from-scratch/01-tokenizers/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/01-tokenizers/code/main.py) that implements byte‑level BPE training:

```python
from phases_10_llms_from_scratch_01_tokenizers_code_main import BPETokenizer

# Initialise with target vocab size (including 256 base bytes)

tokenizer = BPETokenizer(vocab_size=300)

# Train on in-memory corpus

corpus = [
    "lower lowest newest",
    "happiness is a state of mind",
    "the quick brown fox jumps over the lazy dog"
]
tokenizer.train(corpus)

# Encode and decode

ids = tokenizer.encode("the quick brown fox")
text = tokenizer.decode(ids)

```

This implementation handles special token injection and provides deterministic encoding through merge tables stored during training.

### WordPiece with Hugging Face

For production BERT‑style models, the repository demonstrates WordPiece training using the Hugging Face `tokenizers` library:

```python
from tokenizers import Tokenizer, models, pre_tokenizers, trainers

tokenizer = Tokenizer(models.WordPiece(vocab=None, unk_token="[UNK]"))
tokenizer.pre_tokenizer = pre_tokenizers.Whitespace()

trainer = trainers.WordPieceTrainer(
    vocab_size=30522, 
    special_tokens=["[UNK]", "[CLS]", "[SEP]", "[MASK]"]
)

tokenizer.train(["BERT is a transformer model.", "WordPiece merges are likelihood based."], trainer)
output = tokenizer.encode("BERT tokenization")
print(output.tokens)  # Shows "##" continuation markers

```

### Training SentencePiece Models

The SentencePiece implementation supports both BPE and Unigram algorithms through a unified API:

```python
import sentencepiece as spm

# Train BPE model on raw text file

spm.SentencePieceTrainer.train(
    input='data/corpus.txt',
    model_prefix='spm_demo',
    vocab_size=32000,
    character_coverage=0.9995,
    model_type='bpe'  # or 'unigram'

)

sp = spm.SentencePieceProcessor()
sp.load('spm_demo.model')

# Encode with explicit whitespace tokenization

ids = sp.encode_as_ids("SentencePiece works on any language!")
decoded = sp.decode_ids(ids)

```

## Choosing the Right Tokenization Algorithm

| Scenario | Recommended Algorithm | Reason |
|----------|----------------------|--------|
| English‑only models with abundant data | **Byte‑level BPE** | Fixed 256‑byte base eliminates OOV; deterministic merges optimize for speed |
| BERT‑style masked language modeling | **WordPiece** | `"##"` prefixes signal word boundaries explicitly for token‑classification tasks |
| Multilingual corpora without spaces | **SentencePiece** | Raw Unicode processing handles scripts like Chinese or Thai without pre‑tokenization |
| Research and educational purposes | **Scratch BPE** | The `BPETokenizer` implementation in [`phases/10-llms-from-scratch/01-tokenizers/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/01-tokenizers/code/main.py) allows easy instrumentation and merge table visualization |

The curriculum emphasizes that **the tokenizer is a contract**—downstream lessons in metrics computation and model training assume consistent tokenization. Changing algorithms without updating dependent components silently breaks pipeline integrity.

## Summary

- **BPE** merges based on raw pair frequency, with byte‑level variants guaranteeing no unknown tokens through 0‑255 byte vocabularies.
- **WordPiece** uses likelihood ratios `count(AB)/(count(A)·count(B))` to select merges, producing linguistically meaningful sub‑words with explicit continuation markers.
- **SentencePiece** treats whitespace as a regular character (`▁`), enabling language‑agnostic training on Unicode corpora without pre‑tokenization.
- The `rohitg00/ai-engineering-from-scratch` repository provides reference implementations for all three algorithms, with the scratch BPE tokenizer serving as the primary educational resource in [`phases/10-llms-from-scratch/01-tokenizers/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/01-tokenizers/code/main.py).

## Frequently Asked Questions

### What is the main difference between BPE and WordPiece?

BPE selects merges based on the highest frequency count of adjacent symbol pairs, while WordPiece selects merges that maximize the likelihood ratio `count(AB)/(count(A)·count(B))`. This causes WordPiece to prefer statistically surprising co‑occurrences over common pairs, and it introduces `"##"` prefixes to mark continuation sub‑words.

### Why does SentencePiece use the ▁ character?

The `▁` symbol represents whitespace in SentencePiece's vocabulary, allowing the algorithm to treat the input as a raw Unicode stream without pre‑tokenization. This design enables processing of languages that do not use space delimiters (such as Chinese or Thai) while maintaining reversible encoding and decoding operations.

### Can SentencePiece use BPE or only Unigram?

SentencePiece supports both algorithms through the `model_type` parameter. Setting `model_type='bpe'` trains using Byte‑Pair Encoding merges, while `model_type='unigram'` uses the Unigram language model approach that prunes vocabulary items based on marginal likelihood. Both modes operate on the same underlying Unicode stream representation.

### How does byte-level BPE prevent unknown tokens?

Byte‑level BPE initializes its base vocabulary with all 256 possible byte values (0‑255). Since any Unicode character can be expressed as a sequence of UTF‑8 bytes, the tokenizer can encode any input text by falling back to byte representations for rare characters. This guarantees that the encoder never encounters a character it cannot represent, eliminating OOV errors entirely.