Understanding Tokenization Algorithms: BPE, WordPiece, and SentencePiece

Byte‑Pair Encoding (BPE), WordPiece, and SentencePiece are the three dominant sub‑word tokenization algorithms used in modern LLMs, differing primarily in how they select symbol merges—BPE uses raw frequency, WordPiece uses likelihood ratios, and SentencePiece treats text as a raw Unicode stream enabling multilingual support.

Tokenization bridges raw text and neural model vocabularies, making the choice of algorithm a critical design decision in AI engineering. This article explores the mechanics of BPE, WordPiece, and SentencePiece as implemented in the rohitg00/ai-engineering-from-scratch repository, providing reference implementations and practical guidance for selecting the right approach.

How BPE Tokenization Works

Byte‑Pair Encoding (BPE) constructs vocabularies through iterative merging of the most frequent adjacent symbol pairs. According to the curriculum in phases/10-llms-from-scratch/01-tokenizers/docs/en.md, the algorithm initializes with a base vocabulary of individual characters (or bytes), then greedily merges the pair with the highest count(pair) until reaching a target vocabulary size.

The repository provides a byte‑level BPE implementation in phases/10-llms-from-scratch/01-tokenizers/code/main.py that operates on raw bytes (0‑255) rather than characters. This guarantees no out‑of‑vocabulary (OOV) tokens, as any Unicode character can be decomposed into byte sequences.

Key characteristics of BPE include:

  • Deterministic merges based purely on frequency statistics
  • Optional pre‑tokenization that splits on whitespace before merging
  • Cross‑word boundary merges unless explicitly prevented by pre‑processing

WordPiece vs BPE: Likelihood‑Based Merging

WordPiece follows a similar bottom‑up approach to BPE but changes the merge criterion from raw frequency to likelihood maximization. As detailed in the subword tokenization lesson at phases/05-nlp-foundations-to-advanced/19-subword-tokenization/docs/en.md, WordPiece selects merges that maximize the ratio count(AB)/(count(A)·count(B)).

This probabilistic approach produces a more linguistically meaningful vocabulary by preferring surprising co‑occurrences over common adjacent pairs. WordPiece also introduces explicit word‑boundary markers by prefixing continuation sub‑words with "##", which proves beneficial for token‑level classification tasks in BERT‑family models.

SentencePiece: Language‑Agnostic Tokenization

SentencePiece takes a fundamentally different approach by treating the input as a raw Unicode stream, including whitespace characters. This design, documented in phases/10-llms-from-scratch/02-building-a-tokenizer/docs/en.md, enables language‑agnostic training on corpora without spaces (such as Chinese, Thai, or Japanese).

Key differentiators of SentencePiece include:

  • Whitespace encoding as a special token (▁) rather than pre‑tokenization
  • Support for both BPE and Unigram training modes within the same framework
  • Optional byte_fallback=True for handling rare scripts by falling back to UTF‑8 byte sequences

Because SentencePiece requires no language‑specific pre‑tokenizer, it serves as the standard for multilingual models like LLaMA and T5.

Implementation Examples from the Repository

Byte‑Level BPE from Scratch

The curriculum includes a minimal‑dependency BPETokenizer class in phases/10-llms-from-scratch/01-tokenizers/code/main.py that implements byte‑level BPE training:

from phases_10_llms_from_scratch_01_tokenizers_code_main import BPETokenizer

# Initialise with target vocab size (including 256 base bytes)

tokenizer = BPETokenizer(vocab_size=300)

# Train on in-memory corpus

corpus = [
    "lower lowest newest",
    "happiness is a state of mind",
    "the quick brown fox jumps over the lazy dog"
]
tokenizer.train(corpus)

# Encode and decode

ids = tokenizer.encode("the quick brown fox")
text = tokenizer.decode(ids)

This implementation handles special token injection and provides deterministic encoding through merge tables stored during training.

WordPiece with Hugging Face

For production BERT‑style models, the repository demonstrates WordPiece training using the Hugging Face tokenizers library:

from tokenizers import Tokenizer, models, pre_tokenizers, trainers

tokenizer = Tokenizer(models.WordPiece(vocab=None, unk_token="[UNK]"))
tokenizer.pre_tokenizer = pre_tokenizers.Whitespace()

trainer = trainers.WordPieceTrainer(
    vocab_size=30522, 
    special_tokens=["[UNK]", "[CLS]", "[SEP]", "[MASK]"]
)

tokenizer.train(["BERT is a transformer model.", "WordPiece merges are likelihood based."], trainer)
output = tokenizer.encode("BERT tokenization")
print(output.tokens)  # Shows "##" continuation markers

Training SentencePiece Models

The SentencePiece implementation supports both BPE and Unigram algorithms through a unified API:

import sentencepiece as spm

# Train BPE model on raw text file

spm.SentencePieceTrainer.train(
    input='data/corpus.txt',
    model_prefix='spm_demo',
    vocab_size=32000,
    character_coverage=0.9995,
    model_type='bpe'  # or 'unigram'

)

sp = spm.SentencePieceProcessor()
sp.load('spm_demo.model')

# Encode with explicit whitespace tokenization

ids = sp.encode_as_ids("SentencePiece works on any language!")
decoded = sp.decode_ids(ids)

Choosing the Right Tokenization Algorithm

Scenario Recommended Algorithm Reason
English‑only models with abundant data Byte‑level BPE Fixed 256‑byte base eliminates OOV; deterministic merges optimize for speed
BERT‑style masked language modeling WordPiece "##" prefixes signal word boundaries explicitly for token‑classification tasks
Multilingual corpora without spaces SentencePiece Raw Unicode processing handles scripts like Chinese or Thai without pre‑tokenization
Research and educational purposes Scratch BPE The BPETokenizer implementation in phases/10-llms-from-scratch/01-tokenizers/code/main.py allows easy instrumentation and merge table visualization

The curriculum emphasizes that the tokenizer is a contract—downstream lessons in metrics computation and model training assume consistent tokenization. Changing algorithms without updating dependent components silently breaks pipeline integrity.

Summary

  • BPE merges based on raw pair frequency, with byte‑level variants guaranteeing no unknown tokens through 0‑255 byte vocabularies.
  • WordPiece uses likelihood ratios count(AB)/(count(A)·count(B)) to select merges, producing linguistically meaningful sub‑words with explicit continuation markers.
  • SentencePiece treats whitespace as a regular character (▁), enabling language‑agnostic training on Unicode corpora without pre‑tokenization.
  • The rohitg00/ai-engineering-from-scratch repository provides reference implementations for all three algorithms, with the scratch BPE tokenizer serving as the primary educational resource in phases/10-llms-from-scratch/01-tokenizers/code/main.py.

Frequently Asked Questions

What is the main difference between BPE and WordPiece?

BPE selects merges based on the highest frequency count of adjacent symbol pairs, while WordPiece selects merges that maximize the likelihood ratio count(AB)/(count(A)·count(B)). This causes WordPiece to prefer statistically surprising co‑occurrences over common pairs, and it introduces "##" prefixes to mark continuation sub‑words.

Why does SentencePiece use the ▁ character?

The ▁ symbol represents whitespace in SentencePiece's vocabulary, allowing the algorithm to treat the input as a raw Unicode stream without pre‑tokenization. This design enables processing of languages that do not use space delimiters (such as Chinese or Thai) while maintaining reversible encoding and decoding operations.

Can SentencePiece use BPE or only Unigram?

SentencePiece supports both algorithms through the model_type parameter. Setting model_type='bpe' trains using Byte‑Pair Encoding merges, while model_type='unigram' uses the Unigram language model approach that prunes vocabulary items based on marginal likelihood. Both modes operate on the same underlying Unicode stream representation.

How does byte-level BPE prevent unknown tokens?

Byte‑level BPE initializes its base vocabulary with all 256 possible byte values (0‑255). Since any Unicode character can be expressed as a sequence of UTF‑8 bytes, the tokenizer can encode any input text by falling back to byte representations for rare characters. This guarantees that the encoder never encounters a character it cannot represent, eliminating OOV errors entirely.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →