# Why Tokenization Causes Issues in LLMs and How to Debug Them

> Discover why tokenization causes LLM failures and learn effective debugging strategies. Address issues with vocabularies, BPE merges, and non-bijective mappings to improve your models.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**Tokenization causes LLM failures when fixed vocabularies, order-dependent BPE merges, and non-bijective text-to-token mappings inject noise that propagates through the entire model pipeline.**

Tokenization is the mandatory first stage in every LLM pipeline, converting raw text into discrete tokens that the model embeds and processes. In the **nn-zero-to-hero** repository by Andrej Karpathy, the standalone nature of tokenization is emphasized as the root cause of many perplexing model behaviors. Because the tokenizer operates as a separate, pre-trained component sitting between user input and model inference, any flaw in its Byte-Pair Encoding (BPE) algorithm or vocabulary propagates downstream into attention patterns and prediction quality.

## Architectural Reasons Tokenization Breaks LLMs

### Vocabulary Mismatch and Out-of-Vocabulary Tokens

The tokenizer’s vocabulary is **fixed after training**, meaning any character or byte sequence not present during training gets split into sub-optimal sub-tokens or the special "unknown" token. During the `encode()` operation, the model never sees the original character, leading to weak predictions for rare words or code symbols. This issue is highlighted in the **"Let's build the GPT Tokenizer"** lecture referenced in the repository’s [`README.md`](https://github.com/karpathy/nn-zero-to-hero/blob/main/README.md), where the author notes that *"a lot of weird behaviors and problems of LLMs actually trace back to tokenization"*.

### Inconsistent Splitting Across Versions

BPE merges frequent byte-pair sequences, but the merges are **order-dependent**. Small changes in the training corpus can cause identical words to tokenize differently across model versions. At `decode()` time, reconstructed text can contain misplaced spaces or missing punctuation, making evaluation and reproduction difficult.

### Loss of Linguistic Boundaries

Tokenizers operate on raw bytes, ignoring higher-level structure such as morphemes or word boundaries. This creates tokens that cut across meaningful units, affecting **attention patterns** in the transformer layers. The model may attend to "mid-word" tokens, confusing the learned semantics and degrading performance on linguistic tasks.

### Non-Reversible Mappings

The transformation `text → tokens → text` is **not bijective**; many distinct strings map to the same token sequence (e.g., "ﬁ" vs. "f i"). This leads to **ambiguous generation** where the model cannot deterministically decide which original form to output during decoding.

### Training Distribution Bias

The BPE algorithm greedily merges the most frequent byte pairs, biasing the vocabulary toward the training corpus. Out-of-distribution text—such as code, multilingual data, or rare Unicode characters—gets over-fragmented, resulting in **high perplexity** and poor downstream performance on specialized inputs.

## Debugging Tokenization Failures Step by Step

### Inspect Token Sequences Manually

Call `encode()` on problematic sentences and print both the token IDs and their string representations. This reveals whether words are being fragmented into excessive sub-tokens (e.g., "authentication" splitting into `['auth', 'enti', 'cation']`).

### Audit the Vocabulary

Use the tokenizer’s `vocab` (or `itos`/`stoi` maps) to verify if specific tokens exist. In the minimal BPE implementation found in [`minbpe.py`](https://github.com/karpathy/nn-zero-to-hero/blob/main/minbpe.py), you can inspect the `token2id` and `id2token` dictionaries to confirm whether a token is present or falling back to an unknown placeholder.

### Compare Against Reference Tokenizers

Run identical text through a well-known tokenizer (e.g., **tiktoken** for OpenAI models) and compare token counts. Large discrepancies often indicate missing merge rules or an outdated vocabulary in your implementation.

### Verify BPE Merge Rules

Print the list of merges used to build the tokenizer. If a frequent byte pair is missing from the `merges` list, the tokenizer will over-fragment text. Adding the missing merge—or retraining on a larger corpus—can dramatically reduce token count and improve consistency.

### Execute Round-Trip Tests

Encode a string, then decode it back immediately. If `decode(encode(s)) != s`, identify the offending token(s) causing the mismatch. This guarantees that the tokenizer is at least *reversible* for your test set; failures indicate non-bijective mappings that will corrupt generation.

### Visualize Token Boundaries

Plot token IDs over the original characters using visualization libraries. This helps spot off-by-one errors or unexpected splits that manual inspection might miss.

### Profile Token Frequency

Count how often each token appears in a validation set. Rare tokens often cause model instability, allowing you to prune or merge low-frequency entries to improve robustness.

## Practical Code Examples from the Repository

The **minbpe** implementation referenced in the nn-zero-to-hero lectures provides a minimal but complete tokenizer. Below is a practical debugging workflow:

```python
from minbpe import Tokenizer

# Load a pre-trained tokenizer (the repo provides a tiny example)

tok = Tokenizer.from_file("minbpe_vocab.txt")

text = "Tokenization can cause subtle bugs."
ids = tok.encode(text)
print("Token IDs :", ids)
print("Tokens    :", [tok.decode([i]) for i in ids])

# Round-trip check

assert tok.decode(ids) == text, "Round-trip failed!"

```

If the assertion fails, isolate the specific mismatch:

```python
decoded = tok.decode(ids)
for i, (orig, recon) in enumerate(zip(text, decoded)):
    if orig != recon:
        print(f"Mismatch at position {i}: '{orig}' → '{recon}' (token {ids[i]})")

```

To compare against a reference implementation:

```python
import tiktoken

ref = tiktoken.get_encoding('gpt2')
s = "Test string 123"

print('Our tokens:', tokenizer.encode(s))
print('Reference :', ref.encode(s))

```

## Key Repository Files and Source Locations

- **[`README.md`](https://github.com/karpathy/nn-zero-to-hero/blob/main/README.md)**: Contains the overview of Lecture 8 ("Let's build the GPT Tokenizer") and links to the tokenization discussion. Located at the repository root.
- **`lectures/makemore/makemore_part1_bigrams.ipynb`**: Demonstrates early language modeling where tokenization decisions directly affect the count matrix `N` used in bigram prediction.
- **[`minbpe.py`](https://github.com/karpathy/nn-zero-to-hero/blob/main/minbpe.py)** (from the external `minbpe` repository used in the course): Implements `encode()`, `decode`, and the BPE merge logic with `token2id` and `id2token` mappings.

## Summary

- **Tokenization is a separate, pre-trained stage** that sits between raw text and model embedding layers, making it a single point of failure for the entire pipeline.
- **Fixed vocabularies and order-dependent BPE merges** create out-of-vocabulary gaps, inconsistent splitting, and loss of linguistic boundaries that confuse attention mechanisms.
- **Non-reversible mappings** cause ambiguous generation where the model cannot recover the original text format.
- **Debug systematically** by inspecting token sequences, auditing vocabulary files, verifying merge rules, and executing round-trip encode/decode tests.
- **Validate against reference implementations** (like tiktoken) to identify divergence caused by missing merges or corpus bias.

## Frequently Asked Questions

### Why does tokenization cause unknown token errors even for simple words?

Unknown token errors occur when the BPE vocabulary is fixed after training and does not contain the specific byte sequence encountered during inference. In [`minbpe.py`](https://github.com/karpathy/nn-zero-to-hero/blob/main/minbpe.py), if a byte pair never appeared in the training corpus, the `encode()` function falls back to sub-optimal splits or unknown placeholders, causing the model to receive embeddings it never learned to process effectively.

### How can I check if my tokenizer is splitting words correctly?

Print the intermediate token representations using `[(i, tokenizer.decode([i])) for i in ids]` after calling `encode()`. This reveals exactly where splits occur. If common words fragment into excessive sub-tokens (more than 3-4 pieces), check the `merges` list in your BPE implementation to see if frequent pairs are missing from the training data.

### What causes round-trip failures where decode(encode(text)) ≠ text?

Round-trip failures indicate that the tokenizer’s mapping is **non-bijective**. This happens when multiple character sequences (like different Unicode normalization forms of the same character) map to the same token ID, or when whitespace handling during `decode()` introduces or removes spaces. Test with `assert tok.decode(tok.encode(s)) == s` to catch these errors before model training.

### Why do LLMs perform poorly on code or multilingual data?

BPE algorithms bias the vocabulary toward the most frequent byte pairs in the training corpus. If the corpus lacks code symbols or non-English characters, the tokenizer over-fragments these inputs into many small tokens, increasing sequence length and perplexity. Retraining the tokenizer on a diverse corpus or adopting byte-level BPE reduces this fragmentation bias.