Why Tokenization Causes Issues in LLMs and How to Debug Them

Tokenization causes LLM failures when fixed vocabularies, order-dependent BPE merges, and non-bijective text-to-token mappings inject noise that propagates through the entire model pipeline.

Tokenization is the mandatory first stage in every LLM pipeline, converting raw text into discrete tokens that the model embeds and processes. In the nn-zero-to-hero repository by Andrej Karpathy, the standalone nature of tokenization is emphasized as the root cause of many perplexing model behaviors. Because the tokenizer operates as a separate, pre-trained component sitting between user input and model inference, any flaw in its Byte-Pair Encoding (BPE) algorithm or vocabulary propagates downstream into attention patterns and prediction quality.

Architectural Reasons Tokenization Breaks LLMs

Vocabulary Mismatch and Out-of-Vocabulary Tokens

The tokenizer’s vocabulary is fixed after training, meaning any character or byte sequence not present during training gets split into sub-optimal sub-tokens or the special "unknown" token. During the encode() operation, the model never sees the original character, leading to weak predictions for rare words or code symbols. This issue is highlighted in the "Let's build the GPT Tokenizer" lecture referenced in the repository’s README.md, where the author notes that "a lot of weird behaviors and problems of LLMs actually trace back to tokenization".

Inconsistent Splitting Across Versions

BPE merges frequent byte-pair sequences, but the merges are order-dependent. Small changes in the training corpus can cause identical words to tokenize differently across model versions. At decode() time, reconstructed text can contain misplaced spaces or missing punctuation, making evaluation and reproduction difficult.

Loss of Linguistic Boundaries

Tokenizers operate on raw bytes, ignoring higher-level structure such as morphemes or word boundaries. This creates tokens that cut across meaningful units, affecting attention patterns in the transformer layers. The model may attend to "mid-word" tokens, confusing the learned semantics and degrading performance on linguistic tasks.

Non-Reversible Mappings

The transformation text → tokens → text is not bijective; many distinct strings map to the same token sequence (e.g., "fi" vs. "f i"). This leads to ambiguous generation where the model cannot deterministically decide which original form to output during decoding.

Training Distribution Bias

The BPE algorithm greedily merges the most frequent byte pairs, biasing the vocabulary toward the training corpus. Out-of-distribution text—such as code, multilingual data, or rare Unicode characters—gets over-fragmented, resulting in high perplexity and poor downstream performance on specialized inputs.

Debugging Tokenization Failures Step by Step

Inspect Token Sequences Manually

Call encode() on problematic sentences and print both the token IDs and their string representations. This reveals whether words are being fragmented into excessive sub-tokens (e.g., "authentication" splitting into ['auth', 'enti', 'cation']).

Audit the Vocabulary

Use the tokenizer’s vocab (or itos/stoi maps) to verify if specific tokens exist. In the minimal BPE implementation found in minbpe.py, you can inspect the token2id and id2token dictionaries to confirm whether a token is present or falling back to an unknown placeholder.

Compare Against Reference Tokenizers

Run identical text through a well-known tokenizer (e.g., tiktoken for OpenAI models) and compare token counts. Large discrepancies often indicate missing merge rules or an outdated vocabulary in your implementation.

Verify BPE Merge Rules

Print the list of merges used to build the tokenizer. If a frequent byte pair is missing from the merges list, the tokenizer will over-fragment text. Adding the missing merge—or retraining on a larger corpus—can dramatically reduce token count and improve consistency.

Execute Round-Trip Tests

Encode a string, then decode it back immediately. If decode(encode(s)) != s, identify the offending token(s) causing the mismatch. This guarantees that the tokenizer is at least reversible for your test set; failures indicate non-bijective mappings that will corrupt generation.

Visualize Token Boundaries

Plot token IDs over the original characters using visualization libraries. This helps spot off-by-one errors or unexpected splits that manual inspection might miss.

Profile Token Frequency

Count how often each token appears in a validation set. Rare tokens often cause model instability, allowing you to prune or merge low-frequency entries to improve robustness.

Practical Code Examples from the Repository

The minbpe implementation referenced in the nn-zero-to-hero lectures provides a minimal but complete tokenizer. Below is a practical debugging workflow:

from minbpe import Tokenizer

# Load a pre-trained tokenizer (the repo provides a tiny example)

tok = Tokenizer.from_file("minbpe_vocab.txt")

text = "Tokenization can cause subtle bugs."
ids = tok.encode(text)
print("Token IDs :", ids)
print("Tokens    :", [tok.decode([i]) for i in ids])

# Round-trip check

assert tok.decode(ids) == text, "Round-trip failed!"

If the assertion fails, isolate the specific mismatch:

decoded = tok.decode(ids)
for i, (orig, recon) in enumerate(zip(text, decoded)):
    if orig != recon:
        print(f"Mismatch at position {i}: '{orig}' → '{recon}' (token {ids[i]})")

To compare against a reference implementation:

import tiktoken

ref = tiktoken.get_encoding('gpt2')
s = "Test string 123"

print('Our tokens:', tokenizer.encode(s))
print('Reference :', ref.encode(s))

Key Repository Files and Source Locations

  • README.md: Contains the overview of Lecture 8 ("Let's build the GPT Tokenizer") and links to the tokenization discussion. Located at the repository root.
  • lectures/makemore/makemore_part1_bigrams.ipynb: Demonstrates early language modeling where tokenization decisions directly affect the count matrix N used in bigram prediction.
  • minbpe.py (from the external minbpe repository used in the course): Implements encode(), decode, and the BPE merge logic with token2id and id2token mappings.

Summary

  • Tokenization is a separate, pre-trained stage that sits between raw text and model embedding layers, making it a single point of failure for the entire pipeline.
  • Fixed vocabularies and order-dependent BPE merges create out-of-vocabulary gaps, inconsistent splitting, and loss of linguistic boundaries that confuse attention mechanisms.
  • Non-reversible mappings cause ambiguous generation where the model cannot recover the original text format.
  • Debug systematically by inspecting token sequences, auditing vocabulary files, verifying merge rules, and executing round-trip encode/decode tests.
  • Validate against reference implementations (like tiktoken) to identify divergence caused by missing merges or corpus bias.

Frequently Asked Questions

Why does tokenization cause unknown token errors even for simple words?

Unknown token errors occur when the BPE vocabulary is fixed after training and does not contain the specific byte sequence encountered during inference. In minbpe.py, if a byte pair never appeared in the training corpus, the encode() function falls back to sub-optimal splits or unknown placeholders, causing the model to receive embeddings it never learned to process effectively.

How can I check if my tokenizer is splitting words correctly?

Print the intermediate token representations using [(i, tokenizer.decode([i])) for i in ids] after calling encode(). This reveals exactly where splits occur. If common words fragment into excessive sub-tokens (more than 3-4 pieces), check the merges list in your BPE implementation to see if frequent pairs are missing from the training data.

What causes round-trip failures where decode(encode(text)) ≠ text?

Round-trip failures indicate that the tokenizer’s mapping is non-bijective. This happens when multiple character sequences (like different Unicode normalization forms of the same character) map to the same token ID, or when whitespace handling during decode() introduces or removes spaces. Test with assert tok.decode(tok.encode(s)) == s to catch these errors before model training.

Why do LLMs perform poorly on code or multilingual data?

BPE algorithms bias the vocabulary toward the most frequent byte pairs in the training corpus. If the corpus lacks code symbols or non-English characters, the tokenizer over-fragments these inputs into many small tokens, increasing sequence length and perplexity. Retraining the tokenizer on a diverse corpus or adopting byte-level BPE reduces this fragmentation bias.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →