How to Extend the Tiktoken BPE Tokenizer with New Tokens: A Complete Guide
To extend the Tiktoken BPE tokenizer with new tokens, load a custom vocabulary using tiktoken.load.load_tiktoken_bpe, instantiate a new Encoding object with augmented special tokens, and wrap it in a helper class that exposes the custom IDs to your language model.
The tiktoken library provides OpenAI’s high-performance byte-pair encoding (BPE) implementation, but it does not expose a public API for dynamically adding tokens to an existing encoding. In the rasbt/LLMs-from-scratch repository, developers extend the tokenizer by rebuilding the Encoding object from scratch with an augmented vocabulary and merge table, then wrapping it in a model-specific class like Llama3Tokenizer.
Why Tiktoken Requires Rebuilding
Unlike the Hugging Face tokenizers library, tiktoken does not support an add_tokens method on existing encodings. As demonstrated in ch05/09_extending-tokenizers/extend-tiktoken.ipynb, the library’s design treats tokenizers as immutable objects. To add custom tokens, you must re-instantiate tiktoken.Encoding with a new vocabulary dictionary that includes your extra tokens mapped to unique integer IDs.
Step-by-Step Implementation
Load Custom BPE Vocabulary and Merge Rules
Start by loading the base vocabulary and merge rules using load_tiktoken_bpe from the tiktoken.load module. This function expects a directory containing vocab.json and merges.txt files.
from tiktoken.load import load_tiktoken_bpe
import tiktoken
# Path to directory containing vocab.json and merges.txt
model_path = "path/to/custom_bpe"
# Returns a tuple-like structure: (vocab_dict, merge_rules)
mergeable = load_tiktoken_bpe(model_path)
Define Special Tokens and Reserved IDs
Create a dictionary mapping custom token strings to integer IDs. The repository uses IDs above 100,000 to avoid collisions with the base vocabulary. In pkg/llms_from_scratch/llama3.py (lines 340-350), the Llama3Tokenizer constructs this mapping to reserve slots for control tokens.
special = {
"<|bos|>": 100_000,
"<|eos|>": 100_001,
**{f"<|reserved_{i}|>": 100_002 + i for i in range(10)},
}
Instantiate the Encoding Object
Pass the loaded vocabulary and special tokens to the tiktoken.Encoding constructor. Set explicit_vocab to the base vocabulary dictionary and mergeable_ranks to the merge rules loaded previously.
enc = tiktoken.Encoding(
name="custom-gpt2",
pat_str=r"""'(?i:[...])|...""", # Reuse GPT-2 regex pattern or define custom
explicit_vocab=mergeable[0],
mergeable_ranks=mergeable[1],
special_tokens=special,
)
Create a Model-Ready Wrapper Class
The Llama3Tokenizer class in pkg/llms_from_scratch/llama3.py demonstrates the standard wrapper pattern. Store the special token mapping in self.special and implement encode and decode methods that handle the allowed_special parameter (see the encode method implementation around line 361).
class CustomTokenizer:
def __init__(self, encoding: tiktoken.Encoding, special_tokens: dict):
self.tok = encoding
self.special = special_tokens
def encode(self, text: str, allowed_special=None):
"""Encode text, optionally allowing specific special tokens."""
return self.tok.encode(text, allowed_special=allowed_special)
def decode(self, token_ids):
"""Decode token IDs back to string."""
return self.tok.decode(token_ids)
Complete Working Example
Combine the steps above to create a fully functional extended tokenizer. This pattern matches the implementation found in the repository’s extension notebook and llama3.py wrapper.
from tiktoken.load import load_tiktoken_bpe
import tiktoken
# 1. Load base BPE data
model_path = "path/to/custom_bpe"
mergeable = load_tiktoken_bpe(model_path)
# 2. Define extended special tokens
special_tokens = {
"<|bos|>": 100_000,
"<|eos|>": 100_001,
"<|custom|>": 100_002,
}
# 3. Build Encoding with extended vocabulary
encoding = tiktoken.Encoding(
name="custom-llama",
pat_str=r"""'(?i:[...])|...""",
explicit_vocab=mergeable[0],
mergeable_ranks=mergeable[1],
special_tokens=special_tokens,
)
# 4. Create wrapper
tokenizer = CustomTokenizer(encoding, special_tokens)
# 5. Use the extended tokenizer
text = "Hello <|bos|> world <|eos|>"
ids = tokenizer.encode(text, allowed_special={"<|bos|>", "<|eos|>"})
print(ids) # Output includes custom IDs 100000 and 100001
print(tokenizer.decode(ids)) # Reconstructs original text
Integration with Generation Code
When generating text, your model needs to recognize when to emit special tokens. The Llama3Tokenizer wrapper in pkg/llms_from_scratch/llama3.py handles this by checking self.special during the decoding phase. Special tokens inserted during the encode phase (such as <|reserved_i|> markers) are preserved as distinct IDs in the token stream and can trigger specific model behaviors during generation or post-processing.
Summary
- Tiktoken is immutable: You cannot add tokens to an existing
Encoding; you must rebuild it withtiktoken.Encoding. - Use
load_tiktoken_bpe: Load vocabulary and merge rules fromvocab.jsonandmerges.txtfiles to form the base vocabulary. - Reserve high IDs: Assign custom special token IDs above 100,000 to avoid conflicts with the base BPE vocabulary.
- Wrap for convenience: Create a wrapper class (like
Llama3Tokenizerinllama3.py) to managespecial_tokensmappings and provide cleanencode/decodeAPIs. - Reference the notebook: See
ch05/09_extending-tokenizers/extend-tiktoken.ipynbfor the complete step-by-step tutorial with downloadable examples.
Frequently Asked Questions
Can I add tokens to an existing tiktoken encoding without rebuilding it?
No. The tiktoken library does not expose a public add_token or add_special_tokens API. According to the source code analysis in ch05/09_extending-tokenizers/extend-tiktoken.ipynb, the only supported method is to load your base vocabulary, augment it with new token-to-ID mappings, and instantiate a fresh tiktoken.Encoding object.
What ID range should I use for custom special tokens?
Use integer IDs starting at 100,000 or higher. The Llama3Tokenizer implementation in pkg/llms_from_scratch/llama3.py uses this range (100,000+) for special tokens like <|bos|> and <|eos|> to ensure they do not overlap with standard BPE vocabulary IDs, which typically occupy the lower 100,000 integers.
How do I save and reload my extended tokenizer?
The tiktoken.Encoding object itself is stateless regarding the file system, but you can persist your custom vocabulary by saving the augmented vocab.json and merges.txt files. Alternatively, serialize your special_tokens dictionary alongside the base model name. When reloading, simply re-run the load_tiktoken_bpe and Encoding instantiation steps with your saved configuration.
Where can I find a complete working example of extending Tiktoken?
The notebook ch05/09_extending-tokenizers/extend-tiktoken.ipynb in the rasbt/LLMs-from-scratch repository contains a comprehensive tutorial. It demonstrates downloading custom BPE vocabularies, calling load_tiktoken_bpe, building an Encoding with additional special tokens, and serializing the result for reuse. For production patterns, examine pkg/llms_from_scratch/llama3.py to see how the Llama3Tokenizer class wraps the extended encoding for the Llama-3 model architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →