# How to Extend the Tiktoken BPE Tokenizer with New Tokens: A Complete Guide

> Learn to extend the Tiktoken BPE tokenizer with new tokens. Load custom vocabularies, augment special tokens, and wrap in a helper class for your language model. Complete guide.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-12

---

**To extend the Tiktoken BPE tokenizer with new tokens, load a custom vocabulary using `tiktoken.load.load_tiktoken_bpe`, instantiate a new `Encoding` object with augmented special tokens, and wrap it in a helper class that exposes the custom IDs to your language model.**

The **tiktoken** library provides OpenAI’s high-performance byte-pair encoding (BPE) implementation, but it does not expose a public API for dynamically adding tokens to an existing encoding. In the **rasbt/LLMs-from-scratch** repository, developers extend the tokenizer by rebuilding the `Encoding` object from scratch with an augmented vocabulary and merge table, then wrapping it in a model-specific class like `Llama3Tokenizer`.

## Why Tiktoken Requires Rebuilding

Unlike the Hugging Face tokenizers library, **tiktoken** does not support an `add_tokens` method on existing encodings. As demonstrated in `ch05/09_extending-tokenizers/extend-tiktoken.ipynb`, the library’s design treats tokenizers as immutable objects. To add custom tokens, you must re-instantiate `tiktoken.Encoding` with a new vocabulary dictionary that includes your extra tokens mapped to unique integer IDs.

## Step-by-Step Implementation

### Load Custom BPE Vocabulary and Merge Rules

Start by loading the base vocabulary and merge rules using `load_tiktoken_bpe` from the `tiktoken.load` module. This function expects a directory containing [`vocab.json`](https://github.com/rasbt/LLMs-from-scratch/blob/main/vocab.json) and [`merges.txt`](https://github.com/rasbt/LLMs-from-scratch/blob/main/merges.txt) files.

```python
from tiktoken.load import load_tiktoken_bpe
import tiktoken

# Path to directory containing vocab.json and merges.txt

model_path = "path/to/custom_bpe"

# Returns a tuple-like structure: (vocab_dict, merge_rules)

mergeable = load_tiktoken_bpe(model_path)

```

### Define Special Tokens and Reserved IDs

Create a dictionary mapping custom token strings to integer IDs. The repository uses IDs above 100,000 to avoid collisions with the base vocabulary. In [`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py) (lines 340-350), the `Llama3Tokenizer` constructs this mapping to reserve slots for control tokens.

```python
special = {
    "<|bos|>":  100_000,
    "<|eos|>":  100_001,
    **{f"<|reserved_{i}|>": 100_002 + i for i in range(10)},
}

```

### Instantiate the Encoding Object

Pass the loaded vocabulary and special tokens to the `tiktoken.Encoding` constructor. Set `explicit_vocab` to the base vocabulary dictionary and `mergeable_ranks` to the merge rules loaded previously.

```python
enc = tiktoken.Encoding(
    name="custom-gpt2",
    pat_str=r"""'(?i:[...])|...""",      # Reuse GPT-2 regex pattern or define custom

    explicit_vocab=mergeable[0],
    mergeable_ranks=mergeable[1],
    special_tokens=special,
)

```

### Create a Model-Ready Wrapper Class

The `Llama3Tokenizer` class in [`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py) demonstrates the standard wrapper pattern. Store the special token mapping in `self.special` and implement `encode` and `decode` methods that handle the `allowed_special` parameter (see the `encode` method implementation around line 361).

```python
class CustomTokenizer:
    def __init__(self, encoding: tiktoken.Encoding, special_tokens: dict):
        self.tok = encoding
        self.special = special_tokens

    def encode(self, text: str, allowed_special=None):
        """Encode text, optionally allowing specific special tokens."""
        return self.tok.encode(text, allowed_special=allowed_special)

    def decode(self, token_ids):
        """Decode token IDs back to string."""
        return self.tok.decode(token_ids)

```

## Complete Working Example

Combine the steps above to create a fully functional extended tokenizer. This pattern matches the implementation found in the repository’s extension notebook and [`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py) wrapper.

```python
from tiktoken.load import load_tiktoken_bpe
import tiktoken

# 1. Load base BPE data

model_path = "path/to/custom_bpe"
mergeable = load_tiktoken_bpe(model_path)

# 2. Define extended special tokens

special_tokens = {
    "<|bos|>": 100_000,
    "<|eos|>": 100_001,
    "<|custom|>": 100_002,
}

# 3. Build Encoding with extended vocabulary

encoding = tiktoken.Encoding(
    name="custom-llama",
    pat_str=r"""'(?i:[...])|...""",
    explicit_vocab=mergeable[0],
    mergeable_ranks=mergeable[1],
    special_tokens=special_tokens,
)

# 4. Create wrapper

tokenizer = CustomTokenizer(encoding, special_tokens)

# 5. Use the extended tokenizer

text = "Hello <|bos|> world <|eos|>"
ids = tokenizer.encode(text, allowed_special={"<|bos|>", "<|eos|>"})
print(ids)  # Output includes custom IDs 100000 and 100001

print(tokenizer.decode(ids))  # Reconstructs original text

```

## Integration with Generation Code

When generating text, your model needs to recognize when to emit special tokens. The `Llama3Tokenizer` wrapper in [`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py) handles this by checking `self.special` during the decoding phase. Special tokens inserted during the `encode` phase (such as `<|reserved_i|>` markers) are preserved as distinct IDs in the token stream and can trigger specific model behaviors during generation or post-processing.

## Summary

- **Tiktoken is immutable**: You cannot add tokens to an existing `Encoding`; you must rebuild it with `tiktoken.Encoding`.
- **Use `load_tiktoken_bpe`**: Load vocabulary and merge rules from [`vocab.json`](https://github.com/rasbt/LLMs-from-scratch/blob/main/vocab.json) and [`merges.txt`](https://github.com/rasbt/LLMs-from-scratch/blob/main/merges.txt) files to form the base vocabulary.
- **Reserve high IDs**: Assign custom special token IDs above 100,000 to avoid conflicts with the base BPE vocabulary.
- **Wrap for convenience**: Create a wrapper class (like `Llama3Tokenizer` in [`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py)) to manage `special_tokens` mappings and provide clean `encode`/`decode` APIs.
- **Reference the notebook**: See `ch05/09_extending-tokenizers/extend-tiktoken.ipynb` for the complete step-by-step tutorial with downloadable examples.

## Frequently Asked Questions

### Can I add tokens to an existing tiktoken encoding without rebuilding it?

No. The **tiktoken** library does not expose a public `add_token` or `add_special_tokens` API. According to the source code analysis in `ch05/09_extending-tokenizers/extend-tiktoken.ipynb`, the only supported method is to load your base vocabulary, augment it with new token-to-ID mappings, and instantiate a fresh `tiktoken.Encoding` object.

### What ID range should I use for custom special tokens?

Use integer IDs starting at 100,000 or higher. The `Llama3Tokenizer` implementation in [`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py) uses this range (100,000+) for special tokens like `<|bos|>` and `<|eos|>` to ensure they do not overlap with standard BPE vocabulary IDs, which typically occupy the lower 100,000 integers.

### How do I save and reload my extended tokenizer?

The `tiktoken.Encoding` object itself is stateless regarding the file system, but you can persist your custom vocabulary by saving the augmented [`vocab.json`](https://github.com/rasbt/LLMs-from-scratch/blob/main/vocab.json) and [`merges.txt`](https://github.com/rasbt/LLMs-from-scratch/blob/main/merges.txt) files. Alternatively, serialize your `special_tokens` dictionary alongside the base model name. When reloading, simply re-run the `load_tiktoken_bpe` and `Encoding` instantiation steps with your saved configuration.

### Where can I find a complete working example of extending Tiktoken?

The notebook **`ch05/09_extending-tokenizers/extend-tiktoken.ipynb`** in the rasbt/LLMs-from-scratch repository contains a comprehensive tutorial. It demonstrates downloading custom BPE vocabularies, calling `load_tiktoken_bpe`, building an `Encoding` with additional special tokens, and serializing the result for reuse. For production patterns, examine [`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py) to see how the `Llama3Tokenizer` class wraps the extended encoding for the Llama-3 model architecture.