# How to Choose Vocabulary Size for Transformer Models: A Complete Guide

> Learn how to choose vocabulary size for transformer models with this guide. Discover optimal ranges for monolingual and multilingual data using BPE for better NLP performance.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: tutorial
- Published: 2026-05-31

---

**Choose a vocabulary size between 30,000 and 50,000 for monolingual English corpora using Byte-Pair Encoding (BPE), or 100,000 to 200,000 for multilingual datasets, while ensuring the identical value propagates through the tokenizer configuration, model embedding layer, and language modeling head.**

Transformer-based language models rely on fixed-size token indices that directly impact memory consumption, training speed, and sequence representation quality. In the `FareedKhan-dev/train-llm-from-scratch` repository, the vocabulary size is defined centrally in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py) and cascades through the embedding matrices and output layers implemented in [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py), making it a critical hyperparameter to optimize before training begins.

## Why Vocabulary Size Matters

### Memory and Computational Impact

The embedding matrix scales linearly with vocabulary size. In [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py), the `Transformer` class creates `nn.Embedding(vocab_size, n_embed)` for token embeddings and `nn.Linear(n_embed, vocab_size)` for the language modeling head. For a 768-dimensional model with a 50,000 vocabulary, these layers consume approximately 150 MiB of GPU memory in float32 precision. Larger vocabularies increase memory pressure and slow training due to the additional logits computed during each forward pass.

### Token Coverage vs. Sequence Length

Small vocabularies produce more out-of-vocabulary (OOV) tokens and longer sequences after tokenization, forcing the model to process extended token streams for the same text. Large vocabularies capture rare sub-words and shorten sequences but risk overfitting to infrequent tokens and wasting memory on rarely used embeddings.

## Practical Guidelines for Selecting Vocabulary Size

### Dataset Size Heuristics

A rough rule of thumb suggests setting `vocab_size ≈ √(number_of_tokens)` for monolingual corpora. This balances adequate coverage against embedding matrix growth. For a dataset with 2.5 billion tokens, this formula suggests approximately 50,000 vocabulary items.

### Tokenization Method Selection

Your choice of algorithm constrains the effective range:

- **Character-level**: Requires only 100–300 tokens but produces very long sequences.
- **BPE or WordPiece (English)**: Standard implementations use **30,000–50,000** tokens.
- **Multilingual models**: Require **100,000–200,000** tokens to cover diverse scripts and character sets adequately.

### Hardware Constraints

Verify that `vocab_size × n_embed × 4` bytes (for float32) fits within your GPU memory budget alongside other parameters and activations. The embedding table and output weights together double this memory requirement compared to the embedding layer alone.

### Iterative Refinement

Train an initial tokenizer with a conservative size, inspect the frequency distribution of merged tokens, and increment `vocab_size` until perplexity improvements on a validation set plateau. This empirical approach often outperforms theoretical calculations for specialized domains.

## Implementing Vocabulary Size in train-llm-from-scratch

The repository enforces architectural consistency by defining `VOCAB_SIZE` once and propagating it through all components.

### Updating the Configuration

Modify the central constant in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py) before beginning any training run:

```python

# config/config.py

VOCAB_SIZE = 50304  # Adjust this integer before training

```

This value is accessed by training scripts via `config['vocab_size']` and passed to the model initialization.

### Training the Tokenizer

The `sft_rlhf_guide.ipynb` notebook demonstrates training a BPE tokenizer with a custom vocabulary size using the `tokenizers` library:

```python
from tokenizers import Tokenizer, models, trainers, pre_tokenizers

def train_tokenizer(corpus_files, vocab_size, save_path):
    tokenizer = Tokenizer(models.BPE())
    tokenizer.pre_tokenizer = pre_tokenizers.Whitespace()
    trainer = trainers.BpeTrainer(
        vocab_size=vocab_size,
        special_tokens=["<pad>", "<s>", "</s>", "<unk>", "<mask>"]
    )
    tokenizer.train(files=corpus_files, trainer=trainer)
    tokenizer.save(save_path)
    return tokenizer

# Example usage

DEMO_VOCAB_SIZE = 64000
tokenizer = train_tokenizer(
    corpus_files=["data/example.txt"],
    vocab_size=DEMO_VOCAB_SIZE,
    save_path="tokenizer.json"
)

```

### Synchronizing the Model Architecture

The `Transformer` class in [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py) instantiates embeddings and the output head using the configured size:

```python
from src.models.transformer import Transformer
from config import config

model = Transformer(
    n_head=12,
    n_embed=768,
    context_length=1024,
    vocab_size=config['vocab_size'],  # Must match tokenizer training

    N_BLOCKS=12
)

```

Internally, the class constructs the critical layers:
- `self.token_embed = nn.Embedding(vocab_size, n_embed)`
- `self.lm_head = nn.Linear(n_embed, vocab_size)`

### Inference Configuration

Ensure [`scripts/generate_text.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/generate_text.py) uses the identical vocabulary size during inference to prevent index mismatches:

```python

# scripts/generate_text.py

vocab_size = config['vocab_size']

# Load model and tokenizer with matching vocab_size

output = model.generate(input_ids, max_new_tokens=100)

```

## Summary

- **Vocabulary size** determines embedding memory footprint, training throughput, and out-of-vocabulary rates in transformer models.
- **English BPE models** typically use 30,000–50,000 tokens; **multilingual models** require 100,000–200,000 to avoid excessive OOV tokens.
- **Rule of thumb**: Set vocabulary size to approximately the square root of total training tokens for monolingual datasets.
- **Implementation**: Update `VOCAB_SIZE` in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py), retrain the tokenizer using the BPE workflow in `sft_rlhf_guide.ipynb`, and verify that [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py) initializes `nn.Embedding` and `nn.Linear` with matching dimensions.
- **Consistency is mandatory**: All components—the tokenizer, training script, model architecture, and inference pipeline—must reference the identical integer value to prevent runtime dimension errors.

## Frequently Asked Questions

### What happens if my vocabulary size is too small?

A small vocabulary increases out-of-vocabulary tokens and sequence lengths after tokenization, forcing the model to process longer token streams for the same input text. This raises computational costs during self-attention operations and may cause the model to miss morphological nuances, degrading perplexity on diverse corpora as implemented in the repository's `Transformer` class.

### How do I calculate GPU memory requirements for embeddings?

Multiply the vocabulary size by the embedding dimension and 4 bytes for float32 precision. For the 768-dimensional model in [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py), a 50,000-token vocabulary consumes approximately 150 MiB for the embedding matrix, with an equivalent amount required for the `lm_head` output layer, totaling roughly 300 MiB for these two components alone.

### Can I change the vocabulary size after starting training?

No, the vocabulary size fixes the dimensions of `nn.Embedding` and `nn.Linear` layers at model instantiation in [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py). Changing the size requires updating the constant in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py), retraining the tokenizer from scratch using the method shown in `sft_rlhf_guide.ipynb`, and reinitializing the model weights completely.

### Why must the tokenizer and model use identical vocabulary sizes?

Mismatches cause index-out-of-bounds errors when the tokenizer outputs indices that exceed the model's embedding table dimensions. The repository prevents this by centralizing the value in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py), which [`scripts/train_transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_transformer.py) and [`scripts/generate_text.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/generate_text.py) both reference when building the model and loading the tokenizer.