How to Choose Vocabulary Size for Transformer Models: A Complete Guide

Choose a vocabulary size between 30,000 and 50,000 for monolingual English corpora using Byte-Pair Encoding (BPE), or 100,000 to 200,000 for multilingual datasets, while ensuring the identical value propagates through the tokenizer configuration, model embedding layer, and language modeling head.

Transformer-based language models rely on fixed-size token indices that directly impact memory consumption, training speed, and sequence representation quality. In the FareedKhan-dev/train-llm-from-scratch repository, the vocabulary size is defined centrally in config/config.py and cascades through the embedding matrices and output layers implemented in src/models/transformer.py, making it a critical hyperparameter to optimize before training begins.

Why Vocabulary Size Matters

Memory and Computational Impact

The embedding matrix scales linearly with vocabulary size. In src/models/transformer.py, the Transformer class creates nn.Embedding(vocab_size, n_embed) for token embeddings and nn.Linear(n_embed, vocab_size) for the language modeling head. For a 768-dimensional model with a 50,000 vocabulary, these layers consume approximately 150 MiB of GPU memory in float32 precision. Larger vocabularies increase memory pressure and slow training due to the additional logits computed during each forward pass.

Token Coverage vs. Sequence Length

Small vocabularies produce more out-of-vocabulary (OOV) tokens and longer sequences after tokenization, forcing the model to process extended token streams for the same text. Large vocabularies capture rare sub-words and shorten sequences but risk overfitting to infrequent tokens and wasting memory on rarely used embeddings.

Practical Guidelines for Selecting Vocabulary Size

Dataset Size Heuristics

A rough rule of thumb suggests setting vocab_size ≈ √(number_of_tokens) for monolingual corpora. This balances adequate coverage against embedding matrix growth. For a dataset with 2.5 billion tokens, this formula suggests approximately 50,000 vocabulary items.

Tokenization Method Selection

Your choice of algorithm constrains the effective range:

  • Character-level: Requires only 100–300 tokens but produces very long sequences.
  • BPE or WordPiece (English): Standard implementations use 30,000–50,000 tokens.
  • Multilingual models: Require 100,000–200,000 tokens to cover diverse scripts and character sets adequately.

Hardware Constraints

Verify that vocab_size × n_embed × 4 bytes (for float32) fits within your GPU memory budget alongside other parameters and activations. The embedding table and output weights together double this memory requirement compared to the embedding layer alone.

Iterative Refinement

Train an initial tokenizer with a conservative size, inspect the frequency distribution of merged tokens, and increment vocab_size until perplexity improvements on a validation set plateau. This empirical approach often outperforms theoretical calculations for specialized domains.

Implementing Vocabulary Size in train-llm-from-scratch

The repository enforces architectural consistency by defining VOCAB_SIZE once and propagating it through all components.

Updating the Configuration

Modify the central constant in config/config.py before beginning any training run:


# config/config.py

VOCAB_SIZE = 50304  # Adjust this integer before training

This value is accessed by training scripts via config['vocab_size'] and passed to the model initialization.

Training the Tokenizer

The sft_rlhf_guide.ipynb notebook demonstrates training a BPE tokenizer with a custom vocabulary size using the tokenizers library:

from tokenizers import Tokenizer, models, trainers, pre_tokenizers

def train_tokenizer(corpus_files, vocab_size, save_path):
    tokenizer = Tokenizer(models.BPE())
    tokenizer.pre_tokenizer = pre_tokenizers.Whitespace()
    trainer = trainers.BpeTrainer(
        vocab_size=vocab_size,
        special_tokens=["<pad>", "<s>", "</s>", "<unk>", "<mask>"]
    )
    tokenizer.train(files=corpus_files, trainer=trainer)
    tokenizer.save(save_path)
    return tokenizer

# Example usage

DEMO_VOCAB_SIZE = 64000
tokenizer = train_tokenizer(
    corpus_files=["data/example.txt"],
    vocab_size=DEMO_VOCAB_SIZE,
    save_path="tokenizer.json"
)

Synchronizing the Model Architecture

The Transformer class in src/models/transformer.py instantiates embeddings and the output head using the configured size:

from src.models.transformer import Transformer
from config import config

model = Transformer(
    n_head=12,
    n_embed=768,
    context_length=1024,
    vocab_size=config['vocab_size'],  # Must match tokenizer training

    N_BLOCKS=12
)

Internally, the class constructs the critical layers:

  • self.token_embed = nn.Embedding(vocab_size, n_embed)
  • self.lm_head = nn.Linear(n_embed, vocab_size)

Inference Configuration

Ensure scripts/generate_text.py uses the identical vocabulary size during inference to prevent index mismatches:


# scripts/generate_text.py

vocab_size = config['vocab_size']

# Load model and tokenizer with matching vocab_size

output = model.generate(input_ids, max_new_tokens=100)

Summary

  • Vocabulary size determines embedding memory footprint, training throughput, and out-of-vocabulary rates in transformer models.
  • English BPE models typically use 30,000–50,000 tokens; multilingual models require 100,000–200,000 to avoid excessive OOV tokens.
  • Rule of thumb: Set vocabulary size to approximately the square root of total training tokens for monolingual datasets.
  • Implementation: Update VOCAB_SIZE in config/config.py, retrain the tokenizer using the BPE workflow in sft_rlhf_guide.ipynb, and verify that src/models/transformer.py initializes nn.Embedding and nn.Linear with matching dimensions.
  • Consistency is mandatory: All components—the tokenizer, training script, model architecture, and inference pipeline—must reference the identical integer value to prevent runtime dimension errors.

Frequently Asked Questions

What happens if my vocabulary size is too small?

A small vocabulary increases out-of-vocabulary tokens and sequence lengths after tokenization, forcing the model to process longer token streams for the same input text. This raises computational costs during self-attention operations and may cause the model to miss morphological nuances, degrading perplexity on diverse corpora as implemented in the repository's Transformer class.

How do I calculate GPU memory requirements for embeddings?

Multiply the vocabulary size by the embedding dimension and 4 bytes for float32 precision. For the 768-dimensional model in src/models/transformer.py, a 50,000-token vocabulary consumes approximately 150 MiB for the embedding matrix, with an equivalent amount required for the lm_head output layer, totaling roughly 300 MiB for these two components alone.

Can I change the vocabulary size after starting training?

No, the vocabulary size fixes the dimensions of nn.Embedding and nn.Linear layers at model instantiation in src/models/transformer.py. Changing the size requires updating the constant in config/config.py, retraining the tokenizer from scratch using the method shown in sft_rlhf_guide.ipynb, and reinitializing the model weights completely.

Why must the tokenizer and model use identical vocabulary sizes?

Mismatches cause index-out-of-bounds errors when the tokenizer outputs indices that exceed the model's embedding table dimensions. The repository prevents this by centralizing the value in config/config.py, which scripts/train_transformer.py and scripts/generate_text.py both reference when building the model and loading the tokenizer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →