Supported Tokenizer Types for MegaDLMs Data Preprocessing: Complete Guide

MegaDLMs supports ten distinct tokenizer types—including BERT WordPiece variants, GPT-2 BPE, SentencePiece, HuggingFace AutoTokenizer wrappers, Llama2, Tiktoken, and multimodal implementations—selectable via the --tokenizer-type argument for data preprocessing pipelines.

The MegaDLMs framework extends Megatron-LM with a flexible tokenizer abstraction that allows runtime swapping of tokenization strategies. In the jinjieni/megadlms repository, preprocessing scripts like tools/preprocess_data.py and the tokenizer factory in megatron/training/tokenizer/tokenizer.py implement a unified interface supporting multiple backends for large-scale language model training.

Where Tokenizer Types Are Declared in MegaDLMs

Command-Line Argument Parser

In megatron/training/arguments.py, the _add_tokenizer_args function defines the --tokenizer-type command-line flag with an explicit choices parameter that enumerates every supported tokenizer name (lines 30-53).

Tokenizer Factory Implementation

The build_tokenizer function in megatron/training/tokenizer/tokenizer.py (lines 21-98) acts as the central factory, using a chain of if ... elif statements to map the args.tokenizer_type string to its corresponding concrete implementation class such as _GPT2BPETokenizer or _HuggingFaceTokenizer.

Retro-Specific Selectors

For the Retro data-preprocessing pipeline, tools/retro/preprocess_data.py provides specialized helper functions. The get_gpt_tokenizer method only supports GPT2BPETokenizer and GPTSentencePieceTokenizer, while get_bert_tokenizer handles the two BERT variants via a dictionary lookup (lines 152-184).

Full List of Supported Tokenizer Types

The following table lists the exact string values accepted by --tokenizer-type, their implementation classes, and typical use cases:

Tokenizer Type Implementation Class Typical Use Case
BertWordPieceLowerCase _BertWordPieceTokenizer Standard BERT with lowercasing (uncased models)
BertWordPieceCase _BertWordPieceTokenizer Cased BERT tokenization
GPT2BPETokenizer _GPT2BPETokenizer Original GPT-2 byte-pair encoding
SentencePieceTokenizer _SentencePieceTokenizer Generic SentencePiece models
GPTSentencePieceTokenizer _GPTSentencePieceTokenizer GPT-style SentencePiece without extra special tokens
HuggingFaceTokenizer _HuggingFaceTokenizer Wrapper for any 🤗 Transformers AutoTokenizer
Llama2Tokenizer _Llama2Tokenizer Llama-2 specific SentencePiece implementation
TikTokenizer CustomTikTokenizer Tiktoken-compatible tokenizer for OpenAI-style models
MultimodalTokenizer MultimodalTokenizer Combines HuggingFace tokenizer with image-tag handling for vision-language models
NullTokenizer _NullTokenizer Identity tokenizer for debugging or synthetic data

Important Constraint: The Retro pipeline restricts GPT tokenizers to GPT2BPETokenizer and GPTSentencePieceTokenizer only; attempting to use other types raises an exception at line 174 of tools/retro/preprocess_data.py.

Code Examples for MegaDLMs Tokenizer Selection

Command-Line Preprocessing

To preprocess data using the GPT-2 BPE tokenizer:

python tools/preprocess_data.py \
    --input data/my_dataset.jsonl \
    --output-prefix data/processed/my_dataset \
    --workers 8 \
    --partitions 1 \
    --tokenizer-type GPT2BPETokenizer \
    --vocab-file /path/to/gpt2-vocab.json \
    --merge-file /path/to/gpt2-merges.txt

The script forwards the --tokenizer-type argument to build_tokenizer, which instantiates _GPT2BPETokenizer (see tokenizer.py lines 37-41).

Programmatic Tokenizer Construction

To load a HuggingFace tokenizer programmatically:

from megatron.training.tokenizer.tokenizer import build_tokenizer
from megatron.training.arguments import _add_tokenizer_args
import argparse

parser = argparse.ArgumentParser()
parser = _add_tokenizer_args(parser)

args = parser.parse_args([
    '--tokenizer-type', 'HuggingFaceTokenizer',
    '--tokenizer-model', 'meta-llama/Llama-2-7b-hf'
])

tokenizer = build_tokenizer(args)
print(f"Tokenizer vocab size: {tokenizer.vocab_size}")
print(f"EOD token id: {tokenizer.eod}")

Retro Preprocessing Pipeline

To select a tokenizer for Retro GPT preprocessing:

from tools.retro.preprocess_data import get_gpt_tokenizer

class RetroConfig:
    retro_gpt_tokenizer_type = "GPTSentencePieceTokenizer"
    retro_gpt_tokenizer_model = "path/to/gpt_spm.model"

cfg = RetroConfig()
tokenizer = get_gpt_tokenizer(cfg)
print(tokenizer.__class__.__name__)  # Output: _GPTSentencePieceTokenizer

Key Source Files

Summary

  • Ten tokenizer types are officially supported in MegaDLMs, covering BERT WordPiece, GPT-2 BPE, SentencePiece variants, HuggingFace AutoTokenizer wrappers, Llama2, Tiktoken, and multimodal implementations.
  • Selection mechanism relies on the --tokenizer-type argument validated in megatron/training/arguments.py and instantiated via build_tokenizer in megatron/training/tokenizer/tokenizer.py.
  • Retro pipeline restrictions limit GPT-style tokenization to GPT2BPETokenizer and GPTSentencePieceTokenizer only.
  • HuggingFace integration enables loading any compatible model through the HuggingFaceTokenizer type without code changes.
  • Debugging support is available via NullTokenizer, which acts as an identity function for synthetic data testing.

Frequently Asked Questions

How do I add a custom tokenizer to MegaDLMs?

Create a new class inheriting from MegatronTokenizer in megatron/training/tokenizer/tokenizer.py, add the corresponding string to the choices list in _add_tokenizer_args within megatron/training/arguments.py, and implement an elif branch in build_tokenizer to instantiate your class when the matching --tokenizer-type is provided. This extends the supported tokenizer types without modifying the preprocessing scripts.

Can I use Llama2Tokenizer with the Retro preprocessing pipeline?

No. According to the source code in tools/retro/preprocess_data.py, the get_gpt_tokenizer function explicitly checks for GPT2BPETokenizer or GPTSentencePieceTokenizer and raises an exception at line 174 for unrecognized types. You must use one of the two supported GPT tokenizers for Retro data preprocessing.

What is the difference between SentencePieceTokenizer and GPTSentencePieceTokenizer?

SentencePieceTokenizer is a generic implementation that loads any SentencePiece model with full special token support, while GPTSentencePieceTokenizer is specifically tuned for autoregressive GPT-style models and excludes extra special tokens that could interfere with training. Both are instantiated in megatron/training/tokenizer/tokenizer.py but serve different architectural requirements.

How does the NullTokenizer work and when should I use it?

NullTokenizer (implemented as _NullTokenizer) is an identity tokenizer that returns input IDs unchanged, useful for debugging preprocessing pipelines, handling pre-tokenized synthetic data, or testing data loading performance without invoking heavy tokenization logic. It is selectable via --tokenizer-type NullTokenizer and defined in the tokenizer factory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →