Supported Tokenizer Types for MegaDLMs Data Preprocessing: Complete Guide
MegaDLMs supports ten distinct tokenizer types—including BERT WordPiece variants, GPT-2 BPE, SentencePiece, HuggingFace AutoTokenizer wrappers, Llama2, Tiktoken, and multimodal implementations—selectable via the --tokenizer-type argument for data preprocessing pipelines.
The MegaDLMs framework extends Megatron-LM with a flexible tokenizer abstraction that allows runtime swapping of tokenization strategies. In the jinjieni/megadlms repository, preprocessing scripts like tools/preprocess_data.py and the tokenizer factory in megatron/training/tokenizer/tokenizer.py implement a unified interface supporting multiple backends for large-scale language model training.
Where Tokenizer Types Are Declared in MegaDLMs
Command-Line Argument Parser
In megatron/training/arguments.py, the _add_tokenizer_args function defines the --tokenizer-type command-line flag with an explicit choices parameter that enumerates every supported tokenizer name (lines 30-53).
Tokenizer Factory Implementation
The build_tokenizer function in megatron/training/tokenizer/tokenizer.py (lines 21-98) acts as the central factory, using a chain of if ... elif statements to map the args.tokenizer_type string to its corresponding concrete implementation class such as _GPT2BPETokenizer or _HuggingFaceTokenizer.
Retro-Specific Selectors
For the Retro data-preprocessing pipeline, tools/retro/preprocess_data.py provides specialized helper functions. The get_gpt_tokenizer method only supports GPT2BPETokenizer and GPTSentencePieceTokenizer, while get_bert_tokenizer handles the two BERT variants via a dictionary lookup (lines 152-184).
Full List of Supported Tokenizer Types
The following table lists the exact string values accepted by --tokenizer-type, their implementation classes, and typical use cases:
| Tokenizer Type | Implementation Class | Typical Use Case |
|---|---|---|
BertWordPieceLowerCase |
_BertWordPieceTokenizer |
Standard BERT with lowercasing (uncased models) |
BertWordPieceCase |
_BertWordPieceTokenizer |
Cased BERT tokenization |
GPT2BPETokenizer |
_GPT2BPETokenizer |
Original GPT-2 byte-pair encoding |
SentencePieceTokenizer |
_SentencePieceTokenizer |
Generic SentencePiece models |
GPTSentencePieceTokenizer |
_GPTSentencePieceTokenizer |
GPT-style SentencePiece without extra special tokens |
HuggingFaceTokenizer |
_HuggingFaceTokenizer |
Wrapper for any 🤗 Transformers AutoTokenizer |
Llama2Tokenizer |
_Llama2Tokenizer |
Llama-2 specific SentencePiece implementation |
TikTokenizer |
CustomTikTokenizer |
Tiktoken-compatible tokenizer for OpenAI-style models |
MultimodalTokenizer |
MultimodalTokenizer |
Combines HuggingFace tokenizer with image-tag handling for vision-language models |
NullTokenizer |
_NullTokenizer |
Identity tokenizer for debugging or synthetic data |
Important Constraint: The Retro pipeline restricts GPT tokenizers to GPT2BPETokenizer and GPTSentencePieceTokenizer only; attempting to use other types raises an exception at line 174 of tools/retro/preprocess_data.py.
Code Examples for MegaDLMs Tokenizer Selection
Command-Line Preprocessing
To preprocess data using the GPT-2 BPE tokenizer:
python tools/preprocess_data.py \
--input data/my_dataset.jsonl \
--output-prefix data/processed/my_dataset \
--workers 8 \
--partitions 1 \
--tokenizer-type GPT2BPETokenizer \
--vocab-file /path/to/gpt2-vocab.json \
--merge-file /path/to/gpt2-merges.txt
The script forwards the --tokenizer-type argument to build_tokenizer, which instantiates _GPT2BPETokenizer (see tokenizer.py lines 37-41).
Programmatic Tokenizer Construction
To load a HuggingFace tokenizer programmatically:
from megatron.training.tokenizer.tokenizer import build_tokenizer
from megatron.training.arguments import _add_tokenizer_args
import argparse
parser = argparse.ArgumentParser()
parser = _add_tokenizer_args(parser)
args = parser.parse_args([
'--tokenizer-type', 'HuggingFaceTokenizer',
'--tokenizer-model', 'meta-llama/Llama-2-7b-hf'
])
tokenizer = build_tokenizer(args)
print(f"Tokenizer vocab size: {tokenizer.vocab_size}")
print(f"EOD token id: {tokenizer.eod}")
Retro Preprocessing Pipeline
To select a tokenizer for Retro GPT preprocessing:
from tools.retro.preprocess_data import get_gpt_tokenizer
class RetroConfig:
retro_gpt_tokenizer_type = "GPTSentencePieceTokenizer"
retro_gpt_tokenizer_model = "path/to/gpt_spm.model"
cfg = RetroConfig()
tokenizer = get_gpt_tokenizer(cfg)
print(tokenizer.__class__.__name__) # Output: _GPTSentencePieceTokenizer
Key Source Files
megatron/training/tokenizer/tokenizer.py: Contains thebuild_tokenizerfactory and all concrete tokenizer classes.megatron/training/arguments.py: Defines the--tokenizer-typeCLI flag and valid choices (lines 30-53).tools/retro/preprocess_data.py: Implements Retro-specific selectorsget_gpt_tokenizerandget_bert_tokenizer.tools/preprocess_data.py: End-to-end preprocessing pipeline that invokes the tokenizer factory.tests/unit_tests/test_tokenizer.py: Unit tests demonstrating instantiation of each supported tokenizer.
Summary
- Ten tokenizer types are officially supported in MegaDLMs, covering BERT WordPiece, GPT-2 BPE, SentencePiece variants, HuggingFace AutoTokenizer wrappers, Llama2, Tiktoken, and multimodal implementations.
- Selection mechanism relies on the
--tokenizer-typeargument validated inmegatron/training/arguments.pyand instantiated viabuild_tokenizerinmegatron/training/tokenizer/tokenizer.py. - Retro pipeline restrictions limit GPT-style tokenization to
GPT2BPETokenizerandGPTSentencePieceTokenizeronly. - HuggingFace integration enables loading any compatible model through the
HuggingFaceTokenizertype without code changes. - Debugging support is available via
NullTokenizer, which acts as an identity function for synthetic data testing.
Frequently Asked Questions
How do I add a custom tokenizer to MegaDLMs?
Create a new class inheriting from MegatronTokenizer in megatron/training/tokenizer/tokenizer.py, add the corresponding string to the choices list in _add_tokenizer_args within megatron/training/arguments.py, and implement an elif branch in build_tokenizer to instantiate your class when the matching --tokenizer-type is provided. This extends the supported tokenizer types without modifying the preprocessing scripts.
Can I use Llama2Tokenizer with the Retro preprocessing pipeline?
No. According to the source code in tools/retro/preprocess_data.py, the get_gpt_tokenizer function explicitly checks for GPT2BPETokenizer or GPTSentencePieceTokenizer and raises an exception at line 174 for unrecognized types. You must use one of the two supported GPT tokenizers for Retro data preprocessing.
What is the difference between SentencePieceTokenizer and GPTSentencePieceTokenizer?
SentencePieceTokenizer is a generic implementation that loads any SentencePiece model with full special token support, while GPTSentencePieceTokenizer is specifically tuned for autoregressive GPT-style models and excludes extra special tokens that could interfere with training. Both are instantiated in megatron/training/tokenizer/tokenizer.py but serve different architectural requirements.
How does the NullTokenizer work and when should I use it?
NullTokenizer (implemented as _NullTokenizer) is an identity tokenizer that returns input IDs unchanged, useful for debugging preprocessing pipelines, handling pre-tokenized synthetic data, or testing data loading performance without invoking heavy tokenization logic. It is selectable via --tokenizer-type NullTokenizer and defined in the tokenizer factory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →