How MegaDLMs Process and Tokenize Input Data for Training: A Complete Technical Guide
MegaDLMs processes training data through a three-stage pipeline that constructs a padded tokenizer, multiprocesses JSON documents into token IDs, and serializes results into indexed binary files for efficient GPU training.
The jinjieni/megadlms repository implements a high-performance data preprocessing pipeline designed for large-scale language model training. Understanding how MegaDLMs process and tokenize input data for training is essential for configuring custom datasets and optimizing distributed training performance across multiple GPUs.
The Three-Stage Tokenization Pipeline
Stage 1: Tokenizer Construction and Vocabulary Padding
The pipeline begins in megatron/training/tokenizer/tokenizer.py with the build_tokenizer function (lines 21-66). This dispatcher reads command-line arguments and instantiates the appropriate concrete tokenizer class—such as _BertWordPieceTokenizer, _GPT2BPETokenizer, _SentencePieceTokenizer, or HuggingFace variants—based on the args.tokenizer_type parameter.
After instantiation, the vocabulary undergoes mandatory padding via _vocab_size_with_padding (lines 99-102). This function expands the vocabulary size to the nearest multiple of make_vocab_size_divisible_by * tensor_model_parallel_size, ensuring that each GPU slice in a model-parallel configuration receives an equal number of embedding rows. This alignment prevents misaligned matrix multiplications during distributed training.
Stage 2: Multiprocessed Document Encoding
Raw text processing occurs in tools/preprocess_data.py through the Encoder class (lines 48-61). The system initializes a multiprocessing pool where each worker executes Encoder.initializer(), which constructs a global tokenizer instance via build_tokenizer to avoid reloading overhead across processes.
The Encoder.encode method (lines 90-107) implements the core tokenization flow:
- JSON Parsing: Loads each line as JSON and extracts fields specified in
--json-keys. - Sentence Segmentation: If a field contains a list, each element becomes a separate sentence; otherwise, the string becomes a single sentence. When
--split-sentencesis enabled, the system uses NLTK Punkt tokenization for further segmentation. - Tokenization: Calls
Encoder.tokenizer.tokenize(sentence)to convert text into lists of token IDs. - EOD Handling: Optionally appends the end-of-document token (
eod) when--append-eodis set, providing deterministic termination markers for the language model. - Length Tracking: Collects per-sentence lengths into
sentence_lensarrays for later indexing.
Stage 3: Binary Dataset Serialization
The final stage converts tokenized documents into high-performance binary storage. The Partition.process_json_file function creates an IndexedDatasetBuilder instance (lines 66-78) for each JSON key specified in the input.
For every processed document, the builder receives:
- The token ID list (
doc[key]) - The per-sentence length array (
sentence_lens[key])
The builder writes two files per key:
.bin: A contiguous stream of rawint32token IDs optimized for memory-mapped access..idx: An index storing offsets and lengths for each document, enabling O(1) random access during training.
This indexed binary format allows the training loop to stream data from disk without loading the entire dataset into RAM, which is critical when training on multi-terabyte web-text corpora across distributed GPU clusters.
Key Implementation Details
Vocabulary Padding for Model Parallelism
The _vocab_size_with_padding function in tokenizer.py ensures that vocab_size % (make_vocab_size_divisible_by * tensor_model_parallel_size) == 0. This mathematical guarantee allows the embedding matrix to be evenly partitioned across model-parallel GPUs, preventing uneven splits that would cause all-reduce mismatches during the forward and backward passes.
Sentence Splitting and EOD Handling
When preprocessing long documents, the optional --split-sentences flag activates NLTK's Punkt sentence tokenizer, breaking documents into shorter passages. Combined with --append-eod, which appends the special eod token (token ID stored in tokenizer.eod), this creates clear document boundaries that help the model learn when to stop generating text during inference.
Practical Code Examples
Tokenizing a Single String from Python
from megatron.training.tokenizer import build_tokenizer
# Minimal args needed for tokenizer construction
class SimpleArgs:
def __init__(self):
self.rank = 0
self.tensor_model_parallel_size = 1
self.make_vocab_size_divisible_by = 128
self.tokenizer_type = "GPT2BPETokenizer"
self.vocab_file = "path/to/gpt2_vocab.json"
self.merge_file = "path/to/gpt2_merges.txt"
self.padded_vocab_size = None
self.append_eod = True
args = SimpleArgs()
tokenizer = build_tokenizer(args)
text = "Hello world! This is Mega DLM."
ids = tokenizer.tokenize(text)
print("Token IDs:", ids)
print("Detokenized:", tokenizer.detokenize(ids))
This snippet mirrors the logic in Encoder.encode and demonstrates the public API of any tokenizer returned by build_tokenizer.
Full Preprocessing Command (CLI)
python -m tools.preprocess_data \
--input /data/web_text.jsonl \
--json-keys text \
--split-sentences \
--keep-newlines \
--append-eod \
--tokenizer_type GPT2BPETokenizer \
--vocab-file /models/gpt2_vocab.json \
--merge-file /models/gpt2_merges.txt \
--output-prefix /preprocessed/web_text
Under the hood, this command executes Partition.process_json_file, which creates an IndexedDatasetBuilder and writes the binary output files (web_text_text_document.bin and web_text_text_document.idx).
Loading the Binary Dataset During Training
from megatron.core.datasets import indexed_dataset
ds = indexed_dataset.make_dataset(
"/preprocessed/web_text_text_document",
# dtype is inferred from the vocab size automatically
)
# Retrieve the first document (list of token ids)
doc_ids = ds[0] # -> torch.Tensor[int]
print("First doc length:", len(doc_ids))
The IndexedDataset class reads the .idx file to locate the correct slice in the .bin file, matching the format produced by the preprocessing step.
Summary
- Tokenizer Construction: The
build_tokenizerfunction inmegatron/training/tokenizer/tokenizer.pyinstantiates concrete tokenizers and pads vocabulary sizes to be divisible by the model-parallel factor, ensuring even GPU distribution. - Multiprocessed Encoding: The
Encoderclass intools/preprocess_data.pyparallelizes JSON parsing, sentence splitting, and tokenization across CPU cores, producing token ID sequences with optional end-of-document markers. - Binary Serialization:
IndexedDatasetBuildercreates memory-mapped.binand.idxfiles for O(1) random access, enabling terabyte-scale datasets to stream efficiently during distributed GPU training.
Frequently Asked Questions
What tokenizer types does MegaDLMs support?
MegaDLMs supports BERT WordPiece, GPT-2 BPE, SentencePiece, HuggingFace tokenizers, and TikTokenizer through the build_tokenizer dispatcher in megatron/training/tokenizer/tokenizer.py. The args.tokenizer_type parameter determines which concrete class is instantiated.
Why does MegaDLMs pad the vocabulary size?
The _vocab_size_with_padding function ensures the vocabulary size is divisible by make_vocab_size_divisible_by * tensor_model_parallel_size. This guarantees that embedding matrices split evenly across model-parallel GPUs, preventing misaligned tensor operations during distributed training.
How does the preprocessing tool handle large JSON files?
The tools/preprocess_data.py script uses a multiprocessing.Pool with an Encoder.initializer that loads the tokenizer once per worker. This architecture avoids reloading the tokenizer for every document, maximizing CPU throughput when processing multi-terabyte corpora.
What is the difference between the .bin and .idx files?
The .bin file contains a contiguous stream of raw int32 token IDs, while the .idx file stores an index of document offsets and lengths. During training, IndexedDataset uses the .idx file to locate and memory-map specific document slices from the .bin file without loading the entire dataset into RAM.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →