# How MegaDLMs Process and Tokenize Input Data for Training: A Complete Technical Guide

> Discover how MegaDLMs processes and tokenizes input data with a three-stage pipeline. Learn about padded tokenizers, multiprocessed JSON to token IDs, and indexed binary files for efficient GPU training.

- Repository: [Jinjie Ni/megadlms](https://github.com/jinjieni/megadlms)
- Tags: deep-dive
- Published: 2026-03-04

---

**MegaDLMs processes training data through a three-stage pipeline that constructs a padded tokenizer, multiprocesses JSON documents into token IDs, and serializes results into indexed binary files for efficient GPU training.**

The jinjieni/megadlms repository implements a high-performance data preprocessing pipeline designed for large-scale language model training. Understanding how MegaDLMs process and tokenize input data for training is essential for configuring custom datasets and optimizing distributed training performance across multiple GPUs.

## The Three-Stage Tokenization Pipeline

### Stage 1: Tokenizer Construction and Vocabulary Padding

The pipeline begins in [`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py) with the `build_tokenizer` function (lines 21-66). This dispatcher reads command-line arguments and instantiates the appropriate concrete tokenizer class—such as `_BertWordPieceTokenizer`, `_GPT2BPETokenizer`, `_SentencePieceTokenizer`, or HuggingFace variants—based on the `args.tokenizer_type` parameter.

After instantiation, the vocabulary undergoes mandatory padding via `_vocab_size_with_padding` (lines 99-102). This function expands the vocabulary size to the nearest multiple of `make_vocab_size_divisible_by * tensor_model_parallel_size`, ensuring that each GPU slice in a model-parallel configuration receives an equal number of embedding rows. This alignment prevents misaligned matrix multiplications during distributed training.

### Stage 2: Multiprocessed Document Encoding

Raw text processing occurs in [`tools/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/preprocess_data.py) through the `Encoder` class (lines 48-61). The system initializes a multiprocessing pool where each worker executes `Encoder.initializer()`, which constructs a global tokenizer instance via `build_tokenizer` to avoid reloading overhead across processes.

The `Encoder.encode` method (lines 90-107) implements the core tokenization flow:

1. **JSON Parsing**: Loads each line as JSON and extracts fields specified in `--json-keys`.
2. **Sentence Segmentation**: If a field contains a list, each element becomes a separate sentence; otherwise, the string becomes a single sentence. When `--split-sentences` is enabled, the system uses NLTK Punkt tokenization for further segmentation.
3. **Tokenization**: Calls `Encoder.tokenizer.tokenize(sentence)` to convert text into lists of token IDs.
4. **EOD Handling**: Optionally appends the end-of-document token (`eod`) when `--append-eod` is set, providing deterministic termination markers for the language model.
5. **Length Tracking**: Collects per-sentence lengths into `sentence_lens` arrays for later indexing.

### Stage 3: Binary Dataset Serialization

The final stage converts tokenized documents into high-performance binary storage. The `Partition.process_json_file` function creates an `IndexedDatasetBuilder` instance (lines 66-78) for each JSON key specified in the input.

For every processed document, the builder receives:
- The token ID list (`doc[key]`)
- The per-sentence length array (`sentence_lens[key]`)

The builder writes two files per key:
- **`.bin`**: A contiguous stream of raw `int32` token IDs optimized for memory-mapped access.
- **`.idx`**: An index storing offsets and lengths for each document, enabling O(1) random access during training.

This indexed binary format allows the training loop to stream data from disk without loading the entire dataset into RAM, which is critical when training on multi-terabyte web-text corpora across distributed GPU clusters.

## Key Implementation Details

### Vocabulary Padding for Model Parallelism

The `_vocab_size_with_padding` function in [`tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/tokenizer.py) ensures that `vocab_size % (make_vocab_size_divisible_by * tensor_model_parallel_size) == 0`. This mathematical guarantee allows the embedding matrix to be evenly partitioned across model-parallel GPUs, preventing uneven splits that would cause all-reduce mismatches during the forward and backward passes.

### Sentence Splitting and EOD Handling

When preprocessing long documents, the optional `--split-sentences` flag activates NLTK's Punkt sentence tokenizer, breaking documents into shorter passages. Combined with `--append-eod`, which appends the special `eod` token (token ID stored in `tokenizer.eod`), this creates clear document boundaries that help the model learn when to stop generating text during inference.

## Practical Code Examples

### Tokenizing a Single String from Python

```python
from megatron.training.tokenizer import build_tokenizer

# Minimal args needed for tokenizer construction

class SimpleArgs:
    def __init__(self):
        self.rank = 0
        self.tensor_model_parallel_size = 1
        self.make_vocab_size_divisible_by = 128
        self.tokenizer_type = "GPT2BPETokenizer"
        self.vocab_file = "path/to/gpt2_vocab.json"
        self.merge_file = "path/to/gpt2_merges.txt"
        self.padded_vocab_size = None
        self.append_eod = True

args = SimpleArgs()
tokenizer = build_tokenizer(args)

text = "Hello world! This is Mega DLM."
ids = tokenizer.tokenize(text)
print("Token IDs:", ids)
print("Detokenized:", tokenizer.detokenize(ids))

```

This snippet mirrors the logic in `Encoder.encode` and demonstrates the public API of any tokenizer returned by `build_tokenizer`.

### Full Preprocessing Command (CLI)

```bash
python -m tools.preprocess_data \
    --input /data/web_text.jsonl \
    --json-keys text \
    --split-sentences \
    --keep-newlines \
    --append-eod \
    --tokenizer_type GPT2BPETokenizer \
    --vocab-file /models/gpt2_vocab.json \
    --merge-file /models/gpt2_merges.txt \
    --output-prefix /preprocessed/web_text

```

Under the hood, this command executes `Partition.process_json_file`, which creates an `IndexedDatasetBuilder` and writes the binary output files (`web_text_text_document.bin` and `web_text_text_document.idx`).

### Loading the Binary Dataset During Training

```python
from megatron.core.datasets import indexed_dataset

ds = indexed_dataset.make_dataset(
    "/preprocessed/web_text_text_document",
    # dtype is inferred from the vocab size automatically

)

# Retrieve the first document (list of token ids)

doc_ids = ds[0]          # -> torch.Tensor[int]

print("First doc length:", len(doc_ids))

```

The `IndexedDataset` class reads the `.idx` file to locate the correct slice in the `.bin` file, matching the format produced by the preprocessing step.

## Summary

- **Tokenizer Construction**: The `build_tokenizer` function in [`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py) instantiates concrete tokenizers and pads vocabulary sizes to be divisible by the model-parallel factor, ensuring even GPU distribution.
- **Multiprocessed Encoding**: The `Encoder` class in [`tools/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/preprocess_data.py) parallelizes JSON parsing, sentence splitting, and tokenization across CPU cores, producing token ID sequences with optional end-of-document markers.
- **Binary Serialization**: `IndexedDatasetBuilder` creates memory-mapped `.bin` and `.idx` files for O(1) random access, enabling terabyte-scale datasets to stream efficiently during distributed GPU training.

## Frequently Asked Questions

### What tokenizer types does MegaDLMs support?

MegaDLMs supports BERT WordPiece, GPT-2 BPE, SentencePiece, HuggingFace tokenizers, and TikTokenizer through the `build_tokenizer` dispatcher in [`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py). The `args.tokenizer_type` parameter determines which concrete class is instantiated.

### Why does MegaDLMs pad the vocabulary size?

The `_vocab_size_with_padding` function ensures the vocabulary size is divisible by `make_vocab_size_divisible_by * tensor_model_parallel_size`. This guarantees that embedding matrices split evenly across model-parallel GPUs, preventing misaligned tensor operations during distributed training.

### How does the preprocessing tool handle large JSON files?

The [`tools/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/preprocess_data.py) script uses a `multiprocessing.Pool` with an `Encoder.initializer` that loads the tokenizer once per worker. This architecture avoids reloading the tokenizer for every document, maximizing CPU throughput when processing multi-terabyte corpora.

### What is the difference between the .bin and .idx files?

The `.bin` file contains a contiguous stream of raw `int32` token IDs, while the `.idx` file stores an index of document offsets and lengths. During training, `IndexedDataset` uses the `.idx` file to locate and memory-map specific document slices from the `.bin` file without loading the entire dataset into RAM.