# Supported Tokenizer Types for MegaDLMs Data Preprocessing: Complete Guide

> Explore ten supported tokenizer types for MegaDLMs data preprocessing, including BERT, GPT-2, SentencePiece, Llama2 & multimodal options. Select your ideal tokenizer type with --tokenizer-type.

- Repository: [Jinjie Ni/megadlms](https://github.com/jinjieni/megadlms)
- Tags: tutorial
- Published: 2026-03-04

---

**MegaDLMs supports ten distinct tokenizer types—including BERT WordPiece variants, GPT-2 BPE, SentencePiece, HuggingFace AutoTokenizer wrappers, Llama2, Tiktoken, and multimodal implementations—selectable via the `--tokenizer-type` argument for data preprocessing pipelines.**

The MegaDLMs framework extends Megatron-LM with a flexible tokenizer abstraction that allows runtime swapping of tokenization strategies. In the `jinjieni/megadlms` repository, preprocessing scripts like [`tools/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/preprocess_data.py) and the tokenizer factory in [`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py) implement a unified interface supporting multiple backends for large-scale language model training.

## Where Tokenizer Types Are Declared in MegaDLMs

### Command-Line Argument Parser

In [`megatron/training/arguments.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/arguments.py), the `_add_tokenizer_args` function defines the `--tokenizer-type` command-line flag with an explicit `choices` parameter that enumerates every supported tokenizer name (lines 30-53).

### Tokenizer Factory Implementation

The `build_tokenizer` function in [`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py) (lines 21-98) acts as the central factory, using a chain of `if ... elif` statements to map the `args.tokenizer_type` string to its corresponding concrete implementation class such as `_GPT2BPETokenizer` or `_HuggingFaceTokenizer`.

### Retro-Specific Selectors

For the Retro data-preprocessing pipeline, [`tools/retro/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/retro/preprocess_data.py) provides specialized helper functions. The `get_gpt_tokenizer` method only supports `GPT2BPETokenizer` and `GPTSentencePieceTokenizer`, while `get_bert_tokenizer` handles the two BERT variants via a dictionary lookup (lines 152-184).

## Full List of Supported Tokenizer Types

The following table lists the exact string values accepted by `--tokenizer-type`, their implementation classes, and typical use cases:

| Tokenizer Type | Implementation Class | Typical Use Case |
|---------------|---------------------|------------------|
| `BertWordPieceLowerCase` | `_BertWordPieceTokenizer` | Standard BERT with lowercasing (uncased models) |
| `BertWordPieceCase` | `_BertWordPieceTokenizer` | Cased BERT tokenization |
| `GPT2BPETokenizer` | `_GPT2BPETokenizer` | Original GPT-2 byte-pair encoding |
| `SentencePieceTokenizer` | `_SentencePieceTokenizer` | Generic SentencePiece models |
| `GPTSentencePieceTokenizer` | `_GPTSentencePieceTokenizer` | GPT-style SentencePiece without extra special tokens |
| `HuggingFaceTokenizer` | `_HuggingFaceTokenizer` | Wrapper for any 🤗 Transformers `AutoTokenizer` |
| `Llama2Tokenizer` | `_Llama2Tokenizer` | Llama-2 specific SentencePiece implementation |
| `TikTokenizer` | `CustomTikTokenizer` | Tiktoken-compatible tokenizer for OpenAI-style models |
| `MultimodalTokenizer` | `MultimodalTokenizer` | Combines HuggingFace tokenizer with image-tag handling for vision-language models |
| `NullTokenizer` | `_NullTokenizer` | Identity tokenizer for debugging or synthetic data |

**Important Constraint:** The Retro pipeline restricts GPT tokenizers to `GPT2BPETokenizer` and `GPTSentencePieceTokenizer` only; attempting to use other types raises an exception at line 174 of [`tools/retro/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/retro/preprocess_data.py).

## Code Examples for MegaDLMs Tokenizer Selection

### Command-Line Preprocessing

To preprocess data using the GPT-2 BPE tokenizer:

```bash
python tools/preprocess_data.py \
    --input data/my_dataset.jsonl \
    --output-prefix data/processed/my_dataset \
    --workers 8 \
    --partitions 1 \
    --tokenizer-type GPT2BPETokenizer \
    --vocab-file /path/to/gpt2-vocab.json \
    --merge-file /path/to/gpt2-merges.txt

```

The script forwards the `--tokenizer-type` argument to `build_tokenizer`, which instantiates `_GPT2BPETokenizer` (see [`tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/tokenizer.py) lines 37-41).

### Programmatic Tokenizer Construction

To load a HuggingFace tokenizer programmatically:

```python
from megatron.training.tokenizer.tokenizer import build_tokenizer
from megatron.training.arguments import _add_tokenizer_args
import argparse

parser = argparse.ArgumentParser()
parser = _add_tokenizer_args(parser)

args = parser.parse_args([
    '--tokenizer-type', 'HuggingFaceTokenizer',
    '--tokenizer-model', 'meta-llama/Llama-2-7b-hf'
])

tokenizer = build_tokenizer(args)
print(f"Tokenizer vocab size: {tokenizer.vocab_size}")
print(f"EOD token id: {tokenizer.eod}")

```

### Retro Preprocessing Pipeline

To select a tokenizer for Retro GPT preprocessing:

```python
from tools.retro.preprocess_data import get_gpt_tokenizer

class RetroConfig:
    retro_gpt_tokenizer_type = "GPTSentencePieceTokenizer"
    retro_gpt_tokenizer_model = "path/to/gpt_spm.model"

cfg = RetroConfig()
tokenizer = get_gpt_tokenizer(cfg)
print(tokenizer.__class__.__name__)  # Output: _GPTSentencePieceTokenizer

```

## Key Source Files

- **[`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py)**: Contains the `build_tokenizer` factory and all concrete tokenizer classes.
- **[`megatron/training/arguments.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/arguments.py)**: Defines the `--tokenizer-type` CLI flag and valid choices (lines 30-53).
- **[`tools/retro/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/retro/preprocess_data.py)**: Implements Retro-specific selectors `get_gpt_tokenizer` and `get_bert_tokenizer`.
- **[`tools/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/preprocess_data.py)**: End-to-end preprocessing pipeline that invokes the tokenizer factory.
- **[`tests/unit_tests/test_tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/tests/unit_tests/test_tokenizer.py)**: Unit tests demonstrating instantiation of each supported tokenizer.

## Summary

- **Ten tokenizer types** are officially supported in MegaDLMs, covering BERT WordPiece, GPT-2 BPE, SentencePiece variants, HuggingFace AutoTokenizer wrappers, Llama2, Tiktoken, and multimodal implementations.
- **Selection mechanism** relies on the `--tokenizer-type` argument validated in [`megatron/training/arguments.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/arguments.py) and instantiated via `build_tokenizer` in [`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py).
- **Retro pipeline restrictions** limit GPT-style tokenization to `GPT2BPETokenizer` and `GPTSentencePieceTokenizer` only.
- **HuggingFace integration** enables loading any compatible model through the `HuggingFaceTokenizer` type without code changes.
- **Debugging support** is available via `NullTokenizer`, which acts as an identity function for synthetic data testing.

## Frequently Asked Questions

### How do I add a custom tokenizer to MegaDLMs?

Create a new class inheriting from `MegatronTokenizer` in [`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py), add the corresponding string to the `choices` list in `_add_tokenizer_args` within [`megatron/training/arguments.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/arguments.py), and implement an `elif` branch in `build_tokenizer` to instantiate your class when the matching `--tokenizer-type` is provided. This extends the supported tokenizer types without modifying the preprocessing scripts.

### Can I use Llama2Tokenizer with the Retro preprocessing pipeline?

No. According to the source code in [`tools/retro/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/retro/preprocess_data.py), the `get_gpt_tokenizer` function explicitly checks for `GPT2BPETokenizer` or `GPTSentencePieceTokenizer` and raises an exception at line 174 for unrecognized types. You must use one of the two supported GPT tokenizers for Retro data preprocessing.

### What is the difference between SentencePieceTokenizer and GPTSentencePieceTokenizer?

`SentencePieceTokenizer` is a generic implementation that loads any SentencePiece model with full special token support, while `GPTSentencePieceTokenizer` is specifically tuned for autoregressive GPT-style models and excludes extra special tokens that could interfere with training. Both are instantiated in [`megatron/training/tokenizer/tokenizer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/tokenizer/tokenizer.py) but serve different architectural requirements.

### How does the NullTokenizer work and when should I use it?

`NullTokenizer` (implemented as `_NullTokenizer`) is an identity tokenizer that returns input IDs unchanged, useful for debugging preprocessing pipelines, handling pre-tokenized synthetic data, or testing data loading performance without invoking heavy tokenization logic. It is selectable via `--tokenizer-type NullTokenizer` and defined in the tokenizer factory.