# How to Process and Tokenize Data for LLM Training: A Complete Pipeline Guide

> Learn to process and tokenize data for LLM training with a complete pipeline guide. Convert raw text to efficient HDF5 token files using tiktoken and specialized loss masks.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: tutorial
- Published: 2026-06-11

---

**The train-llm-from-scratch repository implements a three-stage pipeline that converts raw text corpora into efficient HDF5 token files using tiktoken, with specialized loss masks for pre-training, supervised fine-tuning, and reinforcement learning.**

Processing and tokenizing data for LLM training requires distinct strategies for each stage of the model lifecycle. In the `FareedKhan-dev/train-llm-from-scratch` repository, the data preparation pipeline handles everything from streaming Pile dataset shards to packing instruction-following examples with per-token loss masks. This implementation uses **tiktoken** with the `r50k_base` encoder and stores all token streams in chunked HDF5 format for efficient random access during training.

## Pre-training Data Preparation

The pre-training stage converts raw Pile shards into a single flat-token HDF5 file that the model can read efficiently. This process is orchestrated by [`scripts/prepare_pretrain_data.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/prepare_pretrain_data.py).

### Downloading and Streaming Pile Shards

The script downloads shards from HuggingFace in `.jsonl.zst` format and stream-decompresses each shard to minimize memory overhead. Rather than holding entire shards in memory, the pipeline processes text in batches, encoding paragraphs as they arrive.

### Tokenization and Special Tokens

Each text batch is encoded using **tiktoken** with the `r50k_base` encoder. The only true special token used is the end-of-text token (`<|endoftext|>`), which maps to `EOT_ID = 50256`. This token is appended to every document and serves as the generation stop token. The implementation uses `tiktoken.encode_ordinary_batch` for efficient batch processing.

### Efficient HDF5 Storage

Tokens are written to HDF5 in chunks to avoid repeated dataset resizing. The implementation uses a write chunk size of `8_000_000` tokens:

```python
with h5py.File(out_path, "w") as f:
    dset = f.create_dataset("tokens", (0,), maxshape=(None,), dtype="i4", chunks=(8_000_000,))

```

This approach appends tokens incrementally and stores files under `/ephemeral/data` by default, leveraging large temporary disks for high-throughput I/O.

## Supervised Fine-Tuning (SFT) Data Preparation

The SFT pipeline in [`scripts/prepare_sft_data.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/prepare_sft_data.py) builds training data from public instruction datasets including Alpaca, Dolly, and GSM8K. Unlike pre-training, this stage requires per-token loss masks to distinguish between user prompts (loss = 0) and assistant responses (loss = 1).

### Chat Format and Role Markers

The repository converts each example to a chat format with `role` and `content` fields. Role markers (`<|user|>`, `<|assistant|>`, `<|system|>`) are plain-text strings that are tokenized like any other text; the model learns to treat them as structural cues during training. These markers are defined in [`src/post_training/chat_template.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/chat_template.py).

### Loss Mask Generation with encode_chat

The core function `encode_chat(messages, add_generation_prompt=False)` returns two aligned arrays: token `ids` and a binary `loss_mask`. Only assistant tokens and the trailing `EOT_ID` are marked with `mask == 1`, ensuring the model learns only to predict the response:

```python
from src.post_training.chat_template import encode_chat

messages = [
    {"role": "user", "content": "Explain Newton's second law."},
    {"role": "assistant", "content": "Force equals mass times acceleration."}
]

ids, loss_mask = encode_chat(messages)

# loss_mask contains 1s for assistant tokens and EOT_ID, 0s elsewhere

```

### Packing Variable-Length Examples

The `pack_examples` function concatenates multiple variable-length conversations into fixed-length rows of length `context_length`. This maximizes GPU utilization by reducing padding waste. The packed HDF5 contains both `tokens` and `loss_mask` datasets with identical shapes.

## Reinforcement Learning (RL) Prompt Preparation

For RL stages such as PPO or GRPO, the pipeline in [`data_loader/prompt_dataset.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/prompt_dataset.py) prepares prompts without completions. The `get_prompt_iterator` function loads a JSONL file containing `{"prompt": ..., "gold": ...}` objects and shuffles them per epoch.

The `encode_prompt` function (also in [`src/post_training/chat_template.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/chat_template.py)) tokenizes the prompt and appends the assistant header, but unlike `encode_chat`, it produces no completion and no loss mask. This creates a zero-loss-mask prompt ready for rollout generation in [`src/post_training/rollout.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/rollout.py).

## Core Tokenization Utilities

All token-related logic resides in [`src/post_training/chat_template.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/chat_template.py), providing a consistent interface across training stages.

### encode_chat vs encode_prompt

- **`encode_chat(messages, add_generation_prompt=False)`**: Returns `(ids, loss_mask)` for supervised training. When `add_generation_prompt` is True, it appends the assistant header to the end, preparing the input for response generation.
- **`encode_prompt(messages)`**: Returns `ids` only, ending with the assistant header. Used for RL rollouts where the model must generate the completion.

### Decoding and Safety

The `decode(ids)` function converts token IDs back to human-readable text, safely dropping any IDs greater than or equal to `EOT_ID` to handle potential special token leakage.

## Practical Code Examples

### Tokenizing Raw Pile Shards for Pre-training

Create a pre-training token file from the validation split:

```python
import subprocess
import h5py

out_path = "/ephemeral/data/pile_val.h5"
cmd = [
    "python", "scripts/prepare_pretrain_data.py",
    "--split", "val",
    "--out", out_path,
    "--max_tokens", "10000000"
]
subprocess.run(cmd, check=True)

# Verify the token count

with h5py.File(out_path, "r") as f:
    print("Token count:", f["tokens"].shape[0])

```

### Building an SFT Dataset with Loss Masks

Convert chat examples to packed HDF5 format:

```python
from src.post_training.chat_template import encode_chat
import h5py
import numpy as np

messages = [
    {"role": "user", "content": "Explain Newton's second law."},
    {"role": "assistant", "content": "Force equals mass times acceleration."}
]

ids, loss_mask = encode_chat(messages)
context_len = 1024

# Initialize fixed-length buffers

tokens = np.zeros((1, context_len), dtype=np.int32)
mask = np.zeros((1, context_len), dtype=np.int8)

# Fill with actual data

tokens[0, :len(ids)] = ids
mask[0, :len(loss_mask)] = loss_mask

with h5py.File("sft_example.h5", "w") as f:
    f.create_dataset("tokens", data=tokens)
    f.create_dataset("loss_mask", data=mask)

```

### Preparing Prompts for RL Rollouts

Generate tokenized prompts for reinforcement learning:

```python
from src.post_training.chat_template import encode_prompt
from data_loader.prompt_dataset import get_prompt_iterator

prompt_iter = get_prompt_iterator(
    path="data/prefs.jsonl", 
    prompts_per_iter=4, 
    rank=0, 
    world_size=1, 
    seed=42
)

batch = next(prompt_iter)
ids_batch = [encode_prompt(p["prompt"]) for p in batch]

print("Prompt token ids:", ids_batch[0][:20])

```

## Summary

- **Pre-training** uses [`scripts/prepare_pretrain_data.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/prepare_pretrain_data.py) to stream Pile shards, encode with tiktoken `r50k_base`, and write flat tokens to HDF5 in 8M-token chunks.
- **SFT preparation** leverages `encode_chat` in [`src/post_training/chat_template.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/chat_template.py) to generate aligned token IDs and loss masks, stored via [`scripts/prepare_sft_data.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/prepare_sft_data.py).
- **RL prompts** use `encode_prompt` and `get_prompt_iterator` from [`data_loader/prompt_dataset.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/prompt_dataset.py) to create generation-ready inputs without loss masks.
- **Storage format** is HDF5 with chunked datasets for efficient random access, supporting both pre-training sequences and SFT packed examples.

## Frequently Asked Questions

### What tokenizer does this repository use?

The repository uses **tiktoken** with the `r50k_base` encoder. The only true special token is `<|endoftext|>` (ID `50256`), which serves as both the document separator and generation stop token. Role markers like `<|user|>` and `<|assistant|>` are plain text, not special tokens.

### How does the loss mask work in supervised fine-tuning?

The `encode_chat` function generates a binary `loss_mask` where assistant tokens and the final `EOT_ID` are marked `1`, while user and system tokens are marked `0`. This ensures that during training in [`src/post_training/sft.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/sft.py), the cross-entropy loss is computed only on the assistant's responses, not the user's prompts.

### Why does the pipeline use HDF5 instead of other formats?

HDF5 provides memory-mapped access to large datasets without loading everything into RAM, supports efficient chunked appending (critical for the 8M-token write buffers), and allows random sampling of sequences during training. This is essential for handling multi-billion token corpora on limited memory systems.

### What is the difference between encode_chat and encode_prompt?

`encode_chat` returns both token IDs and a loss mask, making it suitable for supervised training where the model must learn to predict the assistant's response. `encode_prompt` returns only token IDs and is designed for RL inference, adding the assistant header but leaving the completion empty for the model to generate.