How to Process and Tokenize Data for LLM Training: A Complete Pipeline Guide

The train-llm-from-scratch repository implements a three-stage pipeline that converts raw text corpora into efficient HDF5 token files using tiktoken, with specialized loss masks for pre-training, supervised fine-tuning, and reinforcement learning.

Processing and tokenizing data for LLM training requires distinct strategies for each stage of the model lifecycle. In the FareedKhan-dev/train-llm-from-scratch repository, the data preparation pipeline handles everything from streaming Pile dataset shards to packing instruction-following examples with per-token loss masks. This implementation uses tiktoken with the r50k_base encoder and stores all token streams in chunked HDF5 format for efficient random access during training.

Pre-training Data Preparation

The pre-training stage converts raw Pile shards into a single flat-token HDF5 file that the model can read efficiently. This process is orchestrated by scripts/prepare_pretrain_data.py.

Downloading and Streaming Pile Shards

The script downloads shards from HuggingFace in .jsonl.zst format and stream-decompresses each shard to minimize memory overhead. Rather than holding entire shards in memory, the pipeline processes text in batches, encoding paragraphs as they arrive.

Tokenization and Special Tokens

Each text batch is encoded using tiktoken with the r50k_base encoder. The only true special token used is the end-of-text token (<|endoftext|>), which maps to EOT_ID = 50256. This token is appended to every document and serves as the generation stop token. The implementation uses tiktoken.encode_ordinary_batch for efficient batch processing.

Efficient HDF5 Storage

Tokens are written to HDF5 in chunks to avoid repeated dataset resizing. The implementation uses a write chunk size of 8_000_000 tokens:

with h5py.File(out_path, "w") as f:
    dset = f.create_dataset("tokens", (0,), maxshape=(None,), dtype="i4", chunks=(8_000_000,))

This approach appends tokens incrementally and stores files under /ephemeral/data by default, leveraging large temporary disks for high-throughput I/O.

Supervised Fine-Tuning (SFT) Data Preparation

The SFT pipeline in scripts/prepare_sft_data.py builds training data from public instruction datasets including Alpaca, Dolly, and GSM8K. Unlike pre-training, this stage requires per-token loss masks to distinguish between user prompts (loss = 0) and assistant responses (loss = 1).

Chat Format and Role Markers

The repository converts each example to a chat format with role and content fields. Role markers (<|user|>, <|assistant|>, <|system|>) are plain-text strings that are tokenized like any other text; the model learns to treat them as structural cues during training. These markers are defined in src/post_training/chat_template.py.

Loss Mask Generation with encode_chat

The core function encode_chat(messages, add_generation_prompt=False) returns two aligned arrays: token ids and a binary loss_mask. Only assistant tokens and the trailing EOT_ID are marked with mask == 1, ensuring the model learns only to predict the response:

from src.post_training.chat_template import encode_chat

messages = [
    {"role": "user", "content": "Explain Newton's second law."},
    {"role": "assistant", "content": "Force equals mass times acceleration."}
]

ids, loss_mask = encode_chat(messages)

# loss_mask contains 1s for assistant tokens and EOT_ID, 0s elsewhere

Packing Variable-Length Examples

The pack_examples function concatenates multiple variable-length conversations into fixed-length rows of length context_length. This maximizes GPU utilization by reducing padding waste. The packed HDF5 contains both tokens and loss_mask datasets with identical shapes.

Reinforcement Learning (RL) Prompt Preparation

For RL stages such as PPO or GRPO, the pipeline in data_loader/prompt_dataset.py prepares prompts without completions. The get_prompt_iterator function loads a JSONL file containing {"prompt": ..., "gold": ...} objects and shuffles them per epoch.

The encode_prompt function (also in src/post_training/chat_template.py) tokenizes the prompt and appends the assistant header, but unlike encode_chat, it produces no completion and no loss mask. This creates a zero-loss-mask prompt ready for rollout generation in src/post_training/rollout.py.

Core Tokenization Utilities

All token-related logic resides in src/post_training/chat_template.py, providing a consistent interface across training stages.

encode_chat vs encode_prompt

  • encode_chat(messages, add_generation_prompt=False): Returns (ids, loss_mask) for supervised training. When add_generation_prompt is True, it appends the assistant header to the end, preparing the input for response generation.
  • encode_prompt(messages): Returns ids only, ending with the assistant header. Used for RL rollouts where the model must generate the completion.

Decoding and Safety

The decode(ids) function converts token IDs back to human-readable text, safely dropping any IDs greater than or equal to EOT_ID to handle potential special token leakage.

Practical Code Examples

Tokenizing Raw Pile Shards for Pre-training

Create a pre-training token file from the validation split:

import subprocess
import h5py

out_path = "/ephemeral/data/pile_val.h5"
cmd = [
    "python", "scripts/prepare_pretrain_data.py",
    "--split", "val",
    "--out", out_path,
    "--max_tokens", "10000000"
]
subprocess.run(cmd, check=True)

# Verify the token count

with h5py.File(out_path, "r") as f:
    print("Token count:", f["tokens"].shape[0])

Building an SFT Dataset with Loss Masks

Convert chat examples to packed HDF5 format:

from src.post_training.chat_template import encode_chat
import h5py
import numpy as np

messages = [
    {"role": "user", "content": "Explain Newton's second law."},
    {"role": "assistant", "content": "Force equals mass times acceleration."}
]

ids, loss_mask = encode_chat(messages)
context_len = 1024

# Initialize fixed-length buffers

tokens = np.zeros((1, context_len), dtype=np.int32)
mask = np.zeros((1, context_len), dtype=np.int8)

# Fill with actual data

tokens[0, :len(ids)] = ids
mask[0, :len(loss_mask)] = loss_mask

with h5py.File("sft_example.h5", "w") as f:
    f.create_dataset("tokens", data=tokens)
    f.create_dataset("loss_mask", data=mask)

Preparing Prompts for RL Rollouts

Generate tokenized prompts for reinforcement learning:

from src.post_training.chat_template import encode_prompt
from data_loader.prompt_dataset import get_prompt_iterator

prompt_iter = get_prompt_iterator(
    path="data/prefs.jsonl", 
    prompts_per_iter=4, 
    rank=0, 
    world_size=1, 
    seed=42
)

batch = next(prompt_iter)
ids_batch = [encode_prompt(p["prompt"]) for p in batch]

print("Prompt token ids:", ids_batch[0][:20])

Summary

Frequently Asked Questions

What tokenizer does this repository use?

The repository uses tiktoken with the r50k_base encoder. The only true special token is <|endoftext|> (ID 50256), which serves as both the document separator and generation stop token. Role markers like <|user|> and <|assistant|> are plain text, not special tokens.

How does the loss mask work in supervised fine-tuning?

The encode_chat function generates a binary loss_mask where assistant tokens and the final EOT_ID are marked 1, while user and system tokens are marked 0. This ensures that during training in src/post_training/sft.py, the cross-entropy loss is computed only on the assistant's responses, not the user's prompts.

Why does the pipeline use HDF5 instead of other formats?

HDF5 provides memory-mapped access to large datasets without loading everything into RAM, supports efficient chunked appending (critical for the 8M-token write buffers), and allows random sampling of sequences during training. This is essential for handling multi-billion token corpora on limited memory systems.

What is the difference between encode_chat and encode_prompt?

encode_chat returns both token IDs and a loss mask, making it suitable for supervised training where the model must learn to predict the assistant's response. encode_prompt returns only token IDs and is designed for RL inference, adding the assistant header but leaving the completion empty for the model to generate.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →