How to Load and Preprocess HDF5 Data for LLM Training with train-llm-from-scratch
To load and preprocess HDF5 data for LLM training, convert raw .jsonl.zst files into tokenized HDF5 format using scripts/data_preprocess.py, then stream batched tensors via get_batch_iterator() from data_loader/data_loader.py for PyTorch training loops.
The train-llm-from-scratch repository provides a lightweight, file-system-based pipeline for training large language models from raw text. This guide explains how to load and preprocess HDF5 data for LLM training using its specialized preprocessing scripts and memory-efficient data loaders.
Converting Raw Text to HDF5 Format
The preprocessing stage transforms compressed JSON-Lines datasets into a single HDF5 file containing tokenized integers. This eliminates tokenization overhead during training and enables fast random access.
The data_preprocess.py Script
Located at scripts/data_preprocess.py, this script handles the end-to-end conversion:
- Input: Compressed
.jsonl.zstfiles (e.g., from the PILE dataset) - Tokenization: Uses
tiktoken.get_encoding(tokenizer_name)to encode text fields - Special tokens: Appends the
<|endoftext|>delimiter after each document - Output: Creates resizable HDF5 datasets named
tokensintrain/pile_train.h5andval/pile_dev.h5
The process_files function iterates through input directories, decodes each JSON line, and appends integer token IDs to the HDF5 dataset. You can limit processing scope using the --max_data parameter for quick experiments.
python scripts/data_preprocess.py \
--train_dir data/pile/train \
--val_dir data/pile/val \
--out_train_file data/train/pile_train.h5 \
--out_val_file data/val/pile_dev.h5 \
--tokenizer_name r50k_base \
--max_data 50000
Streaming Batches from HDF5
Once tokenized data resides in HDF5 format, the data_loader/data_loader.py module provides efficient batching without loading the entire dataset into RAM.
The get_batch_iterator Function
This function opens the HDF5 file via h5py.File(data_path, 'r') and returns an infinite iterator that yields (xb, yb) tensor pairs:
- Slicing: Splits the flat
tokensdataset intocontext_length-sized chunks - Shuffling: Randomizes example indices at the start of each epoch
- Batching: Extracts
batch_sizesequences and places them on the specified device - Target generation: Automatically creates target tensors
ybas the next-token prediction labels for inputsxb
The iterator handles epoch boundaries automatically—when indices exhaust, they reshuffle and the epoch counter increments, providing a seamless for xb, yb in iterator: experience.
import torch
from data_loader.data_loader import get_batch_iterator
# Configuration
hdf5_path = "data/train/pile_train.h5"
batch_size = 16
context_len = 128
device = "cuda" if torch.cuda.is_available() else "cpu"
# Create infinite iterator
train_iter = get_batch_iterator(
data_path=hdf5_path,
batch_size=batch_size,
context_length=context_len,
device=device,
)
# Training loop integration
model = ... # Your transformer model
optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4)
for step in range(1000):
xb, yb = next(train_iter) # Shapes: [batch_size, context_length]
logits = model(xb)
loss = torch.nn.functional.cross_entropy(
logits.view(-1, logits.size(-1)),
yb.view(-1)
)
optimizer.zero_grad()
loss.backward()
optimizer.step()
Sanity Check and Validation
Verify your pipeline with a minimal standalone test before launching full training:
if __name__ == "__main__":
import os, h5py, numpy as np
from data_loader.data_loader import get_batch_iterator
# Create dummy data
dummy_path = "dummy.h5"
if not os.path.exists(dummy_path):
with h5py.File(dummy_path, "w") as f:
f.create_dataset("tokens", data=np.arange(1000))
# Test iterator
for xb, yb in get_batch_iterator(dummy_path, batch_size=4, context_length=10):
print(f"xb shape: {xb.shape}, yb shape: {yb.shape}") # Should print [4, 10]
break
This confirms the loader correctly slices the one-dimensional token stream into two-dimensional batches suitable for transformer models.
Summary
- Preprocessing: Use
scripts/data_preprocess.pyto tokenize.jsonl.zstfiles with tiktoken and store results in HDF5 format under thetokensdataset. - Loading: Import
get_batch_iteratorfromdata_loader/data_loader.pyto stream(input, target)tensors directly from disk. - Memory efficiency: The pipeline uses HDF5's memory-mapped access to handle datasets larger than RAM, slicing only required
context_lengthwindows. - Training integration: The iterator yields PyTorch tensors on the specified device and automatically manages epoch shuffling for infinite training loops.
Frequently Asked Questions
What file format does train-llm-from-scratch use for tokenized data?
The repository uses HDF5 (Hierarchical Data Format) to store tokenized training data. Specifically, scripts/data_preprocess.py writes a one-dimensional integer array to a dataset named tokens within .h5 files, enabling fast random access and efficient memory mapping during training.
How does the data loader handle batching and shuffling?
The get_batch_iterator() function in data_loader/data_loader.py creates an index array spanning the full token sequence, splits it into context_length-sized chunks, and shuffles these indices at the start of each epoch. It then yields batches of shape (batch_size, context_length) by slicing the HDF5 dataset at the shuffled indices, ensuring randomization while maintaining sequential token continuity within each example.
Why use HDF5 instead of raw text files for LLM training?
HDF5 provides memory-mapped I/O capabilities that allow the data loader to access arbitrary portions of large datasets without loading the entire file into RAM. This is critical when training on terabyte-scale text corpora, as it enables efficient random batch sampling and eliminates the preprocessing bottleneck during training iterations.
What tokenizer does the preprocessing script use?
The scripts/data_preprocess.py script uses tiktoken, OpenAI's fast BPE tokenizer library. The specific encoding (such as r50k_base or cl100k_base) is configurable via the --tokenizer_name argument, and the script automatically appends the <|endoftext|> special token to mark document boundaries in the token stream.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →