How to Load and Preprocess HDF5 Data for LLM Training with train-llm-from-scratch

To load and preprocess HDF5 data for LLM training, convert raw .jsonl.zst files into tokenized HDF5 format using scripts/data_preprocess.py, then stream batched tensors via get_batch_iterator() from data_loader/data_loader.py for PyTorch training loops.

The train-llm-from-scratch repository provides a lightweight, file-system-based pipeline for training large language models from raw text. This guide explains how to load and preprocess HDF5 data for LLM training using its specialized preprocessing scripts and memory-efficient data loaders.

Converting Raw Text to HDF5 Format

The preprocessing stage transforms compressed JSON-Lines datasets into a single HDF5 file containing tokenized integers. This eliminates tokenization overhead during training and enables fast random access.

The data_preprocess.py Script

Located at scripts/data_preprocess.py, this script handles the end-to-end conversion:

  • Input: Compressed .jsonl.zst files (e.g., from the PILE dataset)
  • Tokenization: Uses tiktoken.get_encoding(tokenizer_name) to encode text fields
  • Special tokens: Appends the <|endoftext|> delimiter after each document
  • Output: Creates resizable HDF5 datasets named tokens in train/pile_train.h5 and val/pile_dev.h5

The process_files function iterates through input directories, decodes each JSON line, and appends integer token IDs to the HDF5 dataset. You can limit processing scope using the --max_data parameter for quick experiments.

python scripts/data_preprocess.py \
  --train_dir data/pile/train \
  --val_dir   data/pile/val   \
  --out_train_file data/train/pile_train.h5 \
  --out_val_file   data/val/pile_dev.h5 \
  --tokenizer_name r50k_base \
  --max_data 50000

Streaming Batches from HDF5

Once tokenized data resides in HDF5 format, the data_loader/data_loader.py module provides efficient batching without loading the entire dataset into RAM.

The get_batch_iterator Function

This function opens the HDF5 file via h5py.File(data_path, 'r') and returns an infinite iterator that yields (xb, yb) tensor pairs:

  • Slicing: Splits the flat tokens dataset into context_length-sized chunks
  • Shuffling: Randomizes example indices at the start of each epoch
  • Batching: Extracts batch_size sequences and places them on the specified device
  • Target generation: Automatically creates target tensors yb as the next-token prediction labels for inputs xb

The iterator handles epoch boundaries automatically—when indices exhaust, they reshuffle and the epoch counter increments, providing a seamless for xb, yb in iterator: experience.

import torch
from data_loader.data_loader import get_batch_iterator

# Configuration

hdf5_path = "data/train/pile_train.h5"
batch_size = 16
context_len = 128
device = "cuda" if torch.cuda.is_available() else "cpu"

# Create infinite iterator

train_iter = get_batch_iterator(
    data_path=hdf5_path,
    batch_size=batch_size,
    context_length=context_len,
    device=device,
)

# Training loop integration

model = ...  # Your transformer model

optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4)

for step in range(1000):
    xb, yb = next(train_iter)  # Shapes: [batch_size, context_length]

    logits = model(xb)
    loss = torch.nn.functional.cross_entropy(
        logits.view(-1, logits.size(-1)),
        yb.view(-1)
    )
    
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

Sanity Check and Validation

Verify your pipeline with a minimal standalone test before launching full training:

if __name__ == "__main__":
    import os, h5py, numpy as np
    from data_loader.data_loader import get_batch_iterator
    
    # Create dummy data

    dummy_path = "dummy.h5"
    if not os.path.exists(dummy_path):
        with h5py.File(dummy_path, "w") as f:
            f.create_dataset("tokens", data=np.arange(1000))
    
    # Test iterator

    for xb, yb in get_batch_iterator(dummy_path, batch_size=4, context_length=10):
        print(f"xb shape: {xb.shape}, yb shape: {yb.shape}")  # Should print [4, 10]

        break

This confirms the loader correctly slices the one-dimensional token stream into two-dimensional batches suitable for transformer models.

Summary

  • Preprocessing: Use scripts/data_preprocess.py to tokenize .jsonl.zst files with tiktoken and store results in HDF5 format under the tokens dataset.
  • Loading: Import get_batch_iterator from data_loader/data_loader.py to stream (input, target) tensors directly from disk.
  • Memory efficiency: The pipeline uses HDF5's memory-mapped access to handle datasets larger than RAM, slicing only required context_length windows.
  • Training integration: The iterator yields PyTorch tensors on the specified device and automatically manages epoch shuffling for infinite training loops.

Frequently Asked Questions

What file format does train-llm-from-scratch use for tokenized data?

The repository uses HDF5 (Hierarchical Data Format) to store tokenized training data. Specifically, scripts/data_preprocess.py writes a one-dimensional integer array to a dataset named tokens within .h5 files, enabling fast random access and efficient memory mapping during training.

How does the data loader handle batching and shuffling?

The get_batch_iterator() function in data_loader/data_loader.py creates an index array spanning the full token sequence, splits it into context_length-sized chunks, and shuffles these indices at the start of each epoch. It then yields batches of shape (batch_size, context_length) by slicing the HDF5 dataset at the shuffled indices, ensuring randomization while maintaining sequential token continuity within each example.

Why use HDF5 instead of raw text files for LLM training?

HDF5 provides memory-mapped I/O capabilities that allow the data loader to access arbitrary portions of large datasets without loading the entire file into RAM. This is critical when training on terabyte-scale text corpora, as it enables efficient random batch sampling and eliminates the preprocessing bottleneck during training iterations.

What tokenizer does the preprocessing script use?

The scripts/data_preprocess.py script uses tiktoken, OpenAI's fast BPE tokenizer library. The specific encoding (such as r50k_base or cl100k_base) is configurable via the --tokenizer_name argument, and the script automatically appends the <|endoftext|> special token to mark document boundaries in the token stream.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →