# How to Load and Preprocess HDF5 Data for LLM Training with train-llm-from-scratch

> Learn to load and preprocess HDF5 data for LLM training. Convert JSONL.ZST to HDF5 using scripts and stream batched tensors efficiently with train-llm-from-scratch for PyTorch.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-31

---

**To load and preprocess HDF5 data for LLM training, convert raw `.jsonl.zst` files into tokenized HDF5 format using [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py), then stream batched tensors via `get_batch_iterator()` from [`data_loader/data_loader.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/data_loader.py) for PyTorch training loops.**

The **train-llm-from-scratch** repository provides a lightweight, file-system-based pipeline for training large language models from raw text. This guide explains how to load and preprocess HDF5 data for LLM training using its specialized preprocessing scripts and memory-efficient data loaders.

## Converting Raw Text to HDF5 Format

The preprocessing stage transforms compressed JSON-Lines datasets into a single HDF5 file containing tokenized integers. This eliminates tokenization overhead during training and enables fast random access.

### The data_preprocess.py Script

Located at [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py), this script handles the end-to-end conversion:

- **Input**: Compressed `.jsonl.zst` files (e.g., from the PILE dataset)
- **Tokenization**: Uses `tiktoken.get_encoding(tokenizer_name)` to encode text fields
- **Special tokens**: Appends the `<|endoftext|>` delimiter after each document
- **Output**: Creates resizable HDF5 datasets named `tokens` in `train/pile_train.h5` and `val/pile_dev.h5`

The `process_files` function iterates through input directories, decodes each JSON line, and appends integer token IDs to the HDF5 dataset. You can limit processing scope using the `--max_data` parameter for quick experiments.

```bash
python scripts/data_preprocess.py \
  --train_dir data/pile/train \
  --val_dir   data/pile/val   \
  --out_train_file data/train/pile_train.h5 \
  --out_val_file   data/val/pile_dev.h5 \
  --tokenizer_name r50k_base \
  --max_data 50000

```

## Streaming Batches from HDF5

Once tokenized data resides in HDF5 format, the [`data_loader/data_loader.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/data_loader.py) module provides efficient batching without loading the entire dataset into RAM.

### The get_batch_iterator Function

This function opens the HDF5 file via `h5py.File(data_path, 'r')` and returns an infinite iterator that yields `(xb, yb)` tensor pairs:

- **Slicing**: Splits the flat `tokens` dataset into `context_length`-sized chunks
- **Shuffling**: Randomizes example indices at the start of each epoch
- **Batching**: Extracts `batch_size` sequences and places them on the specified device
- **Target generation**: Automatically creates target tensors `yb` as the next-token prediction labels for inputs `xb`

The iterator handles epoch boundaries automatically—when indices exhaust, they reshuffle and the epoch counter increments, providing a seamless `for xb, yb in iterator:` experience.

```python
import torch
from data_loader.data_loader import get_batch_iterator

# Configuration

hdf5_path = "data/train/pile_train.h5"
batch_size = 16
context_len = 128
device = "cuda" if torch.cuda.is_available() else "cpu"

# Create infinite iterator

train_iter = get_batch_iterator(
    data_path=hdf5_path,
    batch_size=batch_size,
    context_length=context_len,
    device=device,
)

# Training loop integration

model = ...  # Your transformer model

optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4)

for step in range(1000):
    xb, yb = next(train_iter)  # Shapes: [batch_size, context_length]

    logits = model(xb)
    loss = torch.nn.functional.cross_entropy(
        logits.view(-1, logits.size(-1)),
        yb.view(-1)
    )
    
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

```

## Sanity Check and Validation

Verify your pipeline with a minimal standalone test before launching full training:

```python
if __name__ == "__main__":
    import os, h5py, numpy as np
    from data_loader.data_loader import get_batch_iterator
    
    # Create dummy data

    dummy_path = "dummy.h5"
    if not os.path.exists(dummy_path):
        with h5py.File(dummy_path, "w") as f:
            f.create_dataset("tokens", data=np.arange(1000))
    
    # Test iterator

    for xb, yb in get_batch_iterator(dummy_path, batch_size=4, context_length=10):
        print(f"xb shape: {xb.shape}, yb shape: {yb.shape}")  # Should print [4, 10]

        break

```

This confirms the loader correctly slices the one-dimensional token stream into two-dimensional batches suitable for transformer models.

## Summary

- **Preprocessing**: Use [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py) to tokenize `.jsonl.zst` files with **tiktoken** and store results in HDF5 format under the `tokens` dataset.
- **Loading**: Import `get_batch_iterator` from [`data_loader/data_loader.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/data_loader.py) to stream `(input, target)` tensors directly from disk.
- **Memory efficiency**: The pipeline uses HDF5's memory-mapped access to handle datasets larger than RAM, slicing only required `context_length` windows.
- **Training integration**: The iterator yields PyTorch tensors on the specified device and automatically manages epoch shuffling for infinite training loops.

## Frequently Asked Questions

### What file format does train-llm-from-scratch use for tokenized data?

The repository uses **HDF5** (Hierarchical Data Format) to store tokenized training data. Specifically, [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py) writes a one-dimensional integer array to a dataset named `tokens` within `.h5` files, enabling fast random access and efficient memory mapping during training.

### How does the data loader handle batching and shuffling?

The `get_batch_iterator()` function in [`data_loader/data_loader.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/data_loader.py) creates an index array spanning the full token sequence, splits it into `context_length`-sized chunks, and shuffles these indices at the start of each epoch. It then yields batches of shape `(batch_size, context_length)` by slicing the HDF5 dataset at the shuffled indices, ensuring randomization while maintaining sequential token continuity within each example.

### Why use HDF5 instead of raw text files for LLM training?

HDF5 provides **memory-mapped I/O** capabilities that allow the data loader to access arbitrary portions of large datasets without loading the entire file into RAM. This is critical when training on terabyte-scale text corpora, as it enables efficient random batch sampling and eliminates the preprocessing bottleneck during training iterations.

### What tokenizer does the preprocessing script use?

The [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py) script uses **tiktoken**, OpenAI's fast BPE tokenizer library. The specific encoding (such as `r50k_base` or `cl100k_base`) is configurable via the `--tokenizer_name` argument, and the script automatically appends the `<|endoftext|>` special token to mark document boundaries in the token stream.