# How to Preprocess The Pile Dataset for Language Model Training

> Learn how to preprocess The Pile dataset for language model training with FareedKhan-devs efficient pipeline. Convert raw data to token-encoded HDF5 datasets for faster training.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-31

---

**The `FareedKhan-dev/train-llm-from-scratch` repository provides a self‑contained pipeline that converts raw Pile JSON‑LZ4 archives into token‑encoded HDF5 datasets using streaming decompression and the OpenAI tiktoken library, enabling efficient random access during model training without loading the full corpus into memory.**

Preparing the massive Pile corpus for transformer training requires handling compressed multi‑gigabyte shards, consistent tokenization, and a storage format that supports fast random slicing. The repository solves this by implementing a streaming preprocessor that walks raw `.jsonl.zst` files and emits contiguous token streams with explicit document boundaries.

## Overview of the Preprocessing Pipeline

The core logic resides in [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py), which orchestrates the transformation from compressed JSON lines to numeric token arrays. Unlike in‑memory approaches, this implementation processes data in a streaming fashion, handling 10 GB+ files without full extraction to disk.

The pipeline produces two HDF5 files—`pile_train.h5` and `pile_dev.h5`—that are consumed by [`data_loader/data_loader.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/data_loader.py) during training. Because HDF5 supports random reads with low overhead, the trainer can fetch arbitrarily sized context windows (defined by `CONTEXT_LENGTH` in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py)) without loading the entire dataset into RAM.

## Step‑by‑Step Data Transformation

### File Discovery and Filtering

The process begins by scanning the input directory for Pile shards. In lines 38‑40 of [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py), the script uses `os.listdir` combined with a suffix filter to identify only files ending in `.jsonl.zst`, guaranteeing deterministic ordering and reproducibility across runs.

### Streaming Decompression

Rather than decompressing entire files upfront, the implementation leverages the `zstandard` library to open compressed streams in text mode. As shown in lines 45‑48, `zstd.open(..., 'rt')` enables line‑by‑line iteration with progress tracking via `tqdm`, keeping RAM usage constant regardless of file size.

### JSON Parsing and Text Extraction

Each line is parsed as JSON using `json.loads` (lines 52‑56). The script extracts the mandatory `"text"` field and validates its presence, emitting warnings for malformed entries to prevent pipeline failures. Every valid document is appended with the `<|endoftext|>` special token to provide clear document boundaries for the language model.

### Tokenization with tiktoken

The pipeline uses the OpenAI `tiktoken` library for byte‑pair encoding. Lines 28‑30 initialize the tokenizer via `tiktoken.get_encoding(tokenizer_name)`, defaulting to `r50k_base` but supporting alternatives like `cl100k_base`. The extracted text strings are converted to integer token IDs efficiently in batches.

### HDF5 Dataset Creation and Dynamic Resizing

Token IDs are written to a single expandable HDF5 dataset named `tokens`. Lines 59‑66 create the dataset with `maxshape=(None,)` to allow unbounded growth, then use `resize` to accommodate each new batch while preserving the global token order across all input files. This produces a contiguous stream suitable for sliding window batching.

## Configuration and Command‑Line Usage

The script exposes several arguments for controlling the preprocessing behavior. You can limit the number of processed JSON objects per shard using `--max_data`, which is essential for rapid prototyping on limited hardware.

To download and preprocess a subset of the data:

```bash

# Download three training shards plus validation data

python scripts/data_download.py --train_max 3

# Convert to token HDF5 with a processing cap

python scripts/data_preprocess.py \
    --train_dir data/train \
    --val_dir data/val \
    --out_train_file data/train/pile_train.h5 \
    --out_val_file data/val/pile_dev.h5 \
    --tokenizer_name r50k_base \
    --max_data 2000

```

To verify the output file structure:

```python
import h5py

with h5py.File('data/train/pile_train.h5', 'r') as f:
    tokens = f['tokens']
    print(f'Dataset shape: {tokens.shape}')
    print(f'First 10 token IDs: {tokens[:10]}')

```

### Switching Tokenizers

You can change the encoding scheme by passing a different tokenizer name:

```bash
python scripts/data_preprocess.py --tokenizer_name cl100k_base

```

### Processing the Full Dataset

To process the entire Pile corpus, omit the `--max_data` argument or set it to a very large integer. The streaming architecture ensures memory usage remains constant even when processing hundreds of shards.

## Integration with the Training Pipeline

The generated HDF5 files are consumed by [`data_loader/data_loader.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/data_loader.py), specifically by the `get_batch_iterator` function. Because the HDF5 format supports efficient random reads, the training loop can fetch arbitrarily sized windows without loading the entire dataset into memory.

The preprocessing script respects document boundaries by inserting the `<|endoftext|>` token ID between Pile documents. This allows the data loader to construct training sequences that do not cross unrelated texts, improving model coherence.

## Summary

- **Streaming architecture**: Processes `.jsonl.zst` files line‑by‑line using `zstandard` to handle massive shards without full extraction.
- **Tokenization**: Uses `tiktoken` with configurable encodings (`r50k_base`, `cl100k_base`) to convert text to integer IDs.
- **Storage format**: Writes to expandable HDF5 datasets (`maxshape=(None,)`) for efficient random access and low‑memory training.
- **Document boundaries**: Appends `<|endoftext|>` tokens to separate Pile documents clearly.
- **Prototyping support**: The `--max_data` argument enables quick experiments on subsets of the corpus.
- **Training integration**: Output files feed directly into [`data_loader/data_loader.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/data_loader/data_loader.py) for batched training via random slicing.

## Frequently Asked Questions

### Can I preprocess the entire Pile dataset without memory issues?

Yes. The [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py) script processes files in streaming mode using `zstd.open`, which decompresses data on‑the‑fly. The HDF5 output is written incrementally with dynamic resizing, ensuring memory usage remains constant regardless of input size. Simply omit the `--max_data` argument to process all JSON objects in every shard.

### How do I change the tokenizer for preprocessing?

Modify the `--tokenizer_name` argument when invoking the script. The default is `r50k_base`, but you can specify any encoding supported by `tiktoken`, such as `cl100k_base`. The script initializes the encoder via `tiktoken.get_encoding(tokenizer_name)` at line 28, making the tokenizer swappable without code changes.

### Why does the pipeline append `<|endoftext|>` to every document?

The special token acts as an explicit document boundary marker. During training, the model learns to predict this token when a document ends, preventing it from continuing across unrelated texts. This is implemented in lines 52‑56 of [`scripts/data_preprocess.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/data_preprocess.py) by concatenating the token string to each JSON text field before tokenization.

### Can I parallelize the preprocessing across multiple shards?

While the script is single‑process, you can run multiple instances in parallel on disjoint shard directories. Point each instance to a different `--train_dir` subfolder and specify unique output paths. Later, concatenate the resulting HDF5 files using tools like `h5copy` to create a unified dataset for training.