# Creating Instruction Finetuning Datasets from Scratch: A Complete Guide to the LLMs-from-scratch Pipeline

> Learn to create instruction finetuning datasets from scratch for LLMs. This guide covers data formatting with JSON, duplicate cleaning, and GPT training for custom models.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-12

---

**You can build a custom instruction-finetuning dataset from scratch by formatting your data as JSON objects with `instruction`, `input`, and `output` fields, cleaning duplicates using TF-IDF similarity detection, and feeding the result into a lightweight GPT training pipeline.**

The **LLMs-from-scratch** repository by Sebastian Raschka provides an end-to-end educational framework for creating, cleaning, and training on instruction-finetuning datasets. Unlike black-box APIs, this implementation exposes every step—from JSON data preparation to the supervised finetuning loop—making it ideal for researchers who need full control over their training data and model behavior.

## Dataset Format and Structure

The foundation of instruction finetuning is a properly structured JSON dataset. According to the reference implementation in [`ch07/01_main-chapter-code/instruction-data.json`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/instruction-data.json), each entry must contain three specific fields:

```json
{
  "instruction": "Describe the process of photosynthesis in simple terms.",
  "input": "",
  "output": "Photosynthesis is the process by which plants use sunlight, water, and carbon dioxide to create oxygen and energy in the form of sugar."
}

```

- **`instruction`**: The task description or command for the model.
- **`input`**: Optional context or auxiliary information (can be an empty string).
- **`output`**: The target response that the model should learn to generate.

The repository includes a reference dataset containing approximately 1,100 such examples, providing a template for creating custom datasets for specific domains or languages.

## Cleaning and Deduplicating Your Dataset

Raw scraped or crowdsourced data often contains near-duplicate entries that can bias training. The repository ships with [`ch07/02_dataset-utilities/find-near-duplicates.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/02_dataset-utilities/find-near-duplicates.py), a utility that uses scikit-learn's TF-IDF vectorizer with character-level n-grams to detect semantic duplicates.

### Removing Near-Duplicates with TF-IDF

The `find_print_and_remove_near_duplicates` function computes similarity scores between entries and optionally prunes them based on a configurable threshold:

```python
from find_near_duplicates import find_print_and_remove_near_duplicates

# Load your raw dataset

with open("my_instruction_dataset.json", "r") as f:
    data = json.load(f)

# Detect and remove duplicates with 85% similarity threshold

clean_data = find_print_and_remove_near_duplicates(
    json_data=data,
    remove_duplicates=True,
    threshold=0.85
)

# Save cleaned version

with open("my_instruction_dataset_clean.json", "w") as f:
    json.dump(clean_data, f, indent=4)

```

Key implementation details from the source code:
- Uses character-level n-grams (`ngram_range=(1,3)`) to catch subtle phrasing variations.
- By default checks the `input` and `output` fields for duplicates while preserving the instruction text (lines 66-69 of [`find-near-duplicates.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/find-near-duplicates.py)).
- The threshold parameter accepts values between 0 and 1, where 0.85 provides a conservative balance between removing true duplicates and preserving valid variations.

## The Instruction Finetuning Pipeline

Once your dataset is cleaned, the pipeline in [`ch07/01_main-chapter-code/gpt_instruction_finetuning.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/gpt_instruction_finetuning.py) orchestrates the complete training workflow. This script handles data loading, tokenization, model initialization, and supervised training.

### Data Loading and Tokenization

The pipeline uses OpenAI's `tiktoken` library with the GPT-2 encoding. The `custom_collate_fn` (lines 56-96 in the training script) handles batching by padding sequences to equal length and masking padding tokens with `ignore_index = -100` in the loss computation:

```python
import tiktoken
from functools import partial

tokenizer = tiktoken.get_encoding("gpt2")

# Create collate function that pads batches and masks padding tokens

customized_collate_fn = partial(
    custom_collate_fn, device=device, allowed_max_length=1024
)

```

### Model Configuration Options

The repository supports two configuration modes:
- **Test Mode**: A tiny 12-dimensional GPT model (configurable in lines 24-38) that trains quickly on CPU for debugging.
- **Full GPT-2**: Any standard GPT-2 variant (124M to 1.5B parameters) with pretrained weights downloaded via `download_and_load_gpt2()` from [`gpt_download.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/gpt_download.py).

### Training Loop Implementation

The `train_model_simple` function (imported from [`previous_chapters.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/previous_chapters.py)) executes the supervised finetuning:

```python
optimizer = torch.optim.AdamW(model.parameters(), lr=5e-5, weight_decay=0.1)

train_losses, val_losses, tokens_seen = train_model_simple(
    model, train_loader, val_loader, optimizer, device,
    num_epochs=2, 
    eval_freq=5, 
    eval_iter=5,
    start_context=format_input(val_data[0]), 
    tokenizer=tokenizer
)

```

This function automatically handles validation evaluation every 5 steps and generates sample completions to monitor training progress.

### Generating and Saving Responses

After finetuning, the script generates responses for the test set using the `generate` function and stores results:

```python
for i, entry in enumerate(test_data):
    input_text = format_input(entry)
    token_ids = generate(
        model=model,
        idx=text_to_token_ids(input_text, tokenizer).to(device),
        max_new_tokens=256,
        context_size=BASE_CONFIG["context_length"],
        eos_id=50256  # GPT-2 end-of-sequence token

    )
    generated_text = token_ids_to_text(token_ids, tokenizer)
    response_text = generated_text[len(input_text):].replace("### Response:", "").strip()

    test_data[i]["model_response"] = response_text

```

The final checkpoint is saved using `torch.save(model.state_dict(), filename)` with the naming convention `{MODEL_NAME}-sft-standalone.pth`.

## Step-by-Step Implementation

### End-to-End Finetuning Workflow

Execute the complete pipeline from the repository root:

```bash
cd ch07/01_main-chapter-code

# Quick test run with minimal data and tiny model

python gpt_instruction_finetuning.py --test_mode

# Full training with GPT-2 medium

python gpt_instruction_finetuning.py

```

### Loading a Finetuned Model for Inference

To use your finetuned model after training:

```python
import torch
from pkg.llms_from_scratch.ch05 import GPTModel

# Must match training configuration

config = {
    "vocab_size": 50257,
    "context_length": 1024,
    "drop_rate": 0.0,
    "qkv_bias": True,
    "emb_dim": 768,
    "n_layers": 12,
    "n_heads": 12,
}

model = GPTModel(config)
model.load_state_dict(torch.load("gpt2-medium-sft-standalone.pth"))
model.eval()

# Inference

import tiktoken
enc = tiktoken.get_encoding("gpt2")
prompt = "Below is an instruction that describes a task...\n\n### Instruction:\nSummarize the following text\n\n### Input:\nMachine learning is...\n\n### Response:\n"

input_ids = torch.tensor(enc.encode(prompt)).unsqueeze(0)

output_ids = generate(
    model=model,
    idx=input_ids,
    max_new_tokens=200,
    context_size=config["context_length"],
    eos_id=50256
)
print(enc.decode(output_ids.squeeze().tolist()))

```

## Summary

- **Dataset Format**: Create JSON files with `instruction`, `input`, and `output` fields to define supervised training targets.
- **Data Cleaning**: Use [`find-near-duplicates.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/find-near-duplicates.py) with TF-IDF vectorization to remove duplicate entries with configurable similarity thresholds (default 0.85).
- **Training Pipeline**: The [`gpt_instruction_finetuning.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/gpt_instruction_finetuning.py) script handles tokenization (via `tiktoken`), batching with padding masks (`ignore_index=-100`), and supervised training using `train_model_simple`.
- **Model Flexibility**: Supports both tiny 12-dimensional test models for debugging and full GPT-2 variants for production finetuning.
- **Artifacts**: The pipeline produces JSON files containing model responses and `.pth` checkpoint files compatible with the `GPTModel` class defined in [`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py).

## Frequently Asked Questions

### What is the exact JSON schema required for instruction finetuning?

The repository expects a list of JSON objects where each object contains three string fields: `instruction` (the task description), `input` (optional context or user query), and `output` (the desired response). The `input` field can be an empty string if the instruction contains all necessary context. This schema is defined in [`ch07/01_main-chapter-code/instruction-data.json`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/instruction-data.json) and parsed by the `format_input` function in the training script.

### How do I prevent data leakage between training and validation sets?

Use the `find_print_and_remove_near_duplicates` utility from [`ch07/02_dataset-utilities/find-near-duplicates.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/02_dataset-utilities/find-near-duplicates.py) before splitting your data. This function calculates TF-IDF vectors on character-level n-grams and removes entries with similarity scores above your chosen threshold (typically 0.85) to ensure no near-duplicates exist across splits.

### Can I finetune on aCPU-only machine?

Yes. The repository includes a `test_mode` flag in [`gpt_instruction_finetuning.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/gpt_instruction_finetuning.py) that instantiates a minimal 12-dimensional GPT model instead of the full GPT-2 architecture. This test configuration trains in minutes on CPU and validates the entire pipeline before committing to GPU resources for larger models.

### What token is used to indicate the end of generation?

The pipeline uses GPT-2's end-of-sequence token ID **50256** (referenced as `eos_id` in the `generate` function calls). During training, `custom_collate_fn` appends this token to each sequence, and during inference, generation stops when this token is produced or when `max_new_tokens` is reached.