Creating Instruction Finetuning Datasets from Scratch: A Complete Guide to the LLMs-from-scratch Pipeline

You can build a custom instruction-finetuning dataset from scratch by formatting your data as JSON objects with instruction, input, and output fields, cleaning duplicates using TF-IDF similarity detection, and feeding the result into a lightweight GPT training pipeline.

The LLMs-from-scratch repository by Sebastian Raschka provides an end-to-end educational framework for creating, cleaning, and training on instruction-finetuning datasets. Unlike black-box APIs, this implementation exposes every step—from JSON data preparation to the supervised finetuning loop—making it ideal for researchers who need full control over their training data and model behavior.

Dataset Format and Structure

The foundation of instruction finetuning is a properly structured JSON dataset. According to the reference implementation in ch07/01_main-chapter-code/instruction-data.json, each entry must contain three specific fields:

{
  "instruction": "Describe the process of photosynthesis in simple terms.",
  "input": "",
  "output": "Photosynthesis is the process by which plants use sunlight, water, and carbon dioxide to create oxygen and energy in the form of sugar."
}
  • instruction: The task description or command for the model.
  • input: Optional context or auxiliary information (can be an empty string).
  • output: The target response that the model should learn to generate.

The repository includes a reference dataset containing approximately 1,100 such examples, providing a template for creating custom datasets for specific domains or languages.

Cleaning and Deduplicating Your Dataset

Raw scraped or crowdsourced data often contains near-duplicate entries that can bias training. The repository ships with ch07/02_dataset-utilities/find-near-duplicates.py, a utility that uses scikit-learn's TF-IDF vectorizer with character-level n-grams to detect semantic duplicates.

Removing Near-Duplicates with TF-IDF

The find_print_and_remove_near_duplicates function computes similarity scores between entries and optionally prunes them based on a configurable threshold:

from find_near_duplicates import find_print_and_remove_near_duplicates

# Load your raw dataset

with open("my_instruction_dataset.json", "r") as f:
    data = json.load(f)

# Detect and remove duplicates with 85% similarity threshold

clean_data = find_print_and_remove_near_duplicates(
    json_data=data,
    remove_duplicates=True,
    threshold=0.85
)

# Save cleaned version

with open("my_instruction_dataset_clean.json", "w") as f:
    json.dump(clean_data, f, indent=4)

Key implementation details from the source code:

  • Uses character-level n-grams (ngram_range=(1,3)) to catch subtle phrasing variations.
  • By default checks the input and output fields for duplicates while preserving the instruction text (lines 66-69 of find-near-duplicates.py).
  • The threshold parameter accepts values between 0 and 1, where 0.85 provides a conservative balance between removing true duplicates and preserving valid variations.

The Instruction Finetuning Pipeline

Once your dataset is cleaned, the pipeline in ch07/01_main-chapter-code/gpt_instruction_finetuning.py orchestrates the complete training workflow. This script handles data loading, tokenization, model initialization, and supervised training.

Data Loading and Tokenization

The pipeline uses OpenAI's tiktoken library with the GPT-2 encoding. The custom_collate_fn (lines 56-96 in the training script) handles batching by padding sequences to equal length and masking padding tokens with ignore_index = -100 in the loss computation:

import tiktoken
from functools import partial

tokenizer = tiktoken.get_encoding("gpt2")

# Create collate function that pads batches and masks padding tokens

customized_collate_fn = partial(
    custom_collate_fn, device=device, allowed_max_length=1024
)

Model Configuration Options

The repository supports two configuration modes:

  • Test Mode: A tiny 12-dimensional GPT model (configurable in lines 24-38) that trains quickly on CPU for debugging.
  • Full GPT-2: Any standard GPT-2 variant (124M to 1.5B parameters) with pretrained weights downloaded via download_and_load_gpt2() from gpt_download.py.

Training Loop Implementation

The train_model_simple function (imported from previous_chapters.py) executes the supervised finetuning:

optimizer = torch.optim.AdamW(model.parameters(), lr=5e-5, weight_decay=0.1)

train_losses, val_losses, tokens_seen = train_model_simple(
    model, train_loader, val_loader, optimizer, device,
    num_epochs=2, 
    eval_freq=5, 
    eval_iter=5,
    start_context=format_input(val_data[0]), 
    tokenizer=tokenizer
)

This function automatically handles validation evaluation every 5 steps and generates sample completions to monitor training progress.

Generating and Saving Responses

After finetuning, the script generates responses for the test set using the generate function and stores results:

for i, entry in enumerate(test_data):
    input_text = format_input(entry)
    token_ids = generate(
        model=model,
        idx=text_to_token_ids(input_text, tokenizer).to(device),
        max_new_tokens=256,
        context_size=BASE_CONFIG["context_length"],
        eos_id=50256  # GPT-2 end-of-sequence token

    )
    generated_text = token_ids_to_text(token_ids, tokenizer)
    response_text = generated_text[len(input_text):].replace("### Response:", "").strip()

    test_data[i]["model_response"] = response_text

The final checkpoint is saved using torch.save(model.state_dict(), filename) with the naming convention {MODEL_NAME}-sft-standalone.pth.

Step-by-Step Implementation

End-to-End Finetuning Workflow

Execute the complete pipeline from the repository root:

cd ch07/01_main-chapter-code

# Quick test run with minimal data and tiny model

python gpt_instruction_finetuning.py --test_mode

# Full training with GPT-2 medium

python gpt_instruction_finetuning.py

Loading a Finetuned Model for Inference

To use your finetuned model after training:

import torch
from pkg.llms_from_scratch.ch05 import GPTModel

# Must match training configuration

config = {
    "vocab_size": 50257,
    "context_length": 1024,
    "drop_rate": 0.0,
    "qkv_bias": True,
    "emb_dim": 768,
    "n_layers": 12,
    "n_heads": 12,
}

model = GPTModel(config)
model.load_state_dict(torch.load("gpt2-medium-sft-standalone.pth"))
model.eval()

# Inference

import tiktoken
enc = tiktoken.get_encoding("gpt2")
prompt = "Below is an instruction that describes a task...\n\n### Instruction:\nSummarize the following text\n\n### Input:\nMachine learning is...\n\n### Response:\n"

input_ids = torch.tensor(enc.encode(prompt)).unsqueeze(0)

output_ids = generate(
    model=model,
    idx=input_ids,
    max_new_tokens=200,
    context_size=config["context_length"],
    eos_id=50256
)
print(enc.decode(output_ids.squeeze().tolist()))

Summary

  • Dataset Format: Create JSON files with instruction, input, and output fields to define supervised training targets.
  • Data Cleaning: Use find-near-duplicates.py with TF-IDF vectorization to remove duplicate entries with configurable similarity thresholds (default 0.85).
  • Training Pipeline: The gpt_instruction_finetuning.py script handles tokenization (via tiktoken), batching with padding masks (ignore_index=-100), and supervised training using train_model_simple.
  • Model Flexibility: Supports both tiny 12-dimensional test models for debugging and full GPT-2 variants for production finetuning.
  • Artifacts: The pipeline produces JSON files containing model responses and .pth checkpoint files compatible with the GPTModel class defined in pkg/llms_from_scratch/ch05.py.

Frequently Asked Questions

What is the exact JSON schema required for instruction finetuning?

The repository expects a list of JSON objects where each object contains three string fields: instruction (the task description), input (optional context or user query), and output (the desired response). The input field can be an empty string if the instruction contains all necessary context. This schema is defined in ch07/01_main-chapter-code/instruction-data.json and parsed by the format_input function in the training script.

How do I prevent data leakage between training and validation sets?

Use the find_print_and_remove_near_duplicates utility from ch07/02_dataset-utilities/find-near-duplicates.py before splitting your data. This function calculates TF-IDF vectors on character-level n-grams and removes entries with similarity scores above your chosen threshold (typically 0.85) to ensure no near-duplicates exist across splits.

Can I finetune on aCPU-only machine?

Yes. The repository includes a test_mode flag in gpt_instruction_finetuning.py that instantiates a minimal 12-dimensional GPT model instead of the full GPT-2 architecture. This test configuration trains in minutes on CPU and validates the entire pipeline before committing to GPU resources for larger models.

What token is used to indicate the end of generation?

The pipeline uses GPT-2's end-of-sequence token ID 50256 (referenced as eos_id in the generate function calls). During training, custom_collate_fn appends this token to each sequence, and during inference, generation stops when this token is produced or when max_new_tokens is reached.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →