Creating Instruction Finetuning Datasets from Scratch: A Complete Guide to the LLMs-from-scratch Pipeline
You can build a custom instruction-finetuning dataset from scratch by formatting your data as JSON objects with instruction, input, and output fields, cleaning duplicates using TF-IDF similarity detection, and feeding the result into a lightweight GPT training pipeline.
The LLMs-from-scratch repository by Sebastian Raschka provides an end-to-end educational framework for creating, cleaning, and training on instruction-finetuning datasets. Unlike black-box APIs, this implementation exposes every step—from JSON data preparation to the supervised finetuning loop—making it ideal for researchers who need full control over their training data and model behavior.
Dataset Format and Structure
The foundation of instruction finetuning is a properly structured JSON dataset. According to the reference implementation in ch07/01_main-chapter-code/instruction-data.json, each entry must contain three specific fields:
{
"instruction": "Describe the process of photosynthesis in simple terms.",
"input": "",
"output": "Photosynthesis is the process by which plants use sunlight, water, and carbon dioxide to create oxygen and energy in the form of sugar."
}
instruction: The task description or command for the model.input: Optional context or auxiliary information (can be an empty string).output: The target response that the model should learn to generate.
The repository includes a reference dataset containing approximately 1,100 such examples, providing a template for creating custom datasets for specific domains or languages.
Cleaning and Deduplicating Your Dataset
Raw scraped or crowdsourced data often contains near-duplicate entries that can bias training. The repository ships with ch07/02_dataset-utilities/find-near-duplicates.py, a utility that uses scikit-learn's TF-IDF vectorizer with character-level n-grams to detect semantic duplicates.
Removing Near-Duplicates with TF-IDF
The find_print_and_remove_near_duplicates function computes similarity scores between entries and optionally prunes them based on a configurable threshold:
from find_near_duplicates import find_print_and_remove_near_duplicates
# Load your raw dataset
with open("my_instruction_dataset.json", "r") as f:
data = json.load(f)
# Detect and remove duplicates with 85% similarity threshold
clean_data = find_print_and_remove_near_duplicates(
json_data=data,
remove_duplicates=True,
threshold=0.85
)
# Save cleaned version
with open("my_instruction_dataset_clean.json", "w") as f:
json.dump(clean_data, f, indent=4)
Key implementation details from the source code:
- Uses character-level n-grams (
ngram_range=(1,3)) to catch subtle phrasing variations. - By default checks the
inputandoutputfields for duplicates while preserving the instruction text (lines 66-69 offind-near-duplicates.py). - The threshold parameter accepts values between 0 and 1, where 0.85 provides a conservative balance between removing true duplicates and preserving valid variations.
The Instruction Finetuning Pipeline
Once your dataset is cleaned, the pipeline in ch07/01_main-chapter-code/gpt_instruction_finetuning.py orchestrates the complete training workflow. This script handles data loading, tokenization, model initialization, and supervised training.
Data Loading and Tokenization
The pipeline uses OpenAI's tiktoken library with the GPT-2 encoding. The custom_collate_fn (lines 56-96 in the training script) handles batching by padding sequences to equal length and masking padding tokens with ignore_index = -100 in the loss computation:
import tiktoken
from functools import partial
tokenizer = tiktoken.get_encoding("gpt2")
# Create collate function that pads batches and masks padding tokens
customized_collate_fn = partial(
custom_collate_fn, device=device, allowed_max_length=1024
)
Model Configuration Options
The repository supports two configuration modes:
- Test Mode: A tiny 12-dimensional GPT model (configurable in lines 24-38) that trains quickly on CPU for debugging.
- Full GPT-2: Any standard GPT-2 variant (124M to 1.5B parameters) with pretrained weights downloaded via
download_and_load_gpt2()fromgpt_download.py.
Training Loop Implementation
The train_model_simple function (imported from previous_chapters.py) executes the supervised finetuning:
optimizer = torch.optim.AdamW(model.parameters(), lr=5e-5, weight_decay=0.1)
train_losses, val_losses, tokens_seen = train_model_simple(
model, train_loader, val_loader, optimizer, device,
num_epochs=2,
eval_freq=5,
eval_iter=5,
start_context=format_input(val_data[0]),
tokenizer=tokenizer
)
This function automatically handles validation evaluation every 5 steps and generates sample completions to monitor training progress.
Generating and Saving Responses
After finetuning, the script generates responses for the test set using the generate function and stores results:
for i, entry in enumerate(test_data):
input_text = format_input(entry)
token_ids = generate(
model=model,
idx=text_to_token_ids(input_text, tokenizer).to(device),
max_new_tokens=256,
context_size=BASE_CONFIG["context_length"],
eos_id=50256 # GPT-2 end-of-sequence token
)
generated_text = token_ids_to_text(token_ids, tokenizer)
response_text = generated_text[len(input_text):].replace("### Response:", "").strip()
test_data[i]["model_response"] = response_text
The final checkpoint is saved using torch.save(model.state_dict(), filename) with the naming convention {MODEL_NAME}-sft-standalone.pth.
Step-by-Step Implementation
End-to-End Finetuning Workflow
Execute the complete pipeline from the repository root:
cd ch07/01_main-chapter-code
# Quick test run with minimal data and tiny model
python gpt_instruction_finetuning.py --test_mode
# Full training with GPT-2 medium
python gpt_instruction_finetuning.py
Loading a Finetuned Model for Inference
To use your finetuned model after training:
import torch
from pkg.llms_from_scratch.ch05 import GPTModel
# Must match training configuration
config = {
"vocab_size": 50257,
"context_length": 1024,
"drop_rate": 0.0,
"qkv_bias": True,
"emb_dim": 768,
"n_layers": 12,
"n_heads": 12,
}
model = GPTModel(config)
model.load_state_dict(torch.load("gpt2-medium-sft-standalone.pth"))
model.eval()
# Inference
import tiktoken
enc = tiktoken.get_encoding("gpt2")
prompt = "Below is an instruction that describes a task...\n\n### Instruction:\nSummarize the following text\n\n### Input:\nMachine learning is...\n\n### Response:\n"
input_ids = torch.tensor(enc.encode(prompt)).unsqueeze(0)
output_ids = generate(
model=model,
idx=input_ids,
max_new_tokens=200,
context_size=config["context_length"],
eos_id=50256
)
print(enc.decode(output_ids.squeeze().tolist()))
Summary
- Dataset Format: Create JSON files with
instruction,input, andoutputfields to define supervised training targets. - Data Cleaning: Use
find-near-duplicates.pywith TF-IDF vectorization to remove duplicate entries with configurable similarity thresholds (default 0.85). - Training Pipeline: The
gpt_instruction_finetuning.pyscript handles tokenization (viatiktoken), batching with padding masks (ignore_index=-100), and supervised training usingtrain_model_simple. - Model Flexibility: Supports both tiny 12-dimensional test models for debugging and full GPT-2 variants for production finetuning.
- Artifacts: The pipeline produces JSON files containing model responses and
.pthcheckpoint files compatible with theGPTModelclass defined inpkg/llms_from_scratch/ch05.py.
Frequently Asked Questions
What is the exact JSON schema required for instruction finetuning?
The repository expects a list of JSON objects where each object contains three string fields: instruction (the task description), input (optional context or user query), and output (the desired response). The input field can be an empty string if the instruction contains all necessary context. This schema is defined in ch07/01_main-chapter-code/instruction-data.json and parsed by the format_input function in the training script.
How do I prevent data leakage between training and validation sets?
Use the find_print_and_remove_near_duplicates utility from ch07/02_dataset-utilities/find-near-duplicates.py before splitting your data. This function calculates TF-IDF vectors on character-level n-grams and removes entries with similarity scores above your chosen threshold (typically 0.85) to ensure no near-duplicates exist across splits.
Can I finetune on aCPU-only machine?
Yes. The repository includes a test_mode flag in gpt_instruction_finetuning.py that instantiates a minimal 12-dimensional GPT model instead of the full GPT-2 architecture. This test configuration trains in minutes on CPU and validates the entire pipeline before committing to GPU resources for larger models.
What token is used to indicate the end of generation?
The pipeline uses GPT-2's end-of-sequence token ID 50256 (referenced as eos_id in the generate function calls). During training, custom_collate_fn appends this token to each sequence, and during inference, generation stops when this token is produced or when max_new_tokens is reached.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →