Understanding the Difference Between Pretraining and Instruction Finetuning
Pretraining teaches a model to predict the next token from raw text using self-supervision, while instruction finetuning adapts that pretrained model to follow specific user commands by training on prompt-response pairs with loss masking on prompts.
Large language models are developed in two distinct phases: first building general language capabilities through pretraining, then specializing those capabilities via instruction finetuning. The rasbt/LLMs-from-scratch repository implements both stages using the same underlying transformer architecture, yet they differ fundamentally in data formatting, objectives, and loss computation.
What Is Pretraining?
Pretraining is a self-supervised learning phase where the model learns statistical patterns by predicting the next token in sequences of raw text. This stage equips the network with syntax, semantics, and broad world knowledge without requiring explicit labels.
The Self-Supervised Objective
The training objective maximizes the likelihood of the next token given previous tokens, learning the conditional probability distribution (P(\text{next token} \mid \text{previous tokens})). Every token in the training corpus contributes to the loss function, allowing the model to capture general linguistic patterns from massive text sources like Project Gutenberg.
Implementation in pretraining_simple.py
In ch05/03_bonus_pretraining_on_gutenberg/pretraining_simple.py, the pipeline streams raw books and tokenizes them using tiktoken.get_encoding("gpt2"). The script calls create_dataloader_v1 (from pkg/llms_from_scratch/ch02.py) to construct overlapping text windows of length context_length. The training loop uses calc_loss_batch to compute gradients across every token in the sequence, with no masking applied. This script trains the same GPTModel class defined in ch04/gpt_with_kv_gqa.py or ch04/gpt_with_kv_mha.py from scratch on unstructured text.
What Is Instruction Finetuning?
Instruction finetuning is a supervised learning stage that transforms a generic pretrained model into an assistant capable of following explicit human instructions. Unlike pretraining, this phase uses structured datasets containing specific tasks and desired outputs.
Supervised Learning on Structured Data
The dataset consists of JSON entries formatted as instruction → input → desired output. For example, an entry might instruct the model to translate a sentence or summarize a paragraph, with the corresponding correct response provided as the target. The model learns to map these formatted prompts to appropriate completions.
Masked Loss and the InstructionDataset
In ch07/01_main-chapter-code/gpt_instruction_finetuning.py, the InstructionDataset class formats each entry into a prompt-response string. The critical implementation detail lies in the custom_collate_fn: it applies ignore_index to mask all tokens belonging to the prompt portion, ensuring that loss calculation and gradient updates occur only for the response tokens. This prevents the model from wasting capacity learning to reproduce the instructions verbatim.
Key Differences: A Technical Comparison
Both phases share identical GPTModel weights and transformer blocks, but differ in four critical technical aspects:
Data Source: Pretraining consumes raw text corpora (books, web text) without explicit labels. Instruction finetuning requires structured JSON files containing instruction-response pairs.
Loss Calculation: Pretraining computes next-token prediction loss across the entire sequence. Instruction finetuning masks prompt tokens using ignore_index, calculating loss only on response segments.
Training Objective: Pretraining optimizes for general language modeling—learning syntax and world knowledge. Instruction finetuning optimizes for task adherence—learning to map specific prompts to desired outputs.
Data Pipeline: Pretraining uses sliding window chunks via create_dataloader_v1. Instruction finetuning uses InstructionDataset with custom_collate_fn to handle variable-length prompts and apply padding masks.
Code Examples: From Raw Text to Instruction Following
Running the Pretraining Pipeline
To pretrain a small GPT-2 model on raw text from Project Gutenberg:
python ch05/03_bonus_pretraining_on_gutenberg/pretraining_simple.py \
--data_dir gutenberg/data \
--output_dir model_checkpoints \
--n_epochs 1 \
--batch_size 4 \
--print_sample_iter 1000 \
--debug
This script (lines 14-25) initializes the tokenizer and runs a standard language modeling loop where calc_loss_batch processes every token without masking.
Executing Instruction Finetuning
To finetune the pretrained weights on instruction data:
python ch07/01_main-chapter-code/gpt_instruction_finetuning.py \
--test_mode
As implemented in lines 35-53 and 221-280, this script loads pretrained weights via gpt_download.py, initializes the InstructionDataset, and uses custom_collate_fn to mask prompt tokens. The AdamW optimizer updates the same GPTModel weights, but gradients propagate only from the response segments.
Inference After Finetuning
Load a finetuned checkpoint and generate instruction-following responses:
# Load fine-tuned checkpoint
model = GPTModel(BASE_CONFIG).to(device)
model.load_state_dict(torch.load("gpt2-medium-sft-standalone.pth"))
model.eval()
# Format instruction prompt
prompt = format_input({
"instruction": "Summarize the following paragraph.",
"input": "Artificial intelligence ...",
"output": ""
})
# Generate response
token_ids = generate(
model=model,
idx=text_to_token_ids(prompt, tokenizer).to(device),
max_new_tokens=200,
context_size=BASE_CONFIG["context_length"],
eos_id=50256,
)
response = token_ids_to_text(token_ids, tokenizer)[len(prompt):]
print(response)
The generate function—identical to that used during pretraining—now produces contextually appropriate responses because the weights have been adapted via instruction finetuning.
Summary
- Pretraining is a self-supervised phase using raw text where the model learns next-token prediction across all tokens, implemented in
pretraining_simple.py. - Instruction finetuning is a supervised phase using structured prompt-response pairs where loss is masked on prompts using
ignore_index, implemented ingpt_instruction_finetuning.py. - Both stages share the same
GPTModelarchitecture fromch04/gpt_with_kv_gqa.pybut differ in data formatting, loss masking, and training objectives. - The repository provides modular scripts allowing seamless transition from general language modeling to task-specific instruction following.
Frequently Asked Questions
Do I need to pretrain a model before instruction finetuning?
Yes, typically you begin with a pretrained checkpoint—either trained from scratch using pretraining_simple.py or downloaded via gpt_download.py. Instruction finetuning assumes the model already possesses general language understanding and simply adapts it to follow specific formats and instructions.
Why mask the prompt tokens during instruction finetuning?
Masking prompt tokens with ignore_index prevents the loss function from penalizing the model for not exactly reproducing the instruction text. This focuses gradient updates solely on learning to generate appropriate responses, making training more efficient and preventing the model from simply learning to copy input prompts.
Can I use the same GPTModel class for both phases?
Absolutely. The rasbt/LLMs-from-scratch repository uses the identical GPTModel class (defined in ch04/gpt_with_kv_gqa.py or ch04/gpt_with_kv_mha.py) for both pretraining and instruction finetuning. The architecture remains constant; only the data pipeline and loss computation logic differ between stages.
How does the data format differ between the two stages?
Pretraining consumes raw text files (like books from Project Gutenberg) with no structure, processed into sliding windows of fixed context_length. Instruction finetuning requires JSON entries containing explicit fields for instructions, inputs, and desired outputs, which the InstructionDataset class formats into unified prompt-response strings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →