# Understanding the Difference Between Pretraining and Instruction Finetuning

> Learn the crucial difference between pretraining and instruction finetuning. Discover how models learn from raw text and adapt to user commands. LLMs from scratch explained.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: deep-dive
- Published: 2026-05-12

---

**Pretraining teaches a model to predict the next token from raw text using self-supervision, while instruction finetuning adapts that pretrained model to follow specific user commands by training on prompt-response pairs with loss masking on prompts.**

Large language models are developed in two distinct phases: first building general language capabilities through pretraining, then specializing those capabilities via instruction finetuning. The `rasbt/LLMs-from-scratch` repository implements both stages using the same underlying transformer architecture, yet they differ fundamentally in data formatting, objectives, and loss computation.

## What Is Pretraining?

Pretraining is a **self-supervised learning** phase where the model learns statistical patterns by predicting the next token in sequences of raw text. This stage equips the network with syntax, semantics, and broad world knowledge without requiring explicit labels.

### The Self-Supervised Objective

The training objective maximizes the likelihood of the next token given previous tokens, learning the conditional probability distribution \(P(\text{next token} \mid \text{previous tokens})\). Every token in the training corpus contributes to the loss function, allowing the model to capture general linguistic patterns from massive text sources like Project Gutenberg.

### Implementation in pretraining_simple.py

In [`ch05/03_bonus_pretraining_on_gutenberg/pretraining_simple.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05/03_bonus_pretraining_on_gutenberg/pretraining_simple.py), the pipeline streams raw books and tokenizes them using `tiktoken.get_encoding("gpt2")`. The script calls `create_dataloader_v1` (from [`pkg/llms_from_scratch/ch02.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch02.py)) to construct overlapping text windows of length `context_length`. The training loop uses `calc_loss_batch` to compute gradients across **every** token in the sequence, with no masking applied. This script trains the same `GPTModel` class defined in [`ch04/gpt_with_kv_gqa.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/gpt_with_kv_gqa.py) or [`ch04/gpt_with_kv_mha.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/gpt_with_kv_mha.py) from scratch on unstructured text.

## What Is Instruction Finetuning?

Instruction finetuning is a **supervised learning** stage that transforms a generic pretrained model into an assistant capable of following explicit human instructions. Unlike pretraining, this phase uses structured datasets containing specific tasks and desired outputs.

### Supervised Learning on Structured Data

The dataset consists of JSON entries formatted as **instruction → input → desired output**. For example, an entry might instruct the model to translate a sentence or summarize a paragraph, with the corresponding correct response provided as the target. The model learns to map these formatted prompts to appropriate completions.

### Masked Loss and the InstructionDataset

In [`ch07/01_main-chapter-code/gpt_instruction_finetuning.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/gpt_instruction_finetuning.py), the `InstructionDataset` class formats each entry into a prompt-response string. The critical implementation detail lies in the `custom_collate_fn`: it applies `ignore_index` to mask all tokens belonging to the prompt portion, ensuring that loss calculation and gradient updates occur only for the response tokens. This prevents the model from wasting capacity learning to reproduce the instructions verbatim.

## Key Differences: A Technical Comparison

Both phases share identical `GPTModel` weights and transformer blocks, but differ in four critical technical aspects:

**Data Source:** Pretraining consumes raw text corpora (books, web text) without explicit labels. Instruction finetuning requires structured JSON files containing instruction-response pairs.

**Loss Calculation:** Pretraining computes next-token prediction loss across the entire sequence. Instruction finetuning masks prompt tokens using `ignore_index`, calculating loss only on response segments.

**Training Objective:** Pretraining optimizes for general language modeling—learning syntax and world knowledge. Instruction finetuning optimizes for task adherence—learning to map specific prompts to desired outputs.

**Data Pipeline:** Pretraining uses sliding window chunks via `create_dataloader_v1`. Instruction finetuning uses `InstructionDataset` with `custom_collate_fn` to handle variable-length prompts and apply padding masks.

## Code Examples: From Raw Text to Instruction Following

### Running the Pretraining Pipeline

To pretrain a small GPT-2 model on raw text from Project Gutenberg:

```bash
python ch05/03_bonus_pretraining_on_gutenberg/pretraining_simple.py \
    --data_dir gutenberg/data \
    --output_dir model_checkpoints \
    --n_epochs 1 \
    --batch_size 4 \
    --print_sample_iter 1000 \
    --debug

```

This script (lines 14-25) initializes the tokenizer and runs a standard language modeling loop where `calc_loss_batch` processes **every** token without masking.

### Executing Instruction Finetuning

To finetune the pretrained weights on instruction data:

```bash
python ch07/01_main-chapter-code/gpt_instruction_finetuning.py \
    --test_mode

```

As implemented in lines 35-53 and 221-280, this script loads pretrained weights via [`gpt_download.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/gpt_download.py), initializes the `InstructionDataset`, and uses `custom_collate_fn` to mask prompt tokens. The `AdamW` optimizer updates the same `GPTModel` weights, but gradients propagate only from the response segments.

### Inference After Finetuning

Load a finetuned checkpoint and generate instruction-following responses:

```python

# Load fine-tuned checkpoint

model = GPTModel(BASE_CONFIG).to(device)
model.load_state_dict(torch.load("gpt2-medium-sft-standalone.pth"))
model.eval()

# Format instruction prompt

prompt = format_input({
    "instruction": "Summarize the following paragraph.",
    "input": "Artificial intelligence ...",
    "output": ""
})

# Generate response

token_ids = generate(
    model=model,
    idx=text_to_token_ids(prompt, tokenizer).to(device),
    max_new_tokens=200,
    context_size=BASE_CONFIG["context_length"],
    eos_id=50256,
)
response = token_ids_to_text(token_ids, tokenizer)[len(prompt):]
print(response)

```

The `generate` function—identical to that used during pretraining—now produces contextually appropriate responses because the weights have been adapted via instruction finetuning.

## Summary

- **Pretraining** is a self-supervised phase using raw text where the model learns next-token prediction across all tokens, implemented in [`pretraining_simple.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pretraining_simple.py).
- **Instruction finetuning** is a supervised phase using structured prompt-response pairs where loss is masked on prompts using `ignore_index`, implemented in [`gpt_instruction_finetuning.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/gpt_instruction_finetuning.py).
- Both stages share the same `GPTModel` architecture from [`ch04/gpt_with_kv_gqa.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/gpt_with_kv_gqa.py) but differ in data formatting, loss masking, and training objectives.
- The repository provides modular scripts allowing seamless transition from general language modeling to task-specific instruction following.

## Frequently Asked Questions

### Do I need to pretrain a model before instruction finetuning?

Yes, typically you begin with a pretrained checkpoint—either trained from scratch using [`pretraining_simple.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pretraining_simple.py) or downloaded via [`gpt_download.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/gpt_download.py). Instruction finetuning assumes the model already possesses general language understanding and simply adapts it to follow specific formats and instructions.

### Why mask the prompt tokens during instruction finetuning?

Masking prompt tokens with `ignore_index` prevents the loss function from penalizing the model for not exactly reproducing the instruction text. This focuses gradient updates solely on learning to generate appropriate responses, making training more efficient and preventing the model from simply learning to copy input prompts.

### Can I use the same GPTModel class for both phases?

Absolutely. The `rasbt/LLMs-from-scratch` repository uses the identical `GPTModel` class (defined in [`ch04/gpt_with_kv_gqa.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/gpt_with_kv_gqa.py) or [`ch04/gpt_with_kv_mha.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/gpt_with_kv_mha.py)) for both pretraining and instruction finetuning. The architecture remains constant; only the data pipeline and loss computation logic differ between stages.

### How does the data format differ between the two stages?

Pretraining consumes raw text files (like books from Project Gutenberg) with no structure, processed into sliding windows of fixed `context_length`. Instruction finetuning requires JSON entries containing explicit fields for instructions, inputs, and desired outputs, which the `InstructionDataset` class formats into unified prompt-response strings.