# The Three Phases of Agent Model Post‑Training: Pre‑training, SFT, and RL Explained

> Master agent model post-training. Understand the three key phases: pre-training for language priors, SFT for instruction alignment, and RL for reward optimization and tool use.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-06

---

**The three phases of agent model post‑training are: (1) Pre‑training to learn broad language priors, (2) Supervised Fine‑Tuning (SFT) to align with instructions, and (3) Reinforcement Learning (RL) to optimize for rewards and enable tool use.**

Modern AI agents powering systems like GPT‑4, Claude, and open‑source alternatives follow a structured training pipeline. According to the `bojieli/ai-agent-book` repository, this pipeline separates generic knowledge acquisition from behavioral alignment and reward optimization. Understanding these distinct phases helps practitioners debug failures, select appropriate datasets, and design effective fine‑tuning strategies.

---

## Phase 1: Pre‑training – Building the Foundation

Pre‑training establishes the **broad language (or vision‑language) prior** that all subsequent phases rely upon. The model learns to predict masked or next tokens on massive, unlabeled corpora.

### How Pre‑training Works

- **Data source:** Large text dumps (Common Crawl, Wikipedia) or multimodal datasets
- **Objective:** Unsupervised language modeling loss
- **Outcome:** Statistical regularities, factual knowledge, and basic reasoning capabilities

In [`book-vi/glossary.vi.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-vi/glossary.vi.md), the repository defines *post‑training* explicitly in contrast to this pre‑training step, establishing pre‑training as the foundational phase that precedes all agent‑specific development. For domain‑specific agents, the book describes **continued pre‑training** in [`chapter7/continued-pretraining/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/continued-pretraining/README.md) — a middle ground where the model undergoes additional unsupervised training on specialized corpora before instruction tuning.

---

## Phase 2: Supervised Fine‑Tuning (SFT) – Instruction Alignment

SFT transforms a raw language model into an **instruction‑following assistant**. This phase shapes behavior without fundamentally altering the underlying knowledge base.

### SFT Implementation Details

| Aspect | Description |
|--------|-------------|
| **Typical datasets** | Alpaca, ShareGPT, or custom prompt‑response pairs |
| **Loss function** | Cross‑entropy between model outputs and target responses |
| **Key outcome** | Proper formatting, helpful tone, task comprehension |

The repository emphasizes SFT as a distinct **post‑training step** in [`book-vi/glossary.vi.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-vi/glossary.vi.md), separate from both pre‑training and RL. A cursor‑chat note in `cursor-chats/20251010_113500_continue_pretraining,_belong_to_pretrain_or_posttrain.md` explicitly treats SFT as the bridge between generic pre‑trained models and specialized agent capabilities.

### Minimal SFT Code Example

```python
from transformers import AutoModelForCausalLM, AutoTokenizer, Trainer, TrainingArguments

# Load pre‑trained checkpoint

model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-v0.3")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.3")

# Prepare instruction dataset: {"prompt": ..., "response": ...}

train_dataset = ...  

training_args = TrainingArguments(
    output_dir="sft_model",
    per_device_train_batch_size=4,
    num_train_epochs=2,
    learning_rate=5e-5,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset
)
trainer.train()  # Minimize cross‑entropy on instruction outputs

```

---

## Phase 3: Reinforcement Learning (RL) – Reward Optimization

RL completes the pipeline by **internalizing long‑term goals, tool use, and safety constraints** through reward signals rather than explicit demonstrations.

### RL Training Mechanics

- **Reward source:** Human preference rankings or task‑specific success metrics
- **Algorithms:** Policy‑gradient methods (PPO), Direct Preference Optimization (DPO)
- **Capabilities unlocked:** Multi‑step planning, API calling, refusal of harmful requests

The repository's [`chapter7/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/README.md) overview positions RL as the final refinement stage after SFT. The **MiniMind experiments** documented in [`chapter7/MiniMind-pretrain/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/MiniMind-pretrain/README.md) demonstrate this full three‑stage flow: continued pre‑training on Korean text, followed by SFT, then RL via DPO to optimize against human preferences.

### RL Fine‑Tuning with PPO

```python
from trl import PPOTrainer, PPOConfig

# Initialize from SFT checkpoint

ppo_cfg = PPOConfig(
    batch_size=4,
    ppo_epochs=4,
    learning_rate=1e-5,
)

ppo_trainer = PPOTrainer(
    model=trainer.model,  # SFT‑fine‑tuned checkpoint

    tokenizer=tokenizer,
    reward_model=reward_model,  # Trained on human rankings

    config=ppo_cfg,
)

# Training loop: generate, score, update

for epoch in range(3):
    prompts = ["Summarise the news:", "Translate to French:"]
    responses = ppo_trainer.generate(prompts, max_new_tokens=64)
    rewards = reward_model(prompts, responses)
    ppo_trainer.step(prompts, responses, rewards)

ppo_trainer.model.save_pretrained("final_rl_model")

```

---

## How the Three Phases Connect

The `bojieli/ai-agent-book` repository structures experiments around this progression:

1. **Pre‑training** → Generic capabilities, transferable knowledge
2. **SFT** → Behavior shaping, instruction following
3. **RL** → Reward maximization, tool integration, safety

Chapter 7's **continued‑pretraining** scripts ([`continued-pretrain.py`](https://github.com/bojieli/ai-agent-book/blob/main/continued-pretrain.py), [`compare_models.py`](https://github.com/bojieli/ai-agent-book/blob/main/compare_models.py)) implement this with **LoRA adapters**, allowing efficient parameter updates across all three phases without full model retraining.

---

## Phase Comparison: When to Use Each

| Phase | Use When | Avoid When |
|-------|----------|------------|
| **Pre‑training / continued pre‑training** | Domain shift, new languages, specialized knowledge gaps | You need immediate instruction following |
| **SFT** | Task definition is clear, demonstrations are available | Reward criteria are implicit or multi‑factorial |
| **RL** | Long‑horizon planning, trade‑offs between criteria, safety constraints | You lack reliable reward signal or preference data |

---

## Summary

- **Pre‑training** learns broad priors through unsupervised prediction on massive corpora, as defined in [`book-vi/glossary.vi.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-vi/glossary.vi.md)
- **SFT** aligns models to instructions via supervised cross‑entropy loss on prompt‑response pairs
- **RL** refines behavior through reward optimization, enabling tool use and complex planning as shown in the MiniMind experiments
- The complete **pre‑training → SFT → RL pipeline** appears throughout `chapter7/` with concrete implementations using LoRA and HuggingFace/TRL libraries

---

## Frequently Asked Questions

### Does pre‑training belong to post‑training?

No. Pre‑training is the **foundational stage that precedes post‑training**. According to `cursor-chats/20251010_113500_continue_pretraining,_belong_to_pretrain_or_posttrain.md`, the term *post‑training* specifically encompasses SFT and RL phases that occur after the initial pre‑training. However, **continued pre‑training** on domain‑specific data can blur this boundary when performed as part of an agent development pipeline.

### Can I skip SFT and train RL directly from a pre‑trained model?

Technically possible, but practically inadvisable. The repository notes in [`chapter7/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/README.md) that SFT provides **essential behavioral scaffolding** — output formatting, basic instruction comprehension — that makes RL training stable. Without SFT, RL exploration becomes inefficient and the policy may fail to generate coherent responses worth rewarding.

### What RL algorithms does the ai‑agent‑book repository use?

The **MiniMind experiments** in [`chapter7/MiniMind-pretrain/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/MiniMind-pretrain/README.md) demonstrate **Direct Preference Optimization (DPO)**, a simpler alternative to PPO that directly optimizes from preference pairs without explicit reward modeling. The code examples also reference **PPO** (Proximal Policy Optimization) from the TRL library for full policy‑gradient training.

### How do I know which phase is causing my agent's failure?

Use **probing evaluation** at each checkpoint:
- Factual errors → Pre‑training/continued pre‑training issue
- Ignores instructions, wrong format → SFT problem
- Poor trade‑offs, unsafe outputs, no tool use → RL refinement needed

The repository's [`compare_models.py`](https://github.com/bojieli/ai-agent-book/blob/main/compare_models.py) script implements exactly this diagnostic workflow across the three phases.