The Three Phases of Agent Model Post‑Training: Pre‑training, SFT, and RL Explained

The three phases of agent model post‑training are: (1) Pre‑training to learn broad language priors, (2) Supervised Fine‑Tuning (SFT) to align with instructions, and (3) Reinforcement Learning (RL) to optimize for rewards and enable tool use.

Modern AI agents powering systems like GPT‑4, Claude, and open‑source alternatives follow a structured training pipeline. According to the bojieli/ai-agent-book repository, this pipeline separates generic knowledge acquisition from behavioral alignment and reward optimization. Understanding these distinct phases helps practitioners debug failures, select appropriate datasets, and design effective fine‑tuning strategies.


Phase 1: Pre‑training – Building the Foundation

Pre‑training establishes the broad language (or vision‑language) prior that all subsequent phases rely upon. The model learns to predict masked or next tokens on massive, unlabeled corpora.

How Pre‑training Works

  • Data source: Large text dumps (Common Crawl, Wikipedia) or multimodal datasets
  • Objective: Unsupervised language modeling loss
  • Outcome: Statistical regularities, factual knowledge, and basic reasoning capabilities

In book-vi/glossary.vi.md, the repository defines post‑training explicitly in contrast to this pre‑training step, establishing pre‑training as the foundational phase that precedes all agent‑specific development. For domain‑specific agents, the book describes continued pre‑training in chapter7/continued-pretraining/README.md — a middle ground where the model undergoes additional unsupervised training on specialized corpora before instruction tuning.


Phase 2: Supervised Fine‑Tuning (SFT) – Instruction Alignment

SFT transforms a raw language model into an instruction‑following assistant. This phase shapes behavior without fundamentally altering the underlying knowledge base.

SFT Implementation Details

Aspect Description
Typical datasets Alpaca, ShareGPT, or custom prompt‑response pairs
Loss function Cross‑entropy between model outputs and target responses
Key outcome Proper formatting, helpful tone, task comprehension

The repository emphasizes SFT as a distinct post‑training step in book-vi/glossary.vi.md, separate from both pre‑training and RL. A cursor‑chat note in cursor-chats/20251010_113500_continue_pretraining,_belong_to_pretrain_or_posttrain.md explicitly treats SFT as the bridge between generic pre‑trained models and specialized agent capabilities.

Minimal SFT Code Example

from transformers import AutoModelForCausalLM, AutoTokenizer, Trainer, TrainingArguments

# Load pre‑trained checkpoint

model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-v0.3")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.3")

# Prepare instruction dataset: {"prompt": ..., "response": ...}

train_dataset = ...  

training_args = TrainingArguments(
    output_dir="sft_model",
    per_device_train_batch_size=4,
    num_train_epochs=2,
    learning_rate=5e-5,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset
)
trainer.train()  # Minimize cross‑entropy on instruction outputs

Phase 3: Reinforcement Learning (RL) – Reward Optimization

RL completes the pipeline by internalizing long‑term goals, tool use, and safety constraints through reward signals rather than explicit demonstrations.

RL Training Mechanics

  • Reward source: Human preference rankings or task‑specific success metrics
  • Algorithms: Policy‑gradient methods (PPO), Direct Preference Optimization (DPO)
  • Capabilities unlocked: Multi‑step planning, API calling, refusal of harmful requests

The repository's chapter7/README.md overview positions RL as the final refinement stage after SFT. The MiniMind experiments documented in chapter7/MiniMind-pretrain/README.md demonstrate this full three‑stage flow: continued pre‑training on Korean text, followed by SFT, then RL via DPO to optimize against human preferences.

RL Fine‑Tuning with PPO

from trl import PPOTrainer, PPOConfig

# Initialize from SFT checkpoint

ppo_cfg = PPOConfig(
    batch_size=4,
    ppo_epochs=4,
    learning_rate=1e-5,
)

ppo_trainer = PPOTrainer(
    model=trainer.model,  # SFT‑fine‑tuned checkpoint

    tokenizer=tokenizer,
    reward_model=reward_model,  # Trained on human rankings

    config=ppo_cfg,
)

# Training loop: generate, score, update

for epoch in range(3):
    prompts = ["Summarise the news:", "Translate to French:"]
    responses = ppo_trainer.generate(prompts, max_new_tokens=64)
    rewards = reward_model(prompts, responses)
    ppo_trainer.step(prompts, responses, rewards)

ppo_trainer.model.save_pretrained("final_rl_model")

How the Three Phases Connect

The bojieli/ai-agent-book repository structures experiments around this progression:

  1. Pre‑training → Generic capabilities, transferable knowledge
  2. SFT → Behavior shaping, instruction following
  3. RL → Reward maximization, tool integration, safety

Chapter 7's continued‑pretraining scripts (continued-pretrain.py, compare_models.py) implement this with LoRA adapters, allowing efficient parameter updates across all three phases without full model retraining.


Phase Comparison: When to Use Each

Phase Use When Avoid When
Pre‑training / continued pre‑training Domain shift, new languages, specialized knowledge gaps You need immediate instruction following
SFT Task definition is clear, demonstrations are available Reward criteria are implicit or multi‑factorial
RL Long‑horizon planning, trade‑offs between criteria, safety constraints You lack reliable reward signal or preference data

Summary

  • Pre‑training learns broad priors through unsupervised prediction on massive corpora, as defined in book-vi/glossary.vi.md
  • SFT aligns models to instructions via supervised cross‑entropy loss on prompt‑response pairs
  • RL refines behavior through reward optimization, enabling tool use and complex planning as shown in the MiniMind experiments
  • The complete pre‑training → SFT → RL pipeline appears throughout chapter7/ with concrete implementations using LoRA and HuggingFace/TRL libraries

Frequently Asked Questions

Does pre‑training belong to post‑training?

No. Pre‑training is the foundational stage that precedes post‑training. According to cursor-chats/20251010_113500_continue_pretraining,_belong_to_pretrain_or_posttrain.md, the term post‑training specifically encompasses SFT and RL phases that occur after the initial pre‑training. However, continued pre‑training on domain‑specific data can blur this boundary when performed as part of an agent development pipeline.

Can I skip SFT and train RL directly from a pre‑trained model?

Technically possible, but practically inadvisable. The repository notes in chapter7/README.md that SFT provides essential behavioral scaffolding — output formatting, basic instruction comprehension — that makes RL training stable. Without SFT, RL exploration becomes inefficient and the policy may fail to generate coherent responses worth rewarding.

What RL algorithms does the ai‑agent‑book repository use?

The MiniMind experiments in chapter7/MiniMind-pretrain/README.md demonstrate Direct Preference Optimization (DPO), a simpler alternative to PPO that directly optimizes from preference pairs without explicit reward modeling. The code examples also reference PPO (Proximal Policy Optimization) from the TRL library for full policy‑gradient training.

How do I know which phase is causing my agent's failure?

Use probing evaluation at each checkpoint:

  • Factual errors → Pre‑training/continued pre‑training issue
  • Ignores instructions, wrong format → SFT problem
  • Poor trade‑offs, unsafe outputs, no tool use → RL refinement needed

The repository's compare_models.py script implements exactly this diagnostic workflow across the three phases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →