# SFT vs RL for Agent Model Training: Key Differences and When to Use Each

> Discover the key differences between SFT and RL for agent model training. Learn when to use SFT for memorization and RL for generalizable strategies.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-06

---

**SFT memorizes fixed input-output mappings from teacher demonstrations, while RL learns generalizable strategies through trial-and-error reward optimization.**

Training large language model (LLM)-based agents typically follows a three-stage pipeline: pre-training, **supervised fine-tuning (SFT)**, and **reinforcement learning (RL)**. According to the *ai-agent-book* source code, each stage serves a distinct purpose with fundamentally different optimization objectives.

## What SFT Does: Learning the Teacher's Map

SFT optimizes for **maximum likelihood estimation**—the model learns to maximize the probability of a labeled answer given an input prompt. During training, the model receives input-output demonstration pairs and computes loss only on the response portion, effectively learning to "predict the next token" within the answer.

This approach yields three characteristic outcomes:

- **Memorization**: The model internalizes exact mappings from prompts to demonstrated answers
- **Format consistency**: Excellent instruction-following and output structure adherence
- **Training efficiency**: Converges in hours to days with thousands of examples

In [`book-en/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter7.md) lines 47-53, the authors emphasize that SFT excels at solidifying the **"form"** of an agent—establishing output style, JSON schemas, and tool-call syntax before any exploration-based learning occurs.

## What RL Does: Exploring With a Compass

RL shifts the optimization target to **maximizing expected reward** from a task-specific reward function. Rather than comparing against a single correct answer, the model generates its own responses, receives scalar reward signals (or penalties), and updates its policy accordingly.

The fundamental differences from SFT include:

| Aspect | Behavior |
|--------|----------|
| Learning mechanism | Discovers strategies through self-generated exploration |
| Generalization | Transfers knowledge to unseen scenarios beyond training demonstrations |
| Discovery capacity | Can invent behaviors that never appeared in any demonstration |
| Data requirements | Needs no standard answers—only task specification + reward function |

As documented in [`book-en/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter7.md) lines 70-76, RL shines when a cheap verification signal exists (e.g., test-case passing, numeric correctness) but writing exhaustive demonstrations would be prohibitively expensive.

## Core Technical Comparison

The *ai-agent-book* provides Table 7-2 (lines 78-88) comparing eight critical dimensions:

| Dimension | SFT | RL |
|-----------|-----|-----|
| **Optimization objective** | Maximum likelihood of labeled answer | Maximum expected reward |
| **Training signal** | Single standard answer per token | Multiple self-generated responses + reward |
| **Data form** | "Input → Output" pairs | Task + reward function |
| **What is learned** | Fixed mapping (memorization) | Transferable strategy (generalization) |
| **Distribution shift handling** | Performance degrades when applying old answers | Re-solves with stable strategy |
| **Sample efficiency** | High (thousands of examples) | Low (tens-to-hundreds× SFT data) |
| **Training stability** | High, quick convergence | Low, prone to oscillation |
| **Best suited for** | Formatting, protocols, stable environments | Unseen scenarios, outcome-driven tasks |

## Why SFT Must Precede RL

The "SFT first, then RL" ordering is not arbitrary—it is architecturally necessary. RL requires **parsable outputs** to compute rewards; without SFT, raw model generations are often malformed, making reward estimation impossible.

As stated in [`book-en/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter7.md) lines 64-67: *SFT establishes the **form**, RL refines the **spirit***. The industry-standard pipeline follows this dependency chain:

1. **SFT**: Produce valid, structured outputs
2. **RL**: Optimize decision quality within that structure

## Practical Implementation Examples

### SFT with LoRA (sesame_csm_sft_unsloth.py)

The repository includes a minimal SFT implementation using parameter-efficient fine-tuning:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import LoraConfig, get_peft_model

model_name = "Qwen2.5-32B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")

# LoRA: rank=64, learning rate ~10× full fine-tune

lora_cfg = LoraConfig(r=64, lora_alpha=32, target_modules=["q_proj","v_proj"])
model = get_peft_model(model, lora_cfg)

# Training on demonstration pairs with prompt masking

for sample in train_data:  # {"input": "...", "output": "..."}

    ids = tokenizer(sample["input"], return_tensors="pt").input_ids
    labels = tokenizer(sample["output"], return_tensors="pt").input_ids
    loss_mask = torch.ones_like(labels)  # Mask prompt, train only on response

    loss = model(input_ids=ids, labels=labels, loss_mask=loss_mask).loss
    loss.backward()
    optimizer.step()

```

Key implementation details from [`book-en/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter7.md) lines 58-59: LoRA configuration and loss masking for response-only training.

### RL with verl Framework (aworld_agent_loop.py)

The RL stage uses the `verl` library for PPO-based training:

```python
from verl import Trainer, PPOConfig, RewardModel

base_model = load_sft_checkpoint()  # Must be SFT-ed first

reward = RewardModel.load("path/to/reward_model")

cfg = PPOConfig(epochs=3, lr=1e-5, clip_range=0.2, target_kl=0.01)
trainer = Trainer(model=base_model, reward_model=reward, config=cfg)

# Environment: generate response, compute reward

def env_step(prompt):
    response = base_model.generate(prompt, max_new_tokens=256)
    reward_score = reward.compute(prompt, response)
    return response, reward_score

# RL training loop

for iteration in range(1000):
    prompt = "Solve 24 = ? using given numbers."
    resp, r = env_step(prompt)
    trainer.step(prompt, resp, r)  # Policy gradient update

```

This pattern appears in [`chapter8/gaia-experience/AWorld/train/adapter/verl/aworld_agent_loop.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/gaia-experience/AWorld/train/adapter/verl/aworld_agent_loop.py), demonstrating post-SFT reinforcement learning.

### End-to-End Shell Workflow

```bash

# Stage 1: SFT (establish form)

python -m scripts.sft \
    --model Qwen2.5-32B-Instruct \
    --lora-rank 64 \
    --data ./sft_data.jsonl \
    --output ./sft_adapter

# Stage 2: RL (refine strategy)

python -m verl.train_ppo \
    --base-model ./sft_adapter \
    --reward-model ./reward_checkpoint \
    --env ./my_env.py \
    --output ./rl_adapter

```

## Summary

- **SFT** = Learning a map drawn by a teacher; optimal for format adherence and stable instruction-following with limited data
- **RL** = Exploring with a compass (reward signal) to discover superior routes beyond any training map; essential for generalization and outcome-optimization
- **Pipeline order**: SFT must precede RL to ensure parsable outputs for reward computation
- **Key files**: [`book-en/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter7.md) (concepts), [`chapter7/sesame/sesame_csm_sft_unsloth.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/sesame/sesame_csm_sft_unsloth.py) (SFT code), [`chapter8/gaia-experience/AWorld/train/adapter/verl/aworld_agent_loop.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/gaia-experience/AWorld/train/adapter/verl/aworld_agent_loop.py) (RL code)

## Frequently Asked Questions

### Can you skip SFT and train an agent with RL only?

No—this is architecturally infeasible in practice. RL requires structured, parsable outputs to compute meaningful rewards. Without SFT, the base model generates malformed responses (invalid JSON, incorrect tool syntax) that break the reward pipeline. The *ai-agent-book* explicitly warns that RL without prior SFT produces "generations [that] are often malformed, making reward estimation impossible."

### Why is RL so much more data-hungry than SFT?

SFT learns from explicit teacher demonstrations—every training example directly conveys what to do. RL learns through exploration: the model must generate many candidate responses, evaluate their rewards, and gradually discover high-reward strategies. This trial-and-error process requires tens to hundreds of times more interactions to achieve comparable task mastery, as shown in Table 7-2's sample efficiency comparison.

### When should you use SFT alone without RL?

Use SFT-only training when: (1) demonstrations are cheap to obtain, (2) the task requires strict format adherence rather than creative problem-solving, (3) the environment is stable with no distribution shift, or (4) training time and compute are constrained. The *ai-agent-book* identifies "formatting, protocol, stable environments" as SFT's optimal domain—precisely because memorization outperforms exploration when the mapping never changes.