SFT vs RL for Agent Model Training: Key Differences and When to Use Each

SFT memorizes fixed input-output mappings from teacher demonstrations, while RL learns generalizable strategies through trial-and-error reward optimization.

Training large language model (LLM)-based agents typically follows a three-stage pipeline: pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL). According to the ai-agent-book source code, each stage serves a distinct purpose with fundamentally different optimization objectives.

What SFT Does: Learning the Teacher's Map

SFT optimizes for maximum likelihood estimation—the model learns to maximize the probability of a labeled answer given an input prompt. During training, the model receives input-output demonstration pairs and computes loss only on the response portion, effectively learning to "predict the next token" within the answer.

This approach yields three characteristic outcomes:

  • Memorization: The model internalizes exact mappings from prompts to demonstrated answers
  • Format consistency: Excellent instruction-following and output structure adherence
  • Training efficiency: Converges in hours to days with thousands of examples

In book-en/chapter7.md lines 47-53, the authors emphasize that SFT excels at solidifying the "form" of an agent—establishing output style, JSON schemas, and tool-call syntax before any exploration-based learning occurs.

What RL Does: Exploring With a Compass

RL shifts the optimization target to maximizing expected reward from a task-specific reward function. Rather than comparing against a single correct answer, the model generates its own responses, receives scalar reward signals (or penalties), and updates its policy accordingly.

The fundamental differences from SFT include:

Aspect Behavior
Learning mechanism Discovers strategies through self-generated exploration
Generalization Transfers knowledge to unseen scenarios beyond training demonstrations
Discovery capacity Can invent behaviors that never appeared in any demonstration
Data requirements Needs no standard answers—only task specification + reward function

As documented in book-en/chapter7.md lines 70-76, RL shines when a cheap verification signal exists (e.g., test-case passing, numeric correctness) but writing exhaustive demonstrations would be prohibitively expensive.

Core Technical Comparison

The ai-agent-book provides Table 7-2 (lines 78-88) comparing eight critical dimensions:

Dimension SFT RL
Optimization objective Maximum likelihood of labeled answer Maximum expected reward
Training signal Single standard answer per token Multiple self-generated responses + reward
Data form "Input → Output" pairs Task + reward function
What is learned Fixed mapping (memorization) Transferable strategy (generalization)
Distribution shift handling Performance degrades when applying old answers Re-solves with stable strategy
Sample efficiency High (thousands of examples) Low (tens-to-hundreds× SFT data)
Training stability High, quick convergence Low, prone to oscillation
Best suited for Formatting, protocols, stable environments Unseen scenarios, outcome-driven tasks

Why SFT Must Precede RL

The "SFT first, then RL" ordering is not arbitrary—it is architecturally necessary. RL requires parsable outputs to compute rewards; without SFT, raw model generations are often malformed, making reward estimation impossible.

As stated in book-en/chapter7.md lines 64-67: SFT establishes the form, RL refines the spirit. The industry-standard pipeline follows this dependency chain:

  1. SFT: Produce valid, structured outputs
  2. RL: Optimize decision quality within that structure

Practical Implementation Examples

SFT with LoRA (sesame_csm_sft_unsloth.py)

The repository includes a minimal SFT implementation using parameter-efficient fine-tuning:

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import LoraConfig, get_peft_model

model_name = "Qwen2.5-32B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")

# LoRA: rank=64, learning rate ~10× full fine-tune

lora_cfg = LoraConfig(r=64, lora_alpha=32, target_modules=["q_proj","v_proj"])
model = get_peft_model(model, lora_cfg)

# Training on demonstration pairs with prompt masking

for sample in train_data:  # {"input": "...", "output": "..."}

    ids = tokenizer(sample["input"], return_tensors="pt").input_ids
    labels = tokenizer(sample["output"], return_tensors="pt").input_ids
    loss_mask = torch.ones_like(labels)  # Mask prompt, train only on response

    loss = model(input_ids=ids, labels=labels, loss_mask=loss_mask).loss
    loss.backward()
    optimizer.step()

Key implementation details from book-en/chapter7.md lines 58-59: LoRA configuration and loss masking for response-only training.

RL with verl Framework (aworld_agent_loop.py)

The RL stage uses the verl library for PPO-based training:

from verl import Trainer, PPOConfig, RewardModel

base_model = load_sft_checkpoint()  # Must be SFT-ed first

reward = RewardModel.load("path/to/reward_model")

cfg = PPOConfig(epochs=3, lr=1e-5, clip_range=0.2, target_kl=0.01)
trainer = Trainer(model=base_model, reward_model=reward, config=cfg)

# Environment: generate response, compute reward

def env_step(prompt):
    response = base_model.generate(prompt, max_new_tokens=256)
    reward_score = reward.compute(prompt, response)
    return response, reward_score

# RL training loop

for iteration in range(1000):
    prompt = "Solve 24 = ? using given numbers."
    resp, r = env_step(prompt)
    trainer.step(prompt, resp, r)  # Policy gradient update

This pattern appears in chapter8/gaia-experience/AWorld/train/adapter/verl/aworld_agent_loop.py, demonstrating post-SFT reinforcement learning.

End-to-End Shell Workflow


# Stage 1: SFT (establish form)

python -m scripts.sft \
    --model Qwen2.5-32B-Instruct \
    --lora-rank 64 \
    --data ./sft_data.jsonl \
    --output ./sft_adapter

# Stage 2: RL (refine strategy)

python -m verl.train_ppo \
    --base-model ./sft_adapter \
    --reward-model ./reward_checkpoint \
    --env ./my_env.py \
    --output ./rl_adapter

Summary

Frequently Asked Questions

Can you skip SFT and train an agent with RL only?

No—this is architecturally infeasible in practice. RL requires structured, parsable outputs to compute meaningful rewards. Without SFT, the base model generates malformed responses (invalid JSON, incorrect tool syntax) that break the reward pipeline. The ai-agent-book explicitly warns that RL without prior SFT produces "generations [that] are often malformed, making reward estimation impossible."

Why is RL so much more data-hungry than SFT?

SFT learns from explicit teacher demonstrations—every training example directly conveys what to do. RL learns through exploration: the model must generate many candidate responses, evaluate their rewards, and gradually discover high-reward strategies. This trial-and-error process requires tens to hundreds of times more interactions to achieve comparable task mastery, as shown in Table 7-2's sample efficiency comparison.

When should you use SFT alone without RL?

Use SFT-only training when: (1) demonstrations are cheap to obtain, (2) the task requires strict format adherence rather than creative problem-solving, (3) the environment is stable with no distribution shift, or (4) training time and compute are constrained. The ai-agent-book identifies "formatting, protocol, stable environments" as SFT's optimal domain—precisely because memorization outperforms exploration when the mapping never changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →