Stages of AI Model Capability Development in Post-Training: The Three-Stage Pipeline
AI model capability development in post-training follows a structured three-stage pipeline—Supervised Fine-Tuning (SFT), Reinforcement Learning (RL), and Continuous Evolution—that transforms foundation models from passive knowledge repositories into strategic, tool-aware agents capable of ongoing adaptation.
The bojieli/ai-agent-book repository documents how modern AI systems progress through distinct stages of AI model capability development in post-training after the initial large-scale pre-training phase. These stages function as a capability development ladder, with each phase writing increasingly complex behaviors into the model’s weights. The repository explicitly maps this progression in docs/en/LEARNING.md as the transition from "mid-training" through SFT to RL, while slides/lesson-01.md frames it as the cycle of "Evaluation → post-training → continual evolution."
The Three Stages of Post-Training Capability Development
Post-training capability development is not a monolithic process but a deliberately sequenced pipeline where each stage addresses specific limitations of the previous one.
Stage 1: Supervised Fine-Tuning (SFT)
Supervised Fine-Tuning (SFT) represents the first post-training step, aligning the model to concrete tasks or domains through labeled input-output pairs. During this stage, the model learns to generate desired responses directly from high-quality demonstration data, effectively memorizing specific behaviors and instruction-following patterns. The chapter8/retool/README.md file identifies SFT as the initial phase in their two-stage experimental pipeline, where the model is trained on task-specific datasets containing explicit tool-call examples.
Use SFT when you possess high-quality demonstrations and need the model to memorize explicit behaviors or follow precise instructions. This stage establishes the behavioral foundation upon which subsequent optimization builds.
Stage 2: Reinforcement Learning (RL)
Reinforcement Learning (RL)—including RL from Human Feedback (RLHF)—constitutes the second post-training step, refining SFT-learned behaviors by rewarding desirable outcomes and penalizing undesirable ones. Unlike SFT’s memorization approach, RL teaches the model when to invoke tools, how to plan multi-step actions, and how to generalize beyond training examples through strategic decision-making. The repository references several policy optimization algorithms for this stage, including PPO, DPO, GRPO, and the specialized DAPO algorithm used in the ReTool experiments described in chapter8/retool/README.md.
Apply RL when tasks require strategic decision-making, sophisticated tool-calling capabilities, or when you need the model to improve sample-efficiency and generalization beyond memorized patterns. slides/lesson-02.md explicitly queries "What post-training can write into weights," highlighting that this phase is responsible for encoding complex, conditional behaviors into model parameters.
Stage 3: Continuous Evolution
Continuous Evolution forms the long-term, cyclical stage where models undergo periodic re-evaluation and incremental updates to maintain performance over time. As outlined in slides/lesson-01.md, this stage follows the loop of "Evaluation → Post-Training → Continual Evolution," incorporating new data or environments through domain-adaptive fine-tuning or additional RL rounds. This stage may involve "continual pre-training" (still considered post-training for the model’s existing weights) or online RL driven by evaluation pipelines such as SWE-Bench, τ²-bench, or AndroidWorld.
Engage Continuous Evolution when operating environments change, new tools become available, or you need to prevent capability degradation over extended deployment periods.
Implementing the Post-Training Pipeline
The repository provides concrete implementation patterns using the Hugging Face ecosystem. Below are minimal, self-contained examples for each stage.
Supervised Fine-Tuning Implementation
This snippet demonstrates fine-tuning a pretrained Llama-2 model using the SFTTrainer from the TRL library, corresponding to the first stage described in docs/en/LEARNING.md:
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import SFTTrainer, DataCollatorForCompletionOnlyLM
import datasets
model_name = "meta-llama/Llama-2-7b-chat-hf"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
# Dataset format: {"prompt":"<instruction>", "completion":"<answer>"}
train_dataset = datasets.load_dataset("json", data_files="data/instructions.json")["train"]
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
train_dataset=train_dataset,
max_seq_length=1024,
dataset_text_field="prompt",
packing=False,
data_collator=DataCollatorForCompletionOnlyLM(tokenizer=tokenizer, response_template=""),
)
trainer.train()
Reinforcement Learning Implementation
The following example uses PPOTrainer to improve the SFT checkpoint using a reward function that prefers tool usage, matching the RL stage architecture found in chapter8/retool/README.md:
from trl import PPOTrainer, PPOConfig
import torch
# Assume `model` is the SFT checkpoint from previous stage
def reward_fn(samples, **kwargs):
# Reward: +1 if generated text contains tool invocation
return torch.tensor([1.0 if "tool" in s else 0.0 for s in samples])
ppo_cfg = PPOConfig(
batch_size=8,
ppo_epochs=4,
learning_rate=5e-6,
)
ppo_trainer = PPOTrainer(
config=ppo_cfg,
model=model,
ref_model=model, # Reference model for KL penalty
tokenizer=tokenizer,
reward_fn=reward_fn,
)
ppo_trainer.train()
Continuous Evolution Implementation
This evaluation loop illustrates the Continuous Evolution stage, periodically validating model performance and triggering additional training cycles when accuracy thresholds drop, as conceptualized in chapter9/README.id.md:
from evaluate import load as load_metric
metric = load_metric("accuracy")
def evaluate(model, tokenizer, dataset):
preds, refs = [], []
for example in dataset:
inputs = tokenizer(example["question"], return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=50)
preds.append(tokenizer.decode(out[0], skip_special_tokens=True))
refs.append(example["answer"])
return metric.compute(predictions=preds, references=refs)
# Continuous evolution loop
while True:
scores = evaluate(model, tokenizer, validation_set)
if scores["accuracy"] < 0.90:
# Trigger new SFT or RL round based on evaluation feedback
break
Key Repository References
The bojieli/ai-agent-book repository contains definitive documentation of these stages across several critical files:
docs/en/LEARNING.md: Outlines the three-stage panorama (mid-training / SFT / RL) providing the high-level conceptual map for capability development.slides/lesson-01.md: Introduces the explicit "Evaluation, post-training, and continual evolution" cycle that defines the long-term stage.slides/lesson-02.md: Examines what post-training stages write into model weights, emphasizing parameter modification during these phases.chapter8/retool/README.md: Documents the concrete two-stage pipeline (SFT → RL) used in the ReTool experiments, including the DAPO RL algorithm.chapter8/README.en.md: Catalogues RLVP post-training research as distinct stages 8-16 in the experimentation framework.chapter9/README.id.md: Discusses how post-training connects skill generation to synthetic data creation, illustrating downstream impacts of the pipeline.
Summary
- Supervised Fine-Tuning (SFT) serves as the behavioral foundation, using labeled demonstrations to teach task-specific responses and explicit instruction following.
- Reinforcement Learning (RL) refines these behaviors through reward optimization, enabling tool-use, strategic planning, and generalization beyond training data using algorithms like PPO, DPO, or DAPO.
- Continuous Evolution maintains model capabilities over time through cyclical evaluation and retraining, incorporating new tools and environmental changes via feedback loops from benchmarks like SWE-Bench or AndroidWorld.
- The bojieli/ai-agent-book repository explicitly structures these stages in
docs/en/LEARNING.mdandslides/lesson-01.mdas the canonical path from pre-trained base model to capable AI agent.
Frequently Asked Questions
What distinguishes SFT from RL in the post-training pipeline?
Supervised Fine-Tuning relies on static, labeled demonstrations to teach the model to reproduce specific input-output mappings, effectively memorizing behaviors. Reinforcement Learning, conversely, uses reward signals to optimize for outcomes, teaching the model when to apply behaviors rather than just how, enabling dynamic tool selection and multi-step planning that extends beyond the training distribution.
Why is Continuous Evolution necessary after completing RL training?
Continuous Evolution addresses distribution shift and environmental dynamism; as operating conditions change or new tools emerge (documented in chapter9/README.id.md), models require periodic re-evaluation and updating to prevent capability degradation. This stage implements the feedback loop of "Evaluation → Post-Training → Continual Evolution" described in slides/lesson-01.md.
Which RL algorithms does the AI-Agent-Book repository recommend for post-training?
According to chapter8/retool/README.md and related documentation, the repository references PPO (Proximal Policy Optimization), DPO (Direct Preference Optimization), GRPO, and the specialized DAPO algorithm used specifically in the ReTool experiments for optimizing tool-use behaviors.
How does post-training differ from the initial pre-training phase?
Pre-training instills broad linguistic knowledge and world understanding through unsupervised learning on massive text corpora. Post-training stages—SFT, RL, and Continuous Evolution—modify these pre-trained weights to instill specific capabilities, tool-use patterns, and alignment characteristics that transform a general language model into a specialized agent, as emphasized in slides/lesson-02.md regarding what post-training "writes into weights."
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →