# Stages of AI Model Capability Development in Post-Training: The Three-Stage Pipeline

> Discover the three stages of AI model capability development in post-training SFT, RL, and Continuous Evolution. Transform models into adaptive agents.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-24

---

**AI model capability development in post-training follows a structured three-stage pipeline—Supervised Fine-Tuning (SFT), Reinforcement Learning (RL), and Continuous Evolution—that transforms foundation models from passive knowledge repositories into strategic, tool-aware agents capable of ongoing adaptation.**

The **bojieli/ai-agent-book** repository documents how modern AI systems progress through distinct **stages of AI model capability development in post-training** after the initial large-scale pre-training phase. These stages function as a capability development ladder, with each phase writing increasingly complex behaviors into the model’s weights. The repository explicitly maps this progression in [`docs/en/LEARNING.md`](https://github.com/bojieli/ai-agent-book/blob/main/docs/en/LEARNING.md) as the transition from "mid-training" through SFT to RL, while [`slides/lesson-01.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-01.md) frames it as the cycle of "Evaluation → post-training → continual evolution."

## The Three Stages of Post-Training Capability Development

Post-training capability development is not a monolithic process but a deliberately sequenced pipeline where each stage addresses specific limitations of the previous one.

### Stage 1: Supervised Fine-Tuning (SFT)

**Supervised Fine-Tuning (SFT)** represents the first post-training step, aligning the model to concrete tasks or domains through labeled input-output pairs. During this stage, the model learns to generate desired responses directly from high-quality demonstration data, effectively memorizing specific behaviors and instruction-following patterns. The [`chapter8/retool/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/retool/README.md) file identifies SFT as the initial phase in their two-stage experimental pipeline, where the model is trained on task-specific datasets containing explicit tool-call examples.

Use SFT when you possess high-quality demonstrations and need the model to memorize explicit behaviors or follow precise instructions. This stage establishes the behavioral foundation upon which subsequent optimization builds.

### Stage 2: Reinforcement Learning (RL)

**Reinforcement Learning (RL)**—including RL from Human Feedback (RLHF)—constitutes the second post-training step, refining SFT-learned behaviors by rewarding desirable outcomes and penalizing undesirable ones. Unlike SFT’s memorization approach, RL teaches the model *when* to invoke tools, how to plan multi-step actions, and how to generalize beyond training examples through strategic decision-making. The repository references several policy optimization algorithms for this stage, including PPO, DPO, GRPO, and the specialized **DAPO** algorithm used in the ReTool experiments described in [`chapter8/retool/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/retool/README.md).

Apply RL when tasks require strategic decision-making, sophisticated tool-calling capabilities, or when you need the model to improve sample-efficiency and generalization beyond memorized patterns. [`slides/lesson-02.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-02.md) explicitly queries "What post-training can write into weights," highlighting that this phase is responsible for encoding complex, conditional behaviors into model parameters.

### Stage 3: Continuous Evolution

**Continuous Evolution** forms the long-term, cyclical stage where models undergo periodic re-evaluation and incremental updates to maintain performance over time. As outlined in [`slides/lesson-01.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-01.md), this stage follows the loop of "Evaluation → Post-Training → Continual Evolution," incorporating new data or environments through domain-adaptive fine-tuning or additional RL rounds. This stage may involve "continual pre-training" (still considered post-training for the model’s existing weights) or online RL driven by evaluation pipelines such as SWE-Bench, τ²-bench, or AndroidWorld.

Engage Continuous Evolution when operating environments change, new tools become available, or you need to prevent capability degradation over extended deployment periods.

## Implementing the Post-Training Pipeline

The repository provides concrete implementation patterns using the Hugging Face ecosystem. Below are minimal, self-contained examples for each stage.

### Supervised Fine-Tuning Implementation

This snippet demonstrates fine-tuning a pretrained Llama-2 model using the `SFTTrainer` from the TRL library, corresponding to the first stage described in [`docs/en/LEARNING.md`](https://github.com/bojieli/ai-agent-book/blob/main/docs/en/LEARNING.md):

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import SFTTrainer, DataCollatorForCompletionOnlyLM
import datasets

model_name = "meta-llama/Llama-2-7b-chat-hf"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")

# Dataset format: {"prompt":"<instruction>", "completion":"<answer>"}

train_dataset = datasets.load_dataset("json", data_files="data/instructions.json")["train"]

trainer = SFTTrainer(
    model=model,
    tokenizer=tokenizer,
    train_dataset=train_dataset,
    max_seq_length=1024,
    dataset_text_field="prompt",
    packing=False,
    data_collator=DataCollatorForCompletionOnlyLM(tokenizer=tokenizer, response_template=""),
)
trainer.train()

```

### Reinforcement Learning Implementation

The following example uses `PPOTrainer` to improve the SFT checkpoint using a reward function that prefers tool usage, matching the RL stage architecture found in [`chapter8/retool/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/retool/README.md):

```python
from trl import PPOTrainer, PPOConfig
import torch

# Assume `model` is the SFT checkpoint from previous stage

def reward_fn(samples, **kwargs):
    # Reward: +1 if generated text contains tool invocation

    return torch.tensor([1.0 if "tool" in s else 0.0 for s in samples])

ppo_cfg = PPOConfig(
    batch_size=8,
    ppo_epochs=4,
    learning_rate=5e-6,
)

ppo_trainer = PPOTrainer(
    config=ppo_cfg,
    model=model,
    ref_model=model,  # Reference model for KL penalty

    tokenizer=tokenizer,
    reward_fn=reward_fn,
)
ppo_trainer.train()

```

### Continuous Evolution Implementation

This evaluation loop illustrates the Continuous Evolution stage, periodically validating model performance and triggering additional training cycles when accuracy thresholds drop, as conceptualized in [`chapter9/README.id.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/README.id.md):

```python
from evaluate import load as load_metric

metric = load_metric("accuracy")

def evaluate(model, tokenizer, dataset):
    preds, refs = [], []
    for example in dataset:
        inputs = tokenizer(example["question"], return_tensors="pt").to("cuda")
        out = model.generate(**inputs, max_new_tokens=50)
        preds.append(tokenizer.decode(out[0], skip_special_tokens=True))
        refs.append(example["answer"])
    return metric.compute(predictions=preds, references=refs)

# Continuous evolution loop

while True:
    scores = evaluate(model, tokenizer, validation_set)
    if scores["accuracy"] < 0.90:
        # Trigger new SFT or RL round based on evaluation feedback

        break

```

## Key Repository References

The **bojieli/ai-agent-book** repository contains definitive documentation of these stages across several critical files:

- **[`docs/en/LEARNING.md`](https://github.com/bojieli/ai-agent-book/blob/main/docs/en/LEARNING.md)**: Outlines the three-stage panorama (mid-training / SFT / RL) providing the high-level conceptual map for capability development.
- **[`slides/lesson-01.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-01.md)**: Introduces the explicit "Evaluation, post-training, and continual evolution" cycle that defines the long-term stage.
- **[`slides/lesson-02.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-02.md)**: Examines what post-training stages write into model weights, emphasizing parameter modification during these phases.
- **[`chapter8/retool/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/retool/README.md)**: Documents the concrete two-stage pipeline (SFT → RL) used in the ReTool experiments, including the DAPO RL algorithm.
- **[`chapter8/README.en.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/README.en.md)**: Catalogues RLVP post-training research as distinct stages 8-16 in the experimentation framework.
- **[`chapter9/README.id.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/README.id.md)**: Discusses how post-training connects skill generation to synthetic data creation, illustrating downstream impacts of the pipeline.

## Summary

- **Supervised Fine-Tuning (SFT)** serves as the behavioral foundation, using labeled demonstrations to teach task-specific responses and explicit instruction following.
- **Reinforcement Learning (RL)** refines these behaviors through reward optimization, enabling tool-use, strategic planning, and generalization beyond training data using algorithms like PPO, DPO, or DAPO.
- **Continuous Evolution** maintains model capabilities over time through cyclical evaluation and retraining, incorporating new tools and environmental changes via feedback loops from benchmarks like SWE-Bench or AndroidWorld.
- The **bojieli/ai-agent-book** repository explicitly structures these stages in [`docs/en/LEARNING.md`](https://github.com/bojieli/ai-agent-book/blob/main/docs/en/LEARNING.md) and [`slides/lesson-01.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-01.md) as the canonical path from pre-trained base model to capable AI agent.

## Frequently Asked Questions

### What distinguishes SFT from RL in the post-training pipeline?

**Supervised Fine-Tuning** relies on static, labeled demonstrations to teach the model to reproduce specific input-output mappings, effectively memorizing behaviors. **Reinforcement Learning**, conversely, uses reward signals to optimize for outcomes, teaching the model *when* to apply behaviors rather than just *how*, enabling dynamic tool selection and multi-step planning that extends beyond the training distribution.

### Why is Continuous Evolution necessary after completing RL training?

Continuous Evolution addresses **distribution shift** and **environmental dynamism**; as operating conditions change or new tools emerge (documented in [`chapter9/README.id.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/README.id.md)), models require periodic re-evaluation and updating to prevent capability degradation. This stage implements the feedback loop of "Evaluation → Post-Training → Continual Evolution" described in [`slides/lesson-01.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-01.md).

### Which RL algorithms does the AI-Agent-Book repository recommend for post-training?

According to [`chapter8/retool/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/retool/README.md) and related documentation, the repository references **PPO** (Proximal Policy Optimization), **DPO** (Direct Preference Optimization), **GRPO**, and the specialized **DAPO** algorithm used specifically in the ReTool experiments for optimizing tool-use behaviors.

### How does post-training differ from the initial pre-training phase?

Pre-training instills broad linguistic knowledge and world understanding through unsupervised learning on massive text corpora. **Post-training** stages—SFT, RL, and Continuous Evolution—modify these pre-trained weights to instill specific capabilities, tool-use patterns, and alignment characteristics that transform a general language model into a specialized agent, as emphasized in [`slides/lesson-02.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-02.md) regarding what post-training "writes into weights."