Implementing Reinforcement Learning from Human Feedback (RLHF): The Complete Three-Stage Pipeline

Reinforcement Learning from Human Feedback (RLHF) aligns large language models through a three-stage pipeline—supervised fine-tuning, reward model training with Bradley-Terry pairwise loss, and PPO optimization with KL divergence regularization—that creates helpful, harmless assistants while preventing reward hacking.

The rohitg00/ai-engineering-from-scratch repository provides a production-ready curriculum walking through the complete RLHF implementation, from the mathematical foundations in code/main.py to high-level orchestration with the TRL library. This guide examines how the codebase implements each stage, why the KL penalty serves as the critical safety mechanism, and how modern variants like DPO and GRPO extend these concepts.

The Three-Stage RLHF Architecture

The RLHF pipeline implemented in phases/09-reinforcement-learning/09-reward-modeling-rlhf/code/main.py follows a strict three-stage progression that transforms a pretrained base model into an aligned policy.

Stage 1 – Supervised Fine-Tuning (SFT)

The pipeline begins with Supervised Fine-Tuning (SFT), where a pretrained base model is fine-tuned on human-written demonstrations of desired behavior. This stage produces the initial policy π_SFT, which is biased toward helpful responses but remains unrestricted and potentially unsafe. The repository implements this as the foundation for subsequent stages, noting that π_SFT serves dual purposes: it becomes both the initialization for the trainable policy and the frozen reference model in stage three.

Stage 2 – Reward Model (RM) Training

The second stage trains a scalar reward function R_φ(x, y) to score the quality of responses. The implementation uses the Bradley-Terry pairwise logistic loss:


L(φ) = -E[log σ(R_φ(x,y₊) - R_φ(x,y₋))]

In code/main.py, the rm_train_step function implements this using a linear bag-of-tokens scorer:

def rm_train_step(w, x, y_pos, y_neg, lr):
    r_pos = dot(w, bag(y_pos))
    r_neg = dot(w, bag(y_neg))
    p = sigmoid(r_pos - r_neg)                     # probability y_pos is preferred

    for tok, cnt in bag(y_pos).items():            # gradient ascent on good tokens

        w[tok] += lr * (1 - p) * cnt
    for tok, cnt in bag(y_neg).items():            # gradient descent on bad tokens

        w[tok] -= lr * (1 - p) * cnt

The training relies on human preference pairs (y₊, y₋) for prompts x, where the model learns to assign higher scores to preferred responses. The synthetic data generation in the repository uses GOOD_WORDS and BAD_WORDS sets to create these pairs programmatically:

PROMPTS = ["help me", "answer me", "explain this"]
GOOD_WORDS = {"clear", "specific", "kind", "thorough"}
BAD_WORDS  = {"vague", "rude", "wrong", "short"}

def make_pair(rng):
    x = rng.choice(PROMPTS)
    y_good = rng.choice(list(GOOD_WORDS)) + " " + rng.choice(list(GOOD_WORDS))
    y_bad  = rng.choice(list(BAD_WORDS)) + " " + rng.choice(list(BAD_WORDS))
    return (x, y_good, y_bad)

Stage 3 – PPO with KL Penalty

The final stage optimizes a trainable policy π_θ (initialized from π_SFT) against the reward model while constraining divergence from a frozen reference π_ref = π_SFT. The rlhf_step function in code/main.py implements the core logic, calculating the total reward as:


r_total = R_φ(x,y) - β·KL(π_θ‖π_ref)

The implementation computes the KL divergence between the current policy and reference policy logits, then applies PPO clipped surrogate updates:

def rlhf_step(theta, ref, w, prompt, rng, eps=0.2, beta=0.1, lr=0.05):
    # Compute logits for current policy and reference

    logits_theta = policy_logits(theta, prompt)
    probs_theta  = softmax(logits_theta)
    token = sample(probs_theta, rng)

    logits_ref = policy_logits(ref, prompt)
    probs_ref  = softmax(logits_ref)

    # Reward = RM score – β·KL

    reward = dot(w, bag([token])) - beta * kl(probs_theta, probs_ref)

    # Apply clipped PPO surrogate (omitted for brevity)

    ...

The beta (β) coefficient controls the strength of the regularization, making it the single most important hyperparameter in the RLHF training loop.

Why the KL Penalty Prevents Reward Hacking

Without the KL penalty, the policy can exploit the reward model by generating out-of-distribution outputs that receive high scores but produce undesirable or incoherent text—a phenomenon known as reward hacking. The KL term acts as a safety leash that keeps π_θ near the manifold where the reward model was trained. As documented in the repository's lesson notes, this constraint ensures the optimized policy remains within the distribution of high-quality human demonstrations seen during SFT training, preventing the model from drifting toward adversarial examples that game the scoring function.

From RLHF to DPO and GRPO: Modern Alignment Variants

The repository extends the three-stage framework to cover modern alternatives that simplify or specialize the pipeline:

  • Direct Preference Optimization (DPO) collapses stages two and three into a single supervised loss over preference data, eliminating the explicit reward model and PPO optimization entirely. This approach, introduced by Rafailov et al. (2023), treats the policy as its own implicit reward model.

  • Group Relative Policy Optimization (GRPO) replaces the learned reward model with a verifier that scores code execution or mathematical correctness, making it ideal for reasoning models where outcomes are objectively verifiable.

  • Process Reward Models (PRMs) extend the reward model concept to score intermediate reasoning steps rather than final outcomes, enabling fine-grained alignment for multi-step tasks like chain-of-thought reasoning.

All variants share the same conceptual foundation established in the three-stage RLHF architecture, making the original pipeline the essential mental model for modern alignment engineering.

Production Implementation with the TRL Library

While the main.py implementation demonstrates the core algorithms, production RLHF at scale uses the Transformers Reinforcement Learning (TRL) library. The repository maps the three stages onto high-level TRL calls:

Stage 2 – Reward Model Training:

from trl import RewardTrainer, RewardConfig
trainer = RewardTrainer(
    model=rm, tokenizer=tok,
    train_dataset=preference_data,
    args=RewardConfig(output_dir="./rm", num_train_epochs=1, learning_rate=1e-5),
)
trainer.train()

Stage 3 – PPO with Adaptive KL Control:

from trl import PPOTrainer, PPOConfig
ppo = PPOTrainer(
    model=policy, ref_model=ref, tokenizer=tok,
    config=PPOConfig(learning_rate=1.41e-5, batch_size=64,
                     init_kl_coef=0.05, target_kl=6.0, adap_kl_ctrl=True),
)

The adap_kl_ctrl=True flag enables dynamic adjustment of the KL coefficient during training, automatically increasing regularization if the policy drifts too far from the reference.

Summary

  • RLHF consists of three sequential stages: Supervised Fine-Tuning creates the initial policy, Reward Model training learns human preferences via Bradley-Terry loss, and PPO optimization maximizes rewards while constrained by KL divergence from the reference.

  • The KL penalty is essential for safety: It prevents reward hacking by keeping the optimized policy near the distribution where the reward model remains accurate, acting as the primary safety mechanism in the training loop.

  • Modern variants simplify the pipeline: DPO removes the explicit reward model and PPO steps, while GRPO substitutes verifiers for learned reward models in code and math domains.

  • Production code differs from educational implementations: The TRL library provides RewardTrainer and PPOTrainer classes that implement the same mathematical foundations with adaptive KL control and batched GPU optimization.

Frequently Asked Questions

What is the Bradley-Terry loss function used for in RLHF?

The Bradley-Terry loss function trains the reward model to output higher scalar scores for preferred responses (y₊) than for rejected responses (y₋). The formula L(φ) = -E[log σ(R_φ(x,y₊) - R_φ(x,y₋))] maximizes the log-probability that the preferred response receives a higher score, learning from pairwise human comparisons rather than absolute ratings.

Why is the KL divergence penalty necessary in the PPO stage?

The KL divergence penalty prevents the policy from exploiting the reward model by generating high-scoring but out-of-distribution text. By penalizing deviation from the frozen reference policy π_ref (the SFT model), the term ensures the optimized model remains within the distribution where the reward model is accurate, preventing reward hacking and maintaining output quality.

How does DPO differ from standard RLHF?

Direct Preference Optimization (DPO) eliminates the separate reward model and PPO optimization stages, instead training the language model directly on preference data using a supervised loss. DPO derives the optimal policy under the RLHF objective analytically, allowing the model to implicitly represent its own reward function while remaining computationally simpler than three-stage RLHF.

What file contains the core RLHF implementation in the repository?

The file phases/09-reinforcement-learning/09-reward-modeling-rlhf/code/main.py contains the minimal Python implementation covering synthetic preference generation, Bradley-Terry reward model updates, and PPO-style optimization with KL penalties. The accompanying documentation in phases/09-reinforcement-learning/09-reward-modeling-rlhf/docs/en.md provides the theoretical narrative and production recipes using the TRL library.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →