RLHF vs DPO: Which to Choose for Preference-Based Model Training

Choose RLHF when you need explicit reward modeling and policy regularization via PPO; choose DPO for a streamlined, memory-efficient approach that optimizes preferences directly through log-probability ratios without a separate reward model.

Both Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) align language models with human preferences, but they differ fundamentally in architectural complexity and computational requirements. The rohitg00/ai-engineering-from-scratch repository implements both approaches from first principles using pure NumPy, providing a transparent view of their internal mechanics. This guide examines their implementation details, memory footprints, and decision criteria based on the actual source code.

Core Architectural Differences

The fundamental distinction lies in how each method incorporates preference data into the training loop. RLHF introduces an intermediate reward model and a multi-stage pipeline, while DPO collapses the entire process into a single loss function.

The RLHF Multi-Stage Pipeline

RLHF requires three distinct training stages to align a model:

  1. Supervised Fine-Tuning (SFT) – Train the base model on high-quality demonstrations
  2. Reward Model Training – Train a separate RewardModel to predict human preferences using the Bradley-Terry loss
  3. PPO Fine-Tuning – Optimize the policy using Proximal Policy Optimization while maintaining a KL-divergence penalty against a reference model

This pipeline maintains 2–4 models in memory simultaneously (policy, reference, reward model, and sometimes an auxiliary KL model), creating significant memory overhead. Each PPO epoch samples new completions, evaluates them with the reward model, computes advantage estimates, and updates the policy with clipped gradients.

The DPO Single-Loop Approach

DPO eliminates the separate reward model entirely. Instead, it directly optimizes a loss function comparing the log-probabilities of the policy and reference model on preferred versus rejected responses. The implementation uses only two models (policy and reference) and requires a single training loop.

The core dpo_loss function computes a binary-cross-entropy-like term derived from the preference ratio, back-propagating directly through the policy network without on-policy sampling. This makes DPO more sample-efficient and stable when the reference model already produces reasonable generations.

Implementation Deep Dive in ai-engineering-from-scratch

The repository’s pure NumPy implementations expose the exact architectural differences between these approaches.

RLHF Pipeline Structure

The complete RLHF implementation resides in phases/10-llms-from-scratch/07-rlhf/code/main.py. Key components include:

  • RewardModel class (lines 48-77): Constructs a scalar reward from the final hidden state of the language model
  • bradley_terry_loss: Implements pairwise preference learning to train the reward model
  • compute_kl_divergence: Enforces the KL penalty during PPO to prevent the policy from drifting too far from the reference
  • PPO update loop: Handles clipped gradients and advantage estimation across multiple epochs

DPO Pipeline Structure

The DPO implementation lives in phases/10-llms-from-scratch/08-dpo/code/main.py. The critical functions are:

  • dpo_loss (lines 94-118): Computes the log-probability ratio between policy and reference models for preferred and rejected completions
  • dpo_train: Iterates over preference pairs with simple gradient steps, eliminating the need for reward model inference or PPO clipping

Both implementations utilize the same MiniGPT architecture defined in phases/04-pre-training-mini-gpt/code/main.py, ensuring a fair comparison of alignment strategies on identical base models.

Memory Footprint and Computational Cost

Memory requirements differ substantially between the two approaches. RLHF demands the highest memory footprint because the PPO step requires loading the policy, reference model, and reward model simultaneously, often with large minibatches for on-policy sampling.

DPO maintains a lower memory profile by keeping only the policy and reference models in memory. Since DPO uses the exact log-probabilities of existing preference data rather than sampling new completions, it achieves greater sample efficiency and requires fewer hyperparameters. RLHF requires tuning the KL-coefficient β, PPO clipping ε, learning-rate schedules, and reward-model training settings, whereas DPO primarily needs the β scaling factor (controlling preference margin strength) and a standard learning rate.

When to Choose RLHF vs DPO

Select RLHF when:

  • You possess a well-crafted reward model built from extensive human annotations
  • You need on-policy exploration to discover novel high-reward generations
  • Strict policy regularization is required (the KL term explicitly enforces proximity to the base model)

Select DPO when:

  • You have high-quality preference pairs and want fast, low-overhead fine-tuning
  • Computational resources are limited (GPUs with restricted VRAM)
  • The reference model already yields good generations and you only need modest alignment tweaks
  • You want to avoid the complexity of tuning PPO hyperparameters

Practical Code Examples

Running the RLHF Pipeline


# RLHF example from the repository

from phases/10-llms-from-scratch/07-rlhf/code/main import (
    RewardModel, train_reward_model, compute_kl_divergence
)

# Initialize policy and reference models (MiniGPT architecture)

policy = MiniGPT()
reference = MiniGPT()
copy_model_weights(policy, reference)

# Stage 2: Train the reward model on preference data

reward_model, loss_history, acc_history = train_reward_model(
    RewardModel(),
    PREFERENCE_DATA,
    num_epochs=10,
    lr=1e-4,
)

# Stage 3: PPO fine-tuning (simplified)

for epoch in range(3):
    # Sample completions, compute rewards, KL divergence, and update policy

    # Detailed PPO logic handles clipped gradients and advantage estimation

    ...

Running the DPO Pipeline

from phases/10-llms-from-scratch/08-dpo/code.main import (
    dpo_train, copy_model_weights, PREFERENCE_DATA
)

# Initialize policy and reference models

policy = MiniGPT()
reference = MiniGPT()
copy_model_weights(policy, reference)

# Single-loop Direct Preference Optimization

policy, loss_log, margin_log = dpo_train(
    policy_model=policy,
    reference_model=reference,
    preference_data=PREFERENCE_DATA,
    num_epochs=5,
    lr=5e-6,
    beta=0.1,  # Controls preference margin strength

)

print("Final DPO loss:", loss_log[-1])

Summary

  • RLHF requires a three-stage pipeline (SFT → Reward Model → PPO) with 2–4 models in memory, making it computationally expensive but offering fine-grained control through explicit reward modeling and KL regularization.
  • DPO executes in a single loop with only two models, optimizing directly via log-probability ratios for faster, more memory-efficient training.
  • Choose RLHF when you need on-policy exploration or have an existing high-quality reward model; choose DPO for rapid iteration with limited compute resources.
  • The rohitg00/ai-engineering-from-scratch repository provides pure NumPy implementations in phases/10-llms-from-scratch/07-rlhf/code/main.py and phases/10-llms-from-scratch/08-dpo/code/main.py for direct comparison.

Frequently Asked Questions

Is DPO always faster to train than RLHF?

Generally yes, because DPO eliminates the reward model training stage and the PPO sampling loop. As implemented in the repository, DPO completes in a single epoch over preference pairs, while RLHF requires sequential training of the reward model followed by multiple PPO epochs with new completion sampling. However, convergence speed depends on data quality—DPO requires clean preference pairs to perform well, whereas RLHF can sometimes extract more signal from noisy data through the reward model's generalization.

Can I switch from DPO to RLHF after initial training?

Yes, you can transition between methods. A model fine-tuned with DPO can serve as the base for subsequent RLHF training, using the DPO-tuned model as the starting point for the PPO phase. Conversely, you can extract preference pairs from an RLHF-trained policy's ranked outputs to create training data for DPO. The repository's shared MiniGPT architecture in phases/04-pre-training-mini-gpt/code/main.py ensures compatibility between both approaches.

What hardware requirements differ between RLHF and DPO?

RLHF typically requires GPUs with larger VRAM (24GB+ for moderate-sized models) because it must hold the policy model, reference model, and reward model in memory simultaneously during PPO updates. DPO runs comfortably on smaller GPUs (12-16GB) since it only requires the policy and reference models. According to the source code, DPO also supports smaller batch sizes because it doesn't require on-policy sampling, reducing activation memory.

Which method produces better alignment quality?

Alignment quality depends on your data and objectives. RLHF often achieves better performance on complex, multi-objective tasks when you have a well-trained reward model that captures subtle preferences. DPO excels when you have high-quality human preference annotations and want to stay close to the reference model's distribution. The repository demonstrates that DPO is more stable when the reference model is already strong, while RLHF's exploration mechanism can discover higher-reward behaviors that don't exist in the initial preference data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →