Implementing Direct Preference Optimization (DPO) for LLM Alignment: A Complete Guide
Direct Preference Optimization (DPO) eliminates the need for complex reinforcement learning pipelines by directly fine-tuning large language models on pairwise human preferences using a simple binary classification loss.
Direct Preference Optimization (DPO) offers a streamlined approach to LLM alignment that bypasses the computational overhead of traditional RLHF methods. The rasbt/LLMs-from-scratch repository provides a minimal, from-scratch implementation in chapter 7 that requires only PyTorch and the book’s existing GPT architecture. This guide walks through the exact source code, file paths, and hyperparameters needed to implement DPO for preference tuning without external RL libraries.
What Is Direct Preference Optimization?
Direct Preference Optimization reformulates alignment as a binary classification problem rather than a reinforcement learning task. Instead of training a separate reward model and using PPO, DPO directly optimizes the policy model to increase the likelihood of preferred responses while decreasing the likelihood of rejected ones, relative to a frozen reference model.
The method requires only a pairwise dataset containing chosen (preferred) and rejected (dispreferred) completions for each prompt. According to the implementation in ch07/04_preference-tuning-with-dpo/dpo-from-scratch.ipynb, the core algorithm compactly fits in approximately 30 lines of code and runs entirely within the repository’s existing GPT framework.
Core Architecture and Components
The DPO implementation in the LLMs-from-scratch repository relies on five essential components that work together to compute the preference-based loss.
Policy and Reference Models
The architecture requires two independent copies of the same GPT model:
- Policy model (
policy_model): The trainable network that will learn from preferences. - Reference model (
reference_model): A frozen copy of the original supervised fine-tuned (SFT) checkpoint that provides a stable baseline for comparison.
Both models use the generic GPTModel class defined in the repository’s package structure. As implemented in the notebook, the reference model is set to eval() mode immediately after loading weights to prevent gradient updates:
policy = GPTModel(BASE_CONFIG).to(device)
policy.load_state_dict(torch.load("gpt2-medium355M-sft.pth"))
policy.train()
reference = GPTModel(BASE_CONFIG).to(device)
reference.load_state_dict(torch.load("gpt2-medium355M-sft.pth"))
reference.eval()
Pairwise Preference Dataset
The InstructionDataset class, originally defined in ch07/02_dataset-utilities/previous_chapters.py, is adapted to return tuples of (prompt, chosen, rejected) rather than single instruction-response pairs. This allows the training loop to process both preference options simultaneously for each input prompt.
The compute_dpo_loss Function
The heart of the implementation lives in lines 71-78 of dpo-from-scratch.ipynb. This function implements Equation 7 from the original DPO paper, calculating the log-ratio between policy and reference models for both chosen and rejected responses:
def compute_dpo_loss(chosen_logp, rejected_logp,
ref_chosen_logp, ref_rej_logp,
beta=0.1):
model_logratios = chosen_logp - rejected_logp
ref_logratios = ref_chosen_logp - ref_rej_logp
logits = model_logratios - ref_logratios
loss = -F.logsigmoid(beta * logits).mean()
chosen_reward = (chosen_logp - ref_chosen_logp).mean()
rejected_reward = (rejected_logp - ref_rej_logp).mean()
return loss, chosen_reward, rejected_reward
The beta (β) parameter controls the strength of the preference signal, typically set between 0.1 and 0.5. Smaller values create more aggressive optimization, while larger values constrain the policy to stay closer to the reference model.
Log-Probability Helpers
The compute_logprobs function (lines 74-99 of the notebook) converts model logits into per-token log-probabilities and extracts values only for the target tokens, ignoring padding. This utility handles the critical alignment between logits and labels by shifting the target sequence one position to the left:
def compute_logprobs(logits, labels, mask=None):
labels = labels[:, 1:].clone()
logits = logits[:, :-1, :]
logp = F.log_softmax(logits, dim=-1)
selected = torch.gather(logp, -1, labels.unsqueeze(-1)).squeeze(-1)
if mask is not None:
selected = selected * mask
return selected.sum() / mask.sum()
return selected.mean()
Implementing the Training Pipeline
Dataset Preparation
The training pipeline begins with loading a pairwise preference dataset where each sample contains a prompt and two possible completions. The data loader batches these triplets for efficient GPU processing.
Forward Pass and Loss Calculation
For each training step, the implementation performs forward passes through both the policy and reference models, then calculates log-probabilities for the chosen and rejected sequences separately:
optimizer = torch.optim.AdamW(policy.parameters(), lr=1e-5)
for batch in dpo_loader:
optimizer.zero_grad()
# Forward passes through both models
policy_out = policy(batch["prompt"])
ref_out = reference(batch["prompt"])
# Calculate log-probabilities for chosen and rejected responses
pol_chosen_logp = compute_logprobs(policy_out, batch["chosen"])
pol_rej_logp = compute_logprobs(policy_out, batch["rejected"])
ref_chosen_logp = compute_logprobs(ref_out, batch["chosen"])
ref_rej_logp = compute_logprobs(ref_out, batch["rejected"])
# Compute DPO loss
loss, chosen_r, rejected_r = compute_dpo_loss(
pol_chosen_logp, pol_rej_logp,
ref_chosen_logp, ref_rej_logp,
beta=0.1
)
loss.backward()
optimizer.step()
Monitoring Reward Margins
The compute_dpo_loss function returns chosen_reward and rejected_reward values representing the difference between policy and reference model log-probabilities. Tracking the reward margin (chosen_r - rejected_r) serves as a crucial sanity check—the margin should increase steadily during training, indicating that the policy is learning to distinguish preferred from dispreferred outputs.
Critical Implementation Details
DPO training requires careful attention to hyperparameters to prevent model collapse:
- Limited Epochs: Unlike standard fine-tuning, DPO can quickly overfit and produce degenerate outputs. The notebook recommends a single epoch or even just a few hundred steps to achieve noticeable style shifts without compromising fluency.
- Beta Tuning: Values between 0.1 and 0.5 are standard. If the model becomes too repetitive or collapses to single-word answers, increase beta to strengthen the KL constraint against the reference model.
- Batch Composition: Ensure batches contain diverse prompts to prevent the policy from over-optimizing for specific response patterns.
Key Source Files
| File | Purpose | Location |
|---|---|---|
dpo-from-scratch.ipynb |
Complete implementation (dataset, loss, training) | ch07/04_preference-tuning-with-dpo/ |
generate.py |
Forward pass utilities used by both models | pkg/llms_from_scratch/generate.py |
previous_chapters.py |
Base InstructionDataset class |
ch07/02_dataset-utilities/previous_chapters.py |
ch07.ipynb |
Contextual SFT material and base training loop | ch07/01_main-chapter-code/ch07.ipynb |
Summary
- Direct Preference Optimization aligns LLMs using only pairwise preference data and a binary classification loss, eliminating the need for PPO or reward models.
- The implementation requires two model copies: a trainable policy and a frozen reference model loaded from the same SFT checkpoint.
- The
compute_dpo_lossfunction indpo-from-scratch.ipynb(lines 71-78) implements the core objective: maximizing the log-ratio gap between chosen and rejected responses relative to the reference. - Beta (β) controls optimization strength—typical values range from 0.1 to 0.5, with smaller values creating more aggressive alignment.
- Reward margins (
chosen_reward - rejected_reward) should be monitored to verify that the policy is successfully learning preferences. - DPO is prone to collapse; limit training to one epoch or fewer and use the logging utilities in
pkg/llms_from_scratch/utils.pyto track generation quality.
Frequently Asked Questions
How does DPO differ from RLHF for LLM alignment?
DPO eliminates the separate reward model and PPO optimization loop required by RLHF. Instead of learning a reward function and then optimizing against it, DPO directly optimizes the language model to satisfy preferences using a simple classification loss. This reduces memory requirements and training complexity while often achieving comparable alignment quality, as demonstrated in the rasbt/LLMs-from-scratch chapter 7 implementation.
What is the optimal beta value for DPO training?
The beta parameter typically ranges between 0.1 and 0.5, where smaller values apply stronger preference optimization and larger values keep the policy closer to the reference model. If you observe model collapse or repetitive outputs, increase beta to 0.3 or 0.5 to strengthen the implicit KL divergence constraint. The notebook uses 0.1 as a default starting point for moderate alignment pressure.
Why does DPO require a reference model?
The reference model provides a stable baseline representing the original supervised fine-tuned (SFT) policy. DPO optimizes the relative probability gap between chosen and rejected responses compared to this baseline, preventing the policy from diverging too far from coherent language generation. Without the reference model, the policy could arbitrarily increase probabilities for chosen outputs in ways that degrade overall text quality.
How many training steps are needed for effective DPO alignment?
DPO is highly sample-efficient and prone to overfitting, so one epoch or less is typically sufficient. The implementation in dpo-from-scratch.ipynb recommends training for just a few hundred steps or a single pass through the preference dataset. Monitoring the reward margin (chosen_r - rejected_r) helps determine convergence—once this value plateaus or starts decreasing, training should stop immediately to prevent degenerate outputs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →