Implementing RLHF and DPO Alignment Algorithms: From Theory to TinyGPT Code
Direct Preference Optimization (DPO) eliminates the need for separate reward models and PPO by deriving a closed-form loss that optimizes language model policies directly against human preference data, implemented here using a minimal TinyGPT architecture.
The ai-engineering-from-scratch repository provides a comprehensive, dependency-minimal curriculum for building large language models from first principles. When implementing RLHF and DPO alignment algorithms, the codebase demonstrates how to move from complex three-stage reinforcement learning pipelines to a simplified supervised approach using only torch and numpy. This implementation centers on a decoder-only transformer called TinyGPT, offering a complete reference for understanding alignment mechanics without production framework overhead.
RLHF vs. DPO: Architectural Differences
Traditional Reinforcement Learning from Human Feedback (RLHF) follows a three-stage pipeline: supervised fine-tuning (SFT) to create a reference policy, training an explicit reward model on preference pairs, and optimizing a new policy with PPO while enforcing a KL-divergence penalty against the reference.
Direct Preference Optimization bypasses the reward modeling stage entirely. By deriving the optimal policy under the Bradley-Terry preference model, DPO expresses the RLHF objective as a single supervised loss:
loss = -log(sigmoid(beta * ((log π_θ(y_w|x) - log π_ref(y_w|x)) - (log π_θ(y_l|x) - log π_ref(y_l|x)))))
This closed-form solution embeds the KL-constraint directly into the loss function, removing the need for PPO rollouts and value function estimation. The mathematical derivation and pedagogical explanation are documented in [phases/19-capstone-projects/40-dpo-from-scratch/docs/en.md](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/40-dpo-from-scratch/docs/en.md).
Core Components of the Implementation
TinyGPT Backbone
Both the reference and policy models instantiate the TinyGPT architecture defined in [phases/19-capstone-projects/39-llm-from-scratch/code/tinygpt.py](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/39-llm-from-scratch/code/tinygpt.py). This minimal decoder-only transformer provides the embedding layers, multi-head self-attention, and feed-forward networks necessary for sequence modeling while remaining small enough for educational experimentation.
Tokenization and Preference Data
The InstructionTokenizer adds special INST and RESP tokens to distinguish prompts from completions. The PreferenceDataset contains 12 (prompt, chosen, rejected) triples used for training, where chosen represents the preferred human response and rejected the dispreferred alternative.
Reference and Policy Model Setup
The implementation creates two model instances:
- Reference Model (π_ref): Loaded from the SFT checkpoint and frozen using
torch.no_grad()withrequires_grad=Falseon all parameters - Policy Model (π_θ): Initialized by copying the reference state dict, then optimized during training
This dual-model architecture ensures the KL-divergence term remains stable throughout optimization, as the reference logits never change.
The DPO Loss and Training Mechanics
Computing Sequence Log-Probabilities
The core utility function calculates the sum of log-probabilities for completion tokens while masking out the prompt. This implementation excludes prompt tokens from the loss calculation by identifying the completion start index:
def sequence_log_prob(model, tokenizer, prompt, completion):
# Tokenize full sequence
full_text = prompt + completion
tokens = tokenizer.encode(full_text)
input_ids = torch.tensor(tokens[:-1]).unsqueeze(0)
target_ids = torch.tensor(tokens[1:]).unsqueeze(0)
with torch.no_grad():
logits = model(input_ids) # (1, seq_len, vocab_size)
log_probs = torch.log_softmax(logits, dim=-1)
# Mask prompt tokens: only sum log-probs for completion indices
prompt_len = len(tokenizer.encode(prompt))
completion_log_probs = log_probs[0, prompt_len:, target_ids[0, prompt_len:]]
return completion_log_probs.sum()
Implementing the DPO Objective
The loss function combines log-ratios from both models and preference pairs into a binary classification objective:
def dpo_loss(logp_theta_w, logp_ref_w, logp_theta_l, logp_ref_l, beta=0.1):
"""
logp_theta_w: log π_θ(y_w|x) - policy log-prob of chosen completion
logp_ref_w: log π_ref(y_w|x) - reference log-prob of chosen completion
logp_theta_l: log π_θ(y_l|x) - policy log-prob of rejected completion
logp_ref_l: log π_ref(y_l|x) - reference log-prob of rejected completion
"""
# Log-ratio difference between chosen and rejected
chosen_ratio = logp_theta_w - logp_ref_w
rejected_ratio = logp_theta_l - logp_ref_l
# DPO objective: maximize margin between chosen and rejected
logits = beta * (chosen_ratio - rejected_ratio)
loss = -torch.logsigmoid(logits).mean()
return loss
The beta parameter controls the strength of the implicit KL-constraint, with typical values ranging from 0.1 to 0.5.
Training Loop Structure
The training iteration in [main.py](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/40-dpo-from-scratch/code/main.py) demonstrates the full forward pass:
for step in range(num_steps):
# Compute log-probs for both models and both completions
lp_theta_w = sequence_log_prob(policy, tokenizer, prompt, chosen)
lp_ref_w = sequence_log_prob(reference, tokenizer, prompt, chosen)
lp_theta_l = sequence_log_prob(policy, tokenizer, prompt, rejected)
lp_ref_l = sequence_log_prob(reference, tokenizer, prompt, rejected)
# Calculate DPO loss
loss = dpo_loss(lp_theta_w, lp_ref_w, lp_theta_l, lp_ref_l, beta=0.1)
# Backprop through policy only
loss.backward()
optimizer.step()
optimizer.zero_grad()
# Monitor convergence via chosen-rejected margin
margin = (lp_theta_w - lp_theta_l).mean().item()
print(f"Step {step}: Loss={loss.item():.4f}, Margin={margin:.3f}")
The margin metric indicates alignment quality—positive and increasing values show the policy assigning higher probability to chosen responses compared to rejected ones.
Why DPO Works: Computational and Theoretical Advantages
Elimination of PPO Infrastructure: DPO removes the need for replay buffers, advantage estimation, and clipping mechanisms required by proximal policy optimization. The implementation requires only a single forward pass through the policy and reference models per batch.
Exact Reference Invariance: Unlike RLHF where the KL-penalty is an approximation, DPO embeds the constraint analytically. The reference model remains completely frozen, preventing mode collapse and ensuring stable training dynamics.
Gradient Direction Guarantees: The loss gradient with respect to log π_θ(chosen|x) is always negative, mathematically ensuring the policy increases the likelihood of preferred completions while decreasing rejected ones, as detailed in the gradient sign analysis within the lesson documentation.
Key Source Files for Deep Dive
| File | Description | Link |
|---|---|---|
main.py |
Full implementation including TinyGPT wrapper, tokenizer, dataset, loss functions, and training loop | View Source |
docs/en.md |
Mathematical derivation of DPO from the RLHF objective, Bradley-Terry model explanation, and learning objectives | View Source |
tinygpt.py |
Shared transformer architecture used by both reference and policy models | View Source |
tests/test_main.py |
Unit tests verifying loss mathematics, gradient signs, and reference model invariance | View Source |
dpo-family/docs/en.md |
Overview of DPO variants (IPO, KTO, ORPO) and their relationship to classical RLHF | View Source |
Summary
- DPO simplifies alignment by deriving a supervised loss that directly optimizes against preference pairs without requiring separate reward models or PPO.
- The implementation uses a frozen reference model and a trainable policy model, both instantiated from the TinyGPT architecture.
- Sequence log-probability masking ensures only completion tokens contribute to the loss, excluding prompt tokens from gradient calculations.
- The DPO loss function implements a sigmoid over log-ratio differences, mathematically equivalent to the RLHF objective with embedded KL-constraints.
- This educational codebase demonstrates that alignment algorithms can be implemented with minimal dependencies—only
torchandnumpy—while maintaining theoretical rigor.
Frequently Asked Questions
What is the primary difference between RLHF and DPO in this implementation?
RLHF requires training an explicit reward model and then optimizing the policy with PPO against that reward, while DPO eliminates the intermediate reward model entirely. In the codebase, DPO calculates the loss directly from the log-ratio between policy and reference model probabilities, reducing the pipeline from three stages to a single supervised learning step.
Why does the implementation freeze the reference model during training?
The reference model represents the SFT baseline that the optimized policy should not diverge from significantly. By freezing π_ref with torch.no_grad() and requires_grad=False, the implementation ensures that the KL-divergence term in the DPO objective remains stable and that gradients flow only through the policy model parameters during backpropagation.
How does the sequence_log_prob function handle prompt masking?
The function tokenizes the combined prompt and completion, then identifies the boundary between them using len(tokenizer.encode(prompt)). When summing log-probabilities, it slices the tensor to include only indices corresponding to completion tokens, ensuring that the model is not rewarded or penalized for probability mass assigned to the prompt itself.
Can this TinyGPT-based DPO implementation scale to production-sized models?
While TinyGPT serves as an educational vehicle for understanding alignment mechanics, the DPO loss function and training loop architecture are model-agnostic. The same dpo_loss calculation and reference-policy comparison logic applies to billion-parameter transformers, though production implementations would require distributed training strategies and optimized attention kernels not included in this minimal educational codebase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →