How to Align LLM Preferences Using DPO, GRPO, and PPO

Direct Preference Optimization (DPO), Generalized Reinforcement Learning from Preference Optimization (GRPO), and Proximal Policy Optimization (PPO) are three reinforcement learning from human feedback (RLHF) methods that align large language models to human preferences, differing primarily in whether they require a separate reward model and how they apply KL-regularization to maintain proximity to the original supervised fine-tuned checkpoint.

Aligning LLM preferences is the critical second-stage step that follows supervised fine-tuning (SFT), transforming a base model into a helpful and safe assistant. The mlabonne/llm-course repository provides a comprehensive guide to these algorithms in the Preference Alignment section (lines 235–252 of README.md), detailing when to use each approach based on computational budget and safety requirements.

Three Algorithms for LLM Preference Alignment

Each method optimizes the policy against a dataset of human preferences—pairs of chosen and rejected completions—but they differ in architectural complexity and reward signal derivation.

Direct Preference Optimization (DPO)

DPO eliminates the need for a separate reward model by directly maximizing the likelihood of chosen responses over rejected ones. According to the repository, DPO optimizes the objective log σ(r_chosen – r_rejected) using the model's own logits, making it the fastest method for chat-style assistants where training budgets are limited. This approach is ideal for rapid experimentation when you want to avoid the overhead of training an auxiliary reward model.

Generalized Reinforcement Learning from Preference Optimization (GRPO)

GRPO extends DPO by incorporating a small on-policy reward model alongside a KL-penalty (KL(π_new || π_ref)) to enforce tighter control over the policy deviation. As noted at line 250 of the README, GRPO is best suited for scenarios requiring a careful balance between safety and quality, where a learned scalar reward signal provides finer-grained guidance than DPO's implicit preference margin.

Proximal Policy Optimization (PPO)

PPO represents the classic RLHF approach that first trains a dedicated reward model to score (prompt, response) pairs, then optimizes the policy using a clipped surrogate objective (L^{CLIP}) with advantage estimates. The repository recommends PPO for high-performance reasoning models where computational resources are abundant and fine-grained reward shaping is necessary to achieve state-of-the-art alignment.

Prerequisites: Dataset and Reference Policy Setup

All three methods require a preference dataset containing prompts paired with chosen and rejected completions. The rejection sampling technique described at line 39 of the README suggests generating multiple candidates per prompt before training, labeling the best as chosen and the worst as rejected to stabilize training.

Crucially, each algorithm maintains a reference policy—typically the original SFT checkpoint—to prevent catastrophic drift. This reference model provides the baseline for KL-divergence calculations that keep the aligned model from over-optimizing reward signals at the expense of general capabilities.

Implementation with the TRL Library

The mlabonne/llm-course recommends using the TRL (Transformer Reinforcement Learning) library, which provides ready-made trainers for all three algorithms as cited at line 237. The following examples assume you have installed trl and transformers.

DPO Training without a Reward Model

Use DPOTrainer when you want the simplest alignment pipeline. This example references the DPO tutorial linked at line 49:

from transformers import AutoModelForCausalLM, AutoTokenizer
from datasets import load_dataset
from trl import DPOTrainer

# Load the SFT checkpoint as both policy and reference

model_name = "mistralai/Mistral-7B-Instruct-v0.2"
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Preference dataset: JSON with "prompt", "chosen", "rejected" fields

dataset = load_dataset("json", data_files="preference_data.json")["train"]

# Initialize trainer with KL coefficient β

trainer = DPOTrainer(
    model=model,
    ref_model=model,          # Reference policy = SFT checkpoint

    tokenizer=tokenizer,
    train_dataset=dataset,
    max_steps=1000,
    beta=0.1,                 # KL-regularization strength

    optim="adamw_torch",
    lr_scheduler_type="cosine",
)

trainer.train()
trainer.save_model("./dpo_aligned_mistral")

GRPO with a Lightweight Reward Model

For GRPO, swap DPOTrainer for GRPOTrainer and supply a small reward model (e.g., 350M parameters):

from transformers import AutoModelForCausalLM, AutoTokenizer, AutoModelForSequenceClassification
from datasets import load_dataset
from trl import GRPOTrainer

# Policy and reference

policy = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2")

# Small reward model for scalar scoring

reward = AutoModelForSequenceClassification.from_pretrained(
    "OpenAssistant/reward-model-multilingual-llama-3.1-8b",
    device_map="auto"
)

dataset = load_dataset("json", data_files="preference_data.json")["train"]

trainer = GRPOTrainer(
    model=policy,
    ref_model=policy,
    reward_model=reward,
    tokenizer=tokenizer,
    train_dataset=dataset,
    max_steps=800,
    beta=0.15,                # KL coefficient

    gamma=0.99,               # Discount factor for advantage

)

trainer.train()
trainer.save_model("./grpo_aligned_mistral")

PPO Using Clipped Surrogate Objectives

PPO requires explicit configuration of clipping ranges and advantage estimation:

from transformers import AutoModelForCausalLM, AutoTokenizer, AutoModelForSequenceClassification
from datasets import load_dataset
from trl import PPOTrainer, PPOConfig

policy = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2", device_map="auto")
reward = AutoModelForSequenceClassification.from_pretrained(
    "OpenAssistant/reward-model-multilingual-llama-3.1-8b",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2")

dataset = load_dataset("json", data_files="pref_data_for_ppo.json")["train"]

# PPO-specific hyperparameters

ppo_config = PPOConfig(
    batch_size=32,
    ppo_epochs=4,
    kl_coef=0.2,
    clip_range=0.2,
    learning_rate=1.41e-5,
)

trainer = PPOTrainer(
    model=policy,
    ref_model=policy,
    reward_model=reward,
    tokenizer=tokenizer,
    train_dataset=dataset,
    config=ppo_config,
)

trainer.train()
trainer.save_pretrained("./ppo_aligned_mistral")

Practical Optimization Tips

Based on the implementation guidance in README.md, follow these recommendations to stabilize training:

  • KL-Penalty Tuning — Set beta between 0.1 and 0.2 for medium-sized models (7B–13B). Higher values keep the policy tighter to the SFT checkpoint but may limit helpfulness gains, while lower values allow more behavioral flexibility.
  • Reward Model Size — For GRPO and PPO, a lightweight reward model (350M–1B parameters) is sufficient; you do not need a full-scale LLM to provide adequate scalar rewards.
  • Evaluation Protocol — After alignment, evaluate on human-oriented benchmarks like Chatbot Arena to verify that the model remains both helpful and harmless, confirming that the KL-regularization successfully prevented reward hacking.

Summary

  • DPO is the fastest method to align LLM preferences, requiring no separate reward model and using implicit preference margins derived from model logits.
  • GRPO adds a small on-policy reward model and explicit KL-penalty for scenarios demanding tighter safety-quality trade-offs than DPO provides.
  • PPO offers maximum control through dedicated reward models and clipped surrogate objectives, suitable for resource-intensive reasoning tasks.
  • All methods rely on the TRL library's trainers (DPOTrainer, GRPOTrainer, PPOTrainer) and require a preference dataset with chosen/rejected pairs.
  • Maintain the SFT checkpoint as the reference policy across all three methods to prevent catastrophic drift during optimization.

Frequently Asked Questions

What is the main difference between DPO and PPO for LLM alignment?

DPO directly optimizes the policy against preference pairs using the Bradley-Terry model and the model's own logits, eliminating the need to train a separate reward model first. PPO, conversely, requires training a reward model to provide scalar scores, then uses these scores to compute advantages and optimize a clipped surrogate objective that constrains how far the policy can deviate from the reference.

When should I use GRPO instead of DPO?

Choose GRPO when your application requires a tighter safety-quality trade-off and benefits from a learned reward signal. GRPO combines DPO's data efficiency with an on-policy reward model and explicit KL-regularization, making it more robust than DPO for high-stakes deployments while remaining simpler than full PPO implementations.

Do I need a reward model to align LLM preferences with DPO?

No, DPO does not require a reward model. The algorithm directly maximizes the margin between chosen and rejected responses using the policy's own probability distributions, which makes it computationally efficient and ideal for rapid prototyping or environments with limited GPU resources.

What beta value should I use for KL-regularization?

The mlabonne/llm-course recommends a beta coefficient between 0.1 and 0.2 for KL-regularization when aligning medium-sized models. Values in this range effectively prevent the policy from drifting too far from the helpful SFT baseline without overly constraining the model's ability to adapt to preference data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →