# How to Align LLM Preferences Using DPO, GRPO, and PPO

> Learn how to align LLM preferences using DPO, GRPO, and PPO. Explore these RLHF methods for better language model control and understand their key differences for optimal model alignment.

- Repository: [Maxime Labonne/llm-course](https://github.com/mlabonne/llm-course)
- Tags: deep-dive
- Published: 2026-03-01

---

**Direct Preference Optimization (DPO), Generalized Reinforcement Learning from Preference Optimization (GRPO), and Proximal Policy Optimization (PPO) are three reinforcement learning from human feedback (RLHF) methods that align large language models to human preferences, differing primarily in whether they require a separate reward model and how they apply KL-regularization to maintain proximity to the original supervised fine-tuned checkpoint.**

Aligning LLM preferences is the critical second-stage step that follows supervised fine-tuning (SFT), transforming a base model into a helpful and safe assistant. The **mlabonne/llm-course** repository provides a comprehensive guide to these algorithms in the [Preference Alignment section](https://github.com/mlabonne/llm-course/blob/main/README.md#L235) (lines 235–252 of [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md)), detailing when to use each approach based on computational budget and safety requirements.

## Three Algorithms for LLM Preference Alignment

Each method optimizes the policy against a dataset of human preferences—pairs of chosen and rejected completions—but they differ in architectural complexity and reward signal derivation.

### Direct Preference Optimization (DPO)

**DPO** eliminates the need for a separate reward model by directly maximizing the likelihood of chosen responses over rejected ones. According to the repository, DPO optimizes the objective `log σ(r_chosen – r_rejected)` using the model's own logits, making it the fastest method for chat-style assistants where training budgets are limited. This approach is ideal for rapid experimentation when you want to avoid the overhead of training an auxiliary reward model.

### Generalized Reinforcement Learning from Preference Optimization (GRPO)

**GRPO** extends DPO by incorporating a small on-policy reward model alongside a KL-penalty (`KL(π_new || π_ref)`) to enforce tighter control over the policy deviation. As noted at [line 250 of the README](https://github.com/mlabonne/llm-course/blob/main/README.md#L250), GRPO is best suited for scenarios requiring a careful balance between safety and quality, where a learned scalar reward signal provides finer-grained guidance than DPO's implicit preference margin.

### Proximal Policy Optimization (PPO)

**PPO** represents the classic RLHF approach that first trains a dedicated reward model to score (prompt, response) pairs, then optimizes the policy using a clipped surrogate objective (`L^{CLIP}`) with advantage estimates. The repository recommends PPO for high-performance reasoning models where computational resources are abundant and fine-grained reward shaping is necessary to achieve state-of-the-art alignment.

## Prerequisites: Dataset and Reference Policy Setup

All three methods require a **preference dataset** containing prompts paired with chosen and rejected completions. The [rejection sampling technique](https://github.com/mlabonne/llm-course/blob/main/README.md#L39) described at line 39 of the README suggests generating multiple candidates per prompt before training, labeling the best as *chosen* and the worst as *rejected* to stabilize training.

Crucially, each algorithm maintains a **reference policy**—typically the original SFT checkpoint—to prevent catastrophic drift. This reference model provides the baseline for KL-divergence calculations that keep the aligned model from over-optimizing reward signals at the expense of general capabilities.

## Implementation with the TRL Library

The mlabonne/llm-course recommends using the **TRL** (Transformer Reinforcement Learning) library, which provides ready-made trainers for all three algorithms as cited at [line 237](https://github.com/mlabonne/llm-course/blob/main/README.md#L237). The following examples assume you have installed `trl` and `transformers`.

### DPO Training without a Reward Model

Use `DPOTrainer` when you want the simplest alignment pipeline. This example references the [DPO tutorial](https://github.com/mlabonne/llm-course/blob/main/README.md#L49) linked at line 49:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from datasets import load_dataset
from trl import DPOTrainer

# Load the SFT checkpoint as both policy and reference

model_name = "mistralai/Mistral-7B-Instruct-v0.2"
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Preference dataset: JSON with "prompt", "chosen", "rejected" fields

dataset = load_dataset("json", data_files="preference_data.json")["train"]

# Initialize trainer with KL coefficient β

trainer = DPOTrainer(
    model=model,
    ref_model=model,          # Reference policy = SFT checkpoint

    tokenizer=tokenizer,
    train_dataset=dataset,
    max_steps=1000,
    beta=0.1,                 # KL-regularization strength

    optim="adamw_torch",
    lr_scheduler_type="cosine",
)

trainer.train()
trainer.save_model("./dpo_aligned_mistral")

```

### GRPO with a Lightweight Reward Model

For GRPO, swap `DPOTrainer` for `GRPOTrainer` and supply a small reward model (e.g., 350M parameters):

```python
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoModelForSequenceClassification
from datasets import load_dataset
from trl import GRPOTrainer

# Policy and reference

policy = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2")

# Small reward model for scalar scoring

reward = AutoModelForSequenceClassification.from_pretrained(
    "OpenAssistant/reward-model-multilingual-llama-3.1-8b",
    device_map="auto"
)

dataset = load_dataset("json", data_files="preference_data.json")["train"]

trainer = GRPOTrainer(
    model=policy,
    ref_model=policy,
    reward_model=reward,
    tokenizer=tokenizer,
    train_dataset=dataset,
    max_steps=800,
    beta=0.15,                # KL coefficient

    gamma=0.99,               # Discount factor for advantage

)

trainer.train()
trainer.save_model("./grpo_aligned_mistral")

```

### PPO Using Clipped Surrogate Objectives

PPO requires explicit configuration of clipping ranges and advantage estimation:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoModelForSequenceClassification
from datasets import load_dataset
from trl import PPOTrainer, PPOConfig

policy = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2", device_map="auto")
reward = AutoModelForSequenceClassification.from_pretrained(
    "OpenAssistant/reward-model-multilingual-llama-3.1-8b",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2")

dataset = load_dataset("json", data_files="pref_data_for_ppo.json")["train"]

# PPO-specific hyperparameters

ppo_config = PPOConfig(
    batch_size=32,
    ppo_epochs=4,
    kl_coef=0.2,
    clip_range=0.2,
    learning_rate=1.41e-5,
)

trainer = PPOTrainer(
    model=policy,
    ref_model=policy,
    reward_model=reward,
    tokenizer=tokenizer,
    train_dataset=dataset,
    config=ppo_config,
)

trainer.train()
trainer.save_pretrained("./ppo_aligned_mistral")

```

## Practical Optimization Tips

Based on the implementation guidance in [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md), follow these recommendations to stabilize training:

*   **KL-Penalty Tuning** — Set `beta` between **0.1 and 0.2** for medium-sized models (7B–13B). Higher values keep the policy tighter to the SFT checkpoint but may limit helpfulness gains, while lower values allow more behavioral flexibility.
*   **Reward Model Size** — For GRPO and PPO, a lightweight reward model (350M–1B parameters) is sufficient; you do not need a full-scale LLM to provide adequate scalar rewards.
*   **Evaluation Protocol** — After alignment, evaluate on human-oriented benchmarks like Chatbot Arena to verify that the model remains both helpful and harmless, confirming that the KL-regularization successfully prevented reward hacking.

## Summary

*   **DPO** is the fastest method to align LLM preferences, requiring no separate reward model and using implicit preference margins derived from model logits.
*   **GRPO** adds a small on-policy reward model and explicit KL-penalty for scenarios demanding tighter safety-quality trade-offs than DPO provides.
*   **PPO** offers maximum control through dedicated reward models and clipped surrogate objectives, suitable for resource-intensive reasoning tasks.
*   All methods rely on the **TRL** library's trainers (`DPOTrainer`, `GRPOTrainer`, `PPOTrainer`) and require a preference dataset with chosen/rejected pairs.
*   Maintain the **SFT checkpoint** as the reference policy across all three methods to prevent catastrophic drift during optimization.

## Frequently Asked Questions

### What is the main difference between DPO and PPO for LLM alignment?

DPO directly optimizes the policy against preference pairs using the Bradley-Terry model and the model's own logits, eliminating the need to train a separate reward model first. PPO, conversely, requires training a reward model to provide scalar scores, then uses these scores to compute advantages and optimize a clipped surrogate objective that constrains how far the policy can deviate from the reference.

### When should I use GRPO instead of DPO?

Choose GRPO when your application requires a tighter safety-quality trade-off and benefits from a learned reward signal. GRPO combines DPO's data efficiency with an on-policy reward model and explicit KL-regularization, making it more robust than DPO for high-stakes deployments while remaining simpler than full PPO implementations.

### Do I need a reward model to align LLM preferences with DPO?

No, DPO does not require a reward model. The algorithm directly maximizes the margin between chosen and rejected responses using the policy's own probability distributions, which makes it computationally efficient and ideal for rapid prototyping or environments with limited GPU resources.

### What beta value should I use for KL-regularization?

The mlabonne/llm-course recommends a **beta coefficient between 0.1 and 0.2** for KL-regularization when aligning medium-sized models. Values in this range effectively prevent the policy from drifting too far from the helpful SFT baseline without overly constraining the model's ability to adapt to preference data.