# How to Perform Direct Preference Optimization (DPO) with the Soup CLI

> Learn how to perform Direct Preference Optimization DPO with the Soup CLI. Streamline LoRA, quantization, and iterative training using a simple YAML config.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Direct Preference Optimization (DPO) in the Soup CLI wraps the TRL DPOTrainer with a streaming-compatible wrapper that enables LoRA, quantization, and iterative training workflows via a single YAML configuration file.**

Direct Preference Optimization (DPO) is Soup’s built-in method for fine-tuning large language models on paired preference data without requiring a separate reward model. The Soup repository abstracts the complexity of the underlying TRL library behind a declarative command-line interface. This guide explains the architectural implementation of DPO in Soup and provides complete working examples for standard and iterative training workflows.

## Core DPO Architecture in Soup

### The DPOTrainerWrapper Implementation

The heart of Soup’s DPO implementation is the `DPOTrainerWrapper` class located in [[`src/soup_cli/trainer/dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/dpo.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/dpo.py#L24-L31). This wrapper extends `StreamingSetupMixin` to ensure memory-efficient model loading required for DPO’s three-forward-pass algorithm (policy, reference, and loss computation). The wrapper lazily imports `trl.DPOTrainer` and forwards training hyperparameters including `dpo_beta`, LoRA configurations, and quantization settings.

### Preference Task Dispatch

When you run `soup train --config <yaml>`, the preference dispatcher in [[`src/soup_cli/trainer/preference.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/preference.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/preference.py#L114-L117) instantiates the appropriate trainer based on the `task: dpo` declaration. The dispatcher validates that the data format is set to `"dpo"` as documented in the schema section of [[`docs/training.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md)](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L88-L99).

## Basic DPO Training Configuration

To execute DPO training, create a [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) file specifying the base model, task type, and preference dataset. The configuration schema requires `task: dpo` and `data.format: dpo` to trigger the preference optimization pipeline.

```yaml

# soup.yaml – basic DPO fine‑tuning

base: meta-llama/Llama-3.1-8B-Instruct
task: dpo
data:
  train: ./data/preferences.jsonl
  format: dpo
training:
  epochs: 3
  dpo_beta: 0.1
  lora:
    r: 64
    alpha: 16
  quantization: 4bit

```

Execute the training with:

```bash
soup train --config soup.yaml

```

## Iterative DPO for Continuous Improvement

For continuous improvement, Soup provides the `soup iterative-dpo` command that automates the generation of new preference pairs across multiple training rounds. As detailed in [[`docs/training.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md)](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L1-L10), this driver samples prompts, scores responses with a reward model, and rebuilds the preference dataset before resuming training.

Configuration requires specifying a reward model path alongside your standard DPO settings:

```yaml

# iterative_dpo.yaml

base: meta-llama/Llama-3.1-8B-Instruct
task: dpo
reward-model: ./output_rm
data:
  train: ./data/preferences.jsonl
  format: dpo
training:
  epochs: 1
  dpo_beta: 0.1
  lora:
    r: 64
    alpha: 16

```

Launch the iterative loop with:

```bash
soup iterative-dpo \
    --base-model meta-llama/Llama-3.1-8B-Instruct \
    --reward-model ./output_rm \
    --prompts ./prompts.jsonl \
    --output-dir ./iterative_dpo_out \
    --rounds 5 \
    --pairs-per-round 1000

```

## Advanced DPO Configuration Options

Soup supports KL-controlled schedules and mixed preference losses for research applications. The KL-controlled DPO variants documented in [[`docs/training.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md)](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L89-L107) allow dynamic adjustment of the beta parameter during training.

```yaml
training:
  dpo_beta: 0.1
  dpo_beta_schedule: cosine
  dpo_beta_end: 0.01
  dpo_ref_regen_epochs: 2
  preference_loss: dpo
  preference_loss_weights:
    dpo: 0.7
    simpo: 0.3

```

As noted in the preference variety section of [[`docs/training.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md)](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L110-L119), you can weight multiple preference loss functions simultaneously without changing the core task type.

## Summary

- **Direct Preference Optimization** in Soup is implemented via the `DPOTrainerWrapper` class in [`src/soup_cli/trainer/dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/dpo.py), which handles streaming, LoRA, and quantization.
- Training requires a YAML configuration with `task: dpo` and `data.format: dpo`, validated against the schema in [`docs/training.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md).
- The `soup train` command executes standard DPO, while `soup iterative-dpo` enables continuous learning loops with automatic preference pair generation.
- Advanced features include KL-controlled beta schedules and mixed loss functions for specialized research workflows.

## Frequently Asked Questions

### What data format is required for DPO training in Soup?

Soup requires preference data in JSONL format with `prompt`, `chosen`, and `rejected` fields. The configuration must specify `data.format: dpo` to trigger schema validation, as implemented in the preference validation logic documented in [[`docs/training.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md)](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L88-L99).

### How does Soup handle the reference model in DPO?

The `DPOTrainerWrapper` manages the reference model through the `StreamingSetupMixin`, ensuring it remains in evaluation mode during the three-forward-pass DPO loss calculation. The reference model can be refreshed periodically using the `dpo_ref_regen_epochs` parameter for KL-controlled variants.

### Can I combine DPO with LoRA and quantization?

Yes. Soup's architecture supports 4-bit and 8-bit quantization alongside LoRA adapters during DPO training. Configure these in the `training` block with `quantization: 4bit` and standard LoRA parameters, and the wrapper will initialize the TRL trainer with the appropriate PEFT configuration.

### What is the difference between standard DPO and iterative DPO in Soup?

Standard DPO (`soup train`) trains on a static preference dataset, while iterative DPO (`soup iterative-dpo`) runs a feedback loop that samples new prompts, generates responses, scores them with a reward model, and adds the resulting preference pairs to the training set for subsequent rounds.