How to Perform Direct Preference Optimization (DPO) with the Soup CLI
Direct Preference Optimization (DPO) in the Soup CLI wraps the TRL DPOTrainer with a streaming-compatible wrapper that enables LoRA, quantization, and iterative training workflows via a single YAML configuration file.
Direct Preference Optimization (DPO) is Soup’s built-in method for fine-tuning large language models on paired preference data without requiring a separate reward model. The Soup repository abstracts the complexity of the underlying TRL library behind a declarative command-line interface. This guide explains the architectural implementation of DPO in Soup and provides complete working examples for standard and iterative training workflows.
Core DPO Architecture in Soup
The DPOTrainerWrapper Implementation
The heart of Soup’s DPO implementation is the DPOTrainerWrapper class located in [src/soup_cli/trainer/dpo.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/dpo.py#L24-L31). This wrapper extends StreamingSetupMixin to ensure memory-efficient model loading required for DPO’s three-forward-pass algorithm (policy, reference, and loss computation). The wrapper lazily imports trl.DPOTrainer and forwards training hyperparameters including dpo_beta, LoRA configurations, and quantization settings.
Preference Task Dispatch
When you run soup train --config <yaml>, the preference dispatcher in [src/soup_cli/trainer/preference.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/preference.py#L114-L117) instantiates the appropriate trainer based on the task: dpo declaration. The dispatcher validates that the data format is set to "dpo" as documented in the schema section of [docs/training.md](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L88-L99).
Basic DPO Training Configuration
To execute DPO training, create a soup.yaml file specifying the base model, task type, and preference dataset. The configuration schema requires task: dpo and data.format: dpo to trigger the preference optimization pipeline.
# soup.yaml – basic DPO fine‑tuning
base: meta-llama/Llama-3.1-8B-Instruct
task: dpo
data:
train: ./data/preferences.jsonl
format: dpo
training:
epochs: 3
dpo_beta: 0.1
lora:
r: 64
alpha: 16
quantization: 4bit
Execute the training with:
soup train --config soup.yaml
Iterative DPO for Continuous Improvement
For continuous improvement, Soup provides the soup iterative-dpo command that automates the generation of new preference pairs across multiple training rounds. As detailed in [docs/training.md](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L1-L10), this driver samples prompts, scores responses with a reward model, and rebuilds the preference dataset before resuming training.
Configuration requires specifying a reward model path alongside your standard DPO settings:
# iterative_dpo.yaml
base: meta-llama/Llama-3.1-8B-Instruct
task: dpo
reward-model: ./output_rm
data:
train: ./data/preferences.jsonl
format: dpo
training:
epochs: 1
dpo_beta: 0.1
lora:
r: 64
alpha: 16
Launch the iterative loop with:
soup iterative-dpo \
--base-model meta-llama/Llama-3.1-8B-Instruct \
--reward-model ./output_rm \
--prompts ./prompts.jsonl \
--output-dir ./iterative_dpo_out \
--rounds 5 \
--pairs-per-round 1000
Advanced DPO Configuration Options
Soup supports KL-controlled schedules and mixed preference losses for research applications. The KL-controlled DPO variants documented in [docs/training.md](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L89-L107) allow dynamic adjustment of the beta parameter during training.
training:
dpo_beta: 0.1
dpo_beta_schedule: cosine
dpo_beta_end: 0.01
dpo_ref_regen_epochs: 2
preference_loss: dpo
preference_loss_weights:
dpo: 0.7
simpo: 0.3
As noted in the preference variety section of [docs/training.md](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L110-L119), you can weight multiple preference loss functions simultaneously without changing the core task type.
Summary
- Direct Preference Optimization in Soup is implemented via the
DPOTrainerWrapperclass insrc/soup_cli/trainer/dpo.py, which handles streaming, LoRA, and quantization. - Training requires a YAML configuration with
task: dpoanddata.format: dpo, validated against the schema indocs/training.md. - The
soup traincommand executes standard DPO, whilesoup iterative-dpoenables continuous learning loops with automatic preference pair generation. - Advanced features include KL-controlled beta schedules and mixed loss functions for specialized research workflows.
Frequently Asked Questions
What data format is required for DPO training in Soup?
Soup requires preference data in JSONL format with prompt, chosen, and rejected fields. The configuration must specify data.format: dpo to trigger schema validation, as implemented in the preference validation logic documented in [docs/training.md](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/training.md#L88-L99).
How does Soup handle the reference model in DPO?
The DPOTrainerWrapper manages the reference model through the StreamingSetupMixin, ensuring it remains in evaluation mode during the three-forward-pass DPO loss calculation. The reference model can be refreshed periodically using the dpo_ref_regen_epochs parameter for KL-controlled variants.
Can I combine DPO with LoRA and quantization?
Yes. Soup's architecture supports 4-bit and 8-bit quantization alongside LoRA adapters during DPO training. Configure these in the training block with quantization: 4bit and standard LoRA parameters, and the wrapper will initialize the TRL trainer with the appropriate PEFT configuration.
What is the difference between standard DPO and iterative DPO in Soup?
Standard DPO (soup train) trains on a static preference dataset, while iterative DPO (soup iterative-dpo) runs a feedback loop that samples new prompts, generates responses, scores them with a reward model, and adds the resulting preference pairs to the training set for subsequent rounds.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →