How to Enable Layer Streaming for Preference Losses (DPO, GRPO) in Soup
Layer streaming for preference losses lets you fine-tune models larger than your GPU VRAM by streaming decoder layers from CPU RAM or disk, with support for DPO, KTO, ORPO, and SimPO starting in Soup v0.72.4.
Soup's layer streaming mechanism bridges the gap between memory-constrained hardware and large model fine-tuning. While standard supervised fine-tuning (SFT) has supported streaming since earlier versions, v0.72.4 extends this capability to reference-free preference losses—enabling Direct Preference Optimization (DPO), Kahneman-Tversky Optimization (KTO), Odds Ratio Preference Optimization (ORPO), and SimPO. Understanding how to configure and activate this feature requires navigating specific constraints around model architecture, batch sizing, and trainer implementation.
What Is Layer Streaming for Preference Losses?
Layer streaming keeps the frozen base model in CPU RAM or on disk, loading only one decoder layer at a time into GPU VRAM for the forward and backward passes. For preference losses, this is particularly valuable because these methods traditionally require either:
- A separate reference model (standard DPO), or
- Dual forward passes (comparing chosen vs. rejected completions)
Soup eliminates the separate reference model by re-using the same streamed base with LoRA adapters disabled, saving approximately 730 MB of extra weights according to the "Preference losses over streaming" documentation table【source】.
Why GRPO Is Excluded
GRPO (Group Relative Policy Optimization) and PPO are deliberately excluded from streaming support. These methods require per-token generation with repeated model reads for roll-outs, which would nullify the amortization benefits of layer-by-layer streaming【source】.
Configuration Steps to Enable Streaming for DPO/KTO
Step 1: Activate the Streaming Flag
Set stream_layers: true in the training section of your soup.yaml.
training:
stream_layers: true
This flag activates the BETA streaming runtime as defined in the schema【source】. Without it, the base model loads fully resident in VRAM.
Step 2: Select a Supported Task
Use task: dpo, task: kto, task: orpo, or task: simpo.
base: meta-llama/Llama-3.1-8B-Instruct
task: dpo # ← preference loss with streaming support
Preference-loss trainers only accept streaming for tasks that do not require a second full model instance. The streaming code path in src/soup_cli/trainer/dpo.py re-uses the frozen base with adapters disabled【source】.
Step 3: Observe Streaming Constraints
The config validator _validate_stream_layers_compat enforces these requirements【source】:
| Constraint | Required Value | Rationale |
|---|---|---|
| Backend | transformers |
Streaming only implemented for HF Transformers |
| Modality | text |
Vision/audio layers not yet supported |
| Quantization | none or 4bit (NF4) |
NF4 recommended for streaming efficiency |
| Batch size | 1 for DPO/ORPO/SimPO; ≥2 for KTO |
KTO's batching needs |
| LoRA | Enabled (lora.r ≥ 1) with init_strategy: random |
Required for adapter-based streaming |
Violating any constraint raises a clear error at config-load time.
Step 4: Tune Streaming Source and Buffering
training:
stream_source: auto # auto | ram | disk
stream_buffers: 2 # double-buffering for I/O overlap
autoselects RAM when weights fit, otherwise falls back to NVMe overflow tierstream_bufferscontrols prefetching for compute/transfer overlap
Step 5: Run Training
soup train --config soup.yaml
No extra CLI flag is needed—the streaming behavior triggers purely from configuration.
Complete Configuration Examples
DPO with Layer Streaming (Batch Size 1)
# soup.yaml – DPO with layer streaming
base: meta-llama/Llama-3.1-8B-Instruct
task: dpo
backend: transformers
modality: text
data:
train: ./data/train.jsonl
format: alpaca
max_length: 512
val_split: 0.1
training:
epochs: 3
lr: 2e-5
batch_size: 1
lora:
r: 64
alpha: 16
init_strategy: random
quantization: 4bit
stream_layers: true
stream_source: auto
stream_buffers: 2
KTO with Larger Batch (Requires ≥2)
training:
epochs: 3
lr: 2e-5
batch_size: 2 # KTO minimum batch size
lora:
r: 64
alpha: 16
init_strategy: random
quantization: 4bit
stream_layers: true
stream_source: auto
stream_buffers: 2
Memory and Performance Characteristics
VRAM Preflight Check
The pre-flight check in the streaming planner predicts VRAM usage and refuses configurations that would exceed budget. For preference losses, the estimate is conservative because the loss packs both chosen and rejected tokens—treating this as an upper bound【source】.
Time vs. Memory Trade-off
| Aspect | Behavior |
|---|---|
| Time penalty | ~1.5× layer reads for DPO vs. SFT (extra comparison pass) |
| VRAM overhead | None—primary benefit on low-memory GPUs |
| Reference model | Re-used streamed base, no separate copy |
Optional Tuning Parameters
training:
# stream_vram_probe: true # Measure actual VRAM instead of formula estimate
# stream_disk_kind: nvme # Force NVMe tier if auto-detection fails
Key Source Files and Implementation
| File | Responsibility |
|---|---|
src/soup_cli/config/schema.py |
stream_layers definition and validation schema【source】 |
src/soup_cli/utils/layer_stream.py |
Planner building the buffered streaming pipeline【source】 |
src/soup_cli/utils/layer_stream_runtime.py |
Runtime copying layers from CPU RAM/disk to VRAM【source】 |
src/soup_cli/trainer/dpo.py |
DPO trainer respecting stream_layers flag【source】 |
docs/performance-and-quantization.md |
Performance numbers and compatibility matrix【source】 |
Summary
- Layer streaming for DPO/KTO activates with
stream_layers: truein Soup v0.72.4+ - GRPO is unsupported due to per-token generation requirements that defeat streaming amortization
- No separate reference model—Soup re-uses the streamed base with LoRA disabled, saving ~730 MB
- Strict constraints apply: transformers backend, text modality, 4bit or no quantization, specific batch sizes, and LoRA with random initialization
- Performance trade-off: ~1.5× time penalty for DPO vs. SFT, but zero VRAM overhead for the reference mechanism
Frequently Asked Questions
Does layer streaming work with GRPO in Soup?
No. GRPO and PPO require per-token generation with repeated model reads for roll-outs, which would nullify the layer-by-layer streaming advantage. GRPO is explicitly excluded from the streaming roadmap【source】.
How much VRAM does streaming DPO save compared to standard DPO?
Approximately 730 MB is saved by eliminating the separate reference model. Standard DPO loads both the policy and reference models; Soup's streaming implementation re-uses the same frozen base with LoRA adapters disabled, so only one set of base weights streams through memory【source】.
Why does KTO require batch size ≥2 when DPO uses batch size 1?
KTO's loss formulation mathematically requires paired or grouped examples to estimate the reference point for the KL divergence term. The config validator _validate_stream_layers_compat enforces batch_size ≥ 2 for task: kto specifically【source】.
Can I use 8-bit quantization with layer streaming?
No. The validator restricts quantization to none or 4bit (NF4). NF4 is explicitly recommended for streaming because its dequantization overhead is amortized across the layer computation, whereas 8-bit formats introduce memory alignment constraints that complicate the streaming buffer management【source】.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →