How to Enable Layer Streaming for Preference Losses (DPO, GRPO) in Soup

Layer streaming for preference losses lets you fine-tune models larger than your GPU VRAM by streaming decoder layers from CPU RAM or disk, with support for DPO, KTO, ORPO, and SimPO starting in Soup v0.72.4.

Soup's layer streaming mechanism bridges the gap between memory-constrained hardware and large model fine-tuning. While standard supervised fine-tuning (SFT) has supported streaming since earlier versions, v0.72.4 extends this capability to reference-free preference losses—enabling Direct Preference Optimization (DPO), Kahneman-Tversky Optimization (KTO), Odds Ratio Preference Optimization (ORPO), and SimPO. Understanding how to configure and activate this feature requires navigating specific constraints around model architecture, batch sizing, and trainer implementation.

What Is Layer Streaming for Preference Losses?

Layer streaming keeps the frozen base model in CPU RAM or on disk, loading only one decoder layer at a time into GPU VRAM for the forward and backward passes. For preference losses, this is particularly valuable because these methods traditionally require either:

  • A separate reference model (standard DPO), or
  • Dual forward passes (comparing chosen vs. rejected completions)

Soup eliminates the separate reference model by re-using the same streamed base with LoRA adapters disabled, saving approximately 730 MB of extra weights according to the "Preference losses over streaming" documentation table【source】.

Why GRPO Is Excluded

GRPO (Group Relative Policy Optimization) and PPO are deliberately excluded from streaming support. These methods require per-token generation with repeated model reads for roll-outs, which would nullify the amortization benefits of layer-by-layer streaming【source】.

Configuration Steps to Enable Streaming for DPO/KTO

Step 1: Activate the Streaming Flag

Set stream_layers: true in the training section of your soup.yaml.

training:
  stream_layers: true

This flag activates the BETA streaming runtime as defined in the schema【source】. Without it, the base model loads fully resident in VRAM.

Step 2: Select a Supported Task

Use task: dpo, task: kto, task: orpo, or task: simpo.

base: meta-llama/Llama-3.1-8B-Instruct
task: dpo          # ← preference loss with streaming support

Preference-loss trainers only accept streaming for tasks that do not require a second full model instance. The streaming code path in src/soup_cli/trainer/dpo.py re-uses the frozen base with adapters disabled【source】.

Step 3: Observe Streaming Constraints

The config validator _validate_stream_layers_compat enforces these requirements【source】:

Constraint Required Value Rationale
Backend transformers Streaming only implemented for HF Transformers
Modality text Vision/audio layers not yet supported
Quantization none or 4bit (NF4) NF4 recommended for streaming efficiency
Batch size 1 for DPO/ORPO/SimPO; ≥2 for KTO KTO's batching needs
LoRA Enabled (lora.r ≥ 1) with init_strategy: random Required for adapter-based streaming

Violating any constraint raises a clear error at config-load time.

Step 4: Tune Streaming Source and Buffering

training:
  stream_source: auto      # auto | ram | disk

  stream_buffers: 2        # double-buffering for I/O overlap
  • auto selects RAM when weights fit, otherwise falls back to NVMe overflow tier
  • stream_buffers controls prefetching for compute/transfer overlap

Step 5: Run Training

soup train --config soup.yaml

No extra CLI flag is needed—the streaming behavior triggers purely from configuration.

Complete Configuration Examples

DPO with Layer Streaming (Batch Size 1)


# soup.yaml – DPO with layer streaming

base: meta-llama/Llama-3.1-8B-Instruct
task: dpo
backend: transformers
modality: text

data:
  train: ./data/train.jsonl
  format: alpaca
  max_length: 512
  val_split: 0.1

training:
  epochs: 3
  lr: 2e-5
  batch_size: 1
  lora:
    r: 64
    alpha: 16
    init_strategy: random
  quantization: 4bit
  stream_layers: true
  stream_source: auto
  stream_buffers: 2

KTO with Larger Batch (Requires ≥2)

training:
  epochs: 3
  lr: 2e-5
  batch_size: 2          # KTO minimum batch size

  lora:
    r: 64
    alpha: 16
    init_strategy: random
  quantization: 4bit
  stream_layers: true
  stream_source: auto
  stream_buffers: 2

Memory and Performance Characteristics

VRAM Preflight Check

The pre-flight check in the streaming planner predicts VRAM usage and refuses configurations that would exceed budget. For preference losses, the estimate is conservative because the loss packs both chosen and rejected tokens—treating this as an upper bound【source】.

Time vs. Memory Trade-off

Aspect Behavior
Time penalty ~1.5× layer reads for DPO vs. SFT (extra comparison pass)
VRAM overhead None—primary benefit on low-memory GPUs
Reference model Re-used streamed base, no separate copy

Optional Tuning Parameters

training:
  # stream_vram_probe: true    # Measure actual VRAM instead of formula estimate

  # stream_disk_kind: nvme     # Force NVMe tier if auto-detection fails

Key Source Files and Implementation

File Responsibility
src/soup_cli/config/schema.py stream_layers definition and validation schema【source】
src/soup_cli/utils/layer_stream.py Planner building the buffered streaming pipeline【source】
src/soup_cli/utils/layer_stream_runtime.py Runtime copying layers from CPU RAM/disk to VRAM【source】
src/soup_cli/trainer/dpo.py DPO trainer respecting stream_layers flag【source】
docs/performance-and-quantization.md Performance numbers and compatibility matrix【source】

Summary

  • Layer streaming for DPO/KTO activates with stream_layers: true in Soup v0.72.4+
  • GRPO is unsupported due to per-token generation requirements that defeat streaming amortization
  • No separate reference model—Soup re-uses the streamed base with LoRA disabled, saving ~730 MB
  • Strict constraints apply: transformers backend, text modality, 4bit or no quantization, specific batch sizes, and LoRA with random initialization
  • Performance trade-off: ~1.5× time penalty for DPO vs. SFT, but zero VRAM overhead for the reference mechanism

Frequently Asked Questions

Does layer streaming work with GRPO in Soup?

No. GRPO and PPO require per-token generation with repeated model reads for roll-outs, which would nullify the layer-by-layer streaming advantage. GRPO is explicitly excluded from the streaming roadmap【source】.

How much VRAM does streaming DPO save compared to standard DPO?

Approximately 730 MB is saved by eliminating the separate reference model. Standard DPO loads both the policy and reference models; Soup's streaming implementation re-uses the same frozen base with LoRA adapters disabled, so only one set of base weights streams through memory【source】.

Why does KTO require batch size ≥2 when DPO uses batch size 1?

KTO's loss formulation mathematically requires paired or grouped examples to estimate the reference point for the KL divergence term. The config validator _validate_stream_layers_compat enforces batch_size ≥ 2 for task: kto specifically【source】.

Can I use 8-bit quantization with layer streaming?

No. The validator restricts quantization to none or 4bit (NF4). NF4 is explicitly recommended for streaming because its dequantization overhead is amortized across the layer computation, whereas 8-bit formats introduce memory alignment constraints that complicate the streaming buffer management【source】.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →