# How to Enable Layer Streaming for Preference Losses (DPO, GRPO) in Soup

> Learn to enable layer streaming for preference losses like DPO and GRPO in Soup v0.72.4. Fine-tune larger models by streaming decoder layers from CPU RAM or disk.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-08-16

---

**Layer streaming for preference losses lets you fine-tune models larger than your GPU VRAM by streaming decoder layers from CPU RAM or disk, with support for DPO, KTO, ORPO, and SimPO starting in Soup v0.72.4.**

Soup's layer streaming mechanism bridges the gap between memory-constrained hardware and large model fine-tuning. While standard supervised fine-tuning (SFT) has supported streaming since earlier versions, **v0.72.4** extends this capability to reference-free preference losses—enabling Direct Preference Optimization (DPO), Kahneman-Tversky Optimization (KTO), Odds Ratio Preference Optimization (ORPO), and SimPO. Understanding how to configure and activate this feature requires navigating specific constraints around model architecture, batch sizing, and trainer implementation.

## What Is Layer Streaming for Preference Losses?

Layer streaming keeps the frozen base model in **CPU RAM or on disk**, loading only **one decoder layer at a time** into GPU VRAM for the forward and backward passes. For preference losses, this is particularly valuable because these methods traditionally require either:

- A **separate reference model** (standard DPO), or
- **Dual forward passes** (comparing chosen vs. rejected completions)

Soup eliminates the separate reference model by **re-using the same streamed base with LoRA adapters disabled**, saving approximately **730 MB** of extra weights according to the "Preference losses over streaming" documentation table【source】.

### Why GRPO Is Excluded

GRPO (Group Relative Policy Optimization) and PPO are **deliberately excluded** from streaming support. These methods require **per-token generation** with repeated model reads for roll-outs, which would nullify the amortization benefits of layer-by-layer streaming【source】.

## Configuration Steps to Enable Streaming for DPO/KTO

### Step 1: Activate the Streaming Flag

Set `stream_layers: true` in the `training` section of your [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml).

```yaml
training:
  stream_layers: true

```

This flag activates the BETA streaming runtime as defined in the schema【source】. Without it, the base model loads fully resident in VRAM.

### Step 2: Select a Supported Task

Use `task: dpo`, `task: kto`, `task: orpo`, or `task: simpo`.

```yaml
base: meta-llama/Llama-3.1-8B-Instruct
task: dpo          # ← preference loss with streaming support

```

Preference-loss trainers only accept streaming for tasks that **do not require a second full model instance**. The streaming code path in [`src/soup_cli/trainer/dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/dpo.py) re-uses the frozen base with adapters disabled【source】.

### Step 3: Observe Streaming Constraints

The config validator `_validate_stream_layers_compat` enforces these requirements【source】:

| Constraint | Required Value | Rationale |
|------------|----------------|-----------|
| Backend | `transformers` | Streaming only implemented for HF Transformers |
| Modality | `text` | Vision/audio layers not yet supported |
| Quantization | `none` or `4bit` (NF4) | NF4 recommended for streaming efficiency |
| Batch size | `1` for DPO/ORPO/SimPO; `≥2` for KTO | KTO's batching needs |
| LoRA | Enabled (`lora.r ≥ 1`) with `init_strategy: random` | Required for adapter-based streaming |

Violating any constraint raises a clear error at config-load time.

### Step 4: Tune Streaming Source and Buffering

```yaml
training:
  stream_source: auto      # auto | ram | disk

  stream_buffers: 2        # double-buffering for I/O overlap

```

- `auto` selects RAM when weights fit, otherwise falls back to NVMe overflow tier
- `stream_buffers` controls prefetching for compute/transfer overlap

### Step 5: Run Training

```bash
soup train --config soup.yaml

```

No extra CLI flag is needed—the streaming behavior triggers purely from configuration.

## Complete Configuration Examples

### DPO with Layer Streaming (Batch Size 1)

```yaml

# soup.yaml – DPO with layer streaming

base: meta-llama/Llama-3.1-8B-Instruct
task: dpo
backend: transformers
modality: text

data:
  train: ./data/train.jsonl
  format: alpaca
  max_length: 512
  val_split: 0.1

training:
  epochs: 3
  lr: 2e-5
  batch_size: 1
  lora:
    r: 64
    alpha: 16
    init_strategy: random
  quantization: 4bit
  stream_layers: true
  stream_source: auto
  stream_buffers: 2

```

### KTO with Larger Batch (Requires ≥2)

```yaml
training:
  epochs: 3
  lr: 2e-5
  batch_size: 2          # KTO minimum batch size

  lora:
    r: 64
    alpha: 16
    init_strategy: random
  quantization: 4bit
  stream_layers: true
  stream_source: auto
  stream_buffers: 2

```

## Memory and Performance Characteristics

### VRAM Preflight Check

The **pre-flight** check in the streaming planner predicts VRAM usage and refuses configurations that would exceed budget. For preference losses, the estimate is **conservative** because the loss packs both chosen and rejected tokens—treating this as an upper bound【source】.

### Time vs. Memory Trade-off

| Aspect | Behavior |
|--------|----------|
| Time penalty | ~1.5× layer reads for DPO vs. SFT (extra comparison pass) |
| VRAM overhead | **None**—primary benefit on low-memory GPUs |
| Reference model | Re-used streamed base, no separate copy |

### Optional Tuning Parameters

```yaml
training:
  # stream_vram_probe: true    # Measure actual VRAM instead of formula estimate

  # stream_disk_kind: nvme     # Force NVMe tier if auto-detection fails

```

## Key Source Files and Implementation

| File | Responsibility |
|------|----------------|
| [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) | `stream_layers` definition and validation schema【source】 |
| [`src/soup_cli/utils/layer_stream.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/layer_stream.py) | Planner building the buffered streaming pipeline【source】 |
| [`src/soup_cli/utils/layer_stream_runtime.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/layer_stream_runtime.py) | Runtime copying layers from CPU RAM/disk to VRAM【source】 |
| [`src/soup_cli/trainer/dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/dpo.py) | DPO trainer respecting `stream_layers` flag【source】 |
| [`docs/performance-and-quantization.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/performance-and-quantization.md) | Performance numbers and compatibility matrix【source】 |

## Summary

- **Layer streaming for DPO/KTO** activates with `stream_layers: true` in Soup v0.72.4+
- **GRPO is unsupported** due to per-token generation requirements that defeat streaming amortization
- **No separate reference model**—Soup re-uses the streamed base with LoRA disabled, saving ~730 MB
- **Strict constraints apply**: transformers backend, text modality, 4bit or no quantization, specific batch sizes, and LoRA with random initialization
- **Performance trade-off**: ~1.5× time penalty for DPO vs. SFT, but zero VRAM overhead for the reference mechanism

## Frequently Asked Questions

### Does layer streaming work with GRPO in Soup?

No. GRPO and PPO require per-token generation with repeated model reads for roll-outs, which would nullify the layer-by-layer streaming advantage. GRPO is explicitly excluded from the streaming roadmap【source】.

### How much VRAM does streaming DPO save compared to standard DPO?

Approximately **730 MB** is saved by eliminating the separate reference model. Standard DPO loads both the policy and reference models; Soup's streaming implementation re-uses the same frozen base with LoRA adapters disabled, so only one set of base weights streams through memory【source】.

### Why does KTO require batch size ≥2 when DPO uses batch size 1?

KTO's loss formulation mathematically requires paired or grouped examples to estimate the reference point for the KL divergence term. The config validator `_validate_stream_layers_compat` enforces `batch_size ≥ 2` for `task: kto` specifically【source】.

### Can I use 8-bit quantization with layer streaming?

No. The validator restricts quantization to `none` or `4bit` (NF4). NF4 is explicitly recommended for streaming because its dequantization overhead is amortized across the layer computation, whereas 8-bit formats introduce memory alignment constraints that complicate the streaming buffer management【source】.