# How to Configure Stream Layers for Optimal VRAM Usage in Soup

> Optimize VRAM usage in Soup by configuring stream layers. Learn how to enable layer streaming to load transformer weights on demand and save GPU memory.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-08-16

---

**Set `training.stream_layers: true` in your YAML config to enable layer streaming, which loads transformer-layer weights from disk on-demand instead of keeping the full model resident in GPU memory.**

Configuring **stream layers** for optimal VRAM usage is essential when training large transformer models with limited GPU memory. Soup v0.72+ introduces a **layer streaming mechanism** that dramatically reduces peak VRAM consumption by swapping weights between disk and GPU during the forward pass. This guide explains how to configure stream layers based on the actual implementation in the [MakazhanAlpamys/Soup](https://github.com/MakazhanAlpamys/Soup) repository.

## How Stream Layers Work

### The Core Streaming Runtime

The `StreamRuntime` class in [`src/soup_cli/utils/layer_stream.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/layer_stream.py) implements the weight-streaming mechanism. It creates a temporary buffer on the device and lazily swaps in weight tensors for each layer as the forward pass reaches it.

- Only the currently-active layer (or a small prefetch buffer) lives in VRAM at any moment
- This cuts peak memory footprint roughly by the size of one-to-several transformer layers
- Training and evaluation semantics remain identical to non-streamed runs

### VRAM Estimation Before Training

The `estimate_stream_peak_vram()` function analyzes model topology and batch size to predict maximum VRAM consumption when streaming is active.

```python
from soup_cli.utils.layer_stream import estimate_stream_peak_vram

peak = estimate_stream_peak_vram(cfg)
print(f"Estimated peak VRAM with streaming: {peak / (1024**3):.2f} GB")

```

*Source: [`src/soup_cli/utils/layer_stream.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/layer_stream.py), line 215*

### Trainer Integration

The `StreamingSetupMixin` in [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py) injects the streaming runtime into the trainer's lifecycle—handling initialization, training loops, and finalization automatically.

## Configuration Schema for Stream Layers

The Pydantic configuration schema defines the streaming options in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) (line 78):

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `training.stream_layers` | `bool` | `false` | Master switch to enable layer streaming |
| `training.streaming_buffer_size` | `int` | `1` | Number of layers to prefetch for I/O smoothing |

Set `stream_layers: true` to activate the streaming path. When `false`, the model uses standard non-streamed behavior.

## Recommended Settings for Optimal VRAM

### Essential Configuration

- **`training.stream_layers`** — Set to `true` (or omit to let the CLI auto-detect). This activates the streaming path and is the single most effective VRAM-saving setting.

- **`training.gradient_checkpointing`** — Set to `false` or leave at default. When streaming is enabled, Hugging Face gradient checkpointing is automatically disabled per the test suite in [`tests/test_v07204.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_v07204.py). Both mechanisms trade computation for memory; combining them causes unnecessary overhead.

- **`training.batch_size`** — Keep as high as your **non-streamed** VRAM allows. Streaming reduces weight memory, not activation memory per step.

- **`training.max_seq_len`** — Unchanged from your standard configuration. Sequence length still determines activation memory, which streaming does not affect.

### Advanced Tuning

- **`training.streaming_buffer_size`** — Default `1` layer is optimal for tight VRAM budgets. Raising to `2` can smooth I/O stalls but increases resident memory proportionally. Monitor the runtime log and adjust if estimates appear optimistic.

- **Mixed precision** — Combine streaming with `torch.float16` or `bfloat16`. Streaming saves weight memory; activations still benefit from reduced precision.

## Practical Configuration Examples

### YAML Configuration

```yaml

# my_config.yaml

training:
  device: cuda
  batch_size: 32
  max_seq_len: 2048
  stream_layers: true          # Enable layer streaming

  streaming_buffer_size: 1     # Prefetch 1 layer (default)

  gradient_checkpointing: false  # Automatically disabled when streaming

```

### CLI Invocation

```bash
soup train --config my_config.yaml

```

### Python API

```python
from soup_cli.config.schema import Config

cfg = Config.load("my_config.yaml")
cfg.training.stream_layers = True
cfg.training.streaming_buffer_size = 1
cfg.validate()  # Sanity-check configuration

```

## When to Enable Stream Layers

Enable streaming only when you encounter VRAM limitations. If a run fits comfortably without streaming, disabling it removes minor I/O overhead.

Monitor the runtime log for estimated peak VRAM and buffer size reports from `StreamingSetupMixin`. Adjust `streaming_buffer_size` if actual consumption exceeds estimates.

For NF4-quantized models, additional streamed weight handling is validated in [`tests/test_v07300.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_v07300.py).

## Summary

- **Enable streaming** with `training.stream_layers: true` to reduce peak VRAM by layer weight size
- **Keep buffer size low** (`1` layer) for maximum memory savings; increase to `2` only if I/O bottlenecks appear
- **Disable gradient checkpointing** when streaming—Soup handles this automatically
- **Estimate VRAM upfront** using `estimate_stream_peak_vram()` to validate configurations
- **Preserve batch size and sequence length** settings; streaming does not affect activation memory

## Frequently Asked Questions

### Does stream_layers affect training speed?

Layer streaming introduces modest I/O overhead from disk-to-GPU weight transfers. The `streaming_buffer_size` setting mitigates this by prefetching layers. For NVMe storage with good sequential read performance, overhead is typically 5-15%. Slower storage or very large models may see more significant impacts.

### Can I use stream_layers with CPU training?

Yes. The `StreamRuntime` works on both CUDA and CPU devices. On CUDA, the runtime prefers `float32` for intermediate buffers. CPU streaming benefits from reduced system RAM pressure rather than VRAM constraints.

### How does streaming interact with model quantization?

Stream layers operate independently of weight quantization. The [`tests/test_v07300.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_v07300.py) suite validates streamed NF4 weight handling for large quantized models. Quantization reduces per-layer size, allowing smaller `streaming_buffer_size` values or enabling streaming on hardware that couldn't otherwise fit the model.

### Why is gradient_checkpointing disabled when streaming is enabled?

Both mechanisms trade computation for memory. Gradient checkpointing saves activation memory by recomputing forward passes during backpropagation. Layer streaming saves weight memory by loading layers on-demand. Enabling both creates redundant computation without additive memory benefits, so Soup automatically disables HF gradient checkpointing when `stream_layers: true`.