How to Configure Stream Layers for Optimal VRAM Usage in Soup

Set training.stream_layers: true in your YAML config to enable layer streaming, which loads transformer-layer weights from disk on-demand instead of keeping the full model resident in GPU memory.

Configuring stream layers for optimal VRAM usage is essential when training large transformer models with limited GPU memory. Soup v0.72+ introduces a layer streaming mechanism that dramatically reduces peak VRAM consumption by swapping weights between disk and GPU during the forward pass. This guide explains how to configure stream layers based on the actual implementation in the MakazhanAlpamys/Soup repository.

How Stream Layers Work

The Core Streaming Runtime

The StreamRuntime class in src/soup_cli/utils/layer_stream.py implements the weight-streaming mechanism. It creates a temporary buffer on the device and lazily swaps in weight tensors for each layer as the forward pass reaches it.

  • Only the currently-active layer (or a small prefetch buffer) lives in VRAM at any moment
  • This cuts peak memory footprint roughly by the size of one-to-several transformer layers
  • Training and evaluation semantics remain identical to non-streamed runs

VRAM Estimation Before Training

The estimate_stream_peak_vram() function analyzes model topology and batch size to predict maximum VRAM consumption when streaming is active.

from soup_cli.utils.layer_stream import estimate_stream_peak_vram

peak = estimate_stream_peak_vram(cfg)
print(f"Estimated peak VRAM with streaming: {peak / (1024**3):.2f} GB")

Source: src/soup_cli/utils/layer_stream.py, line 215

Trainer Integration

The StreamingSetupMixin in src/soup_cli/trainer/stream_setup.py injects the streaming runtime into the trainer's lifecycle—handling initialization, training loops, and finalization automatically.

Configuration Schema for Stream Layers

The Pydantic configuration schema defines the streaming options in src/soup_cli/config/schema.py (line 78):

Parameter Type Default Description
training.stream_layers bool false Master switch to enable layer streaming
training.streaming_buffer_size int 1 Number of layers to prefetch for I/O smoothing

Set stream_layers: true to activate the streaming path. When false, the model uses standard non-streamed behavior.

Essential Configuration

  • training.stream_layers — Set to true (or omit to let the CLI auto-detect). This activates the streaming path and is the single most effective VRAM-saving setting.

  • training.gradient_checkpointing — Set to false or leave at default. When streaming is enabled, Hugging Face gradient checkpointing is automatically disabled per the test suite in tests/test_v07204.py. Both mechanisms trade computation for memory; combining them causes unnecessary overhead.

  • training.batch_size — Keep as high as your non-streamed VRAM allows. Streaming reduces weight memory, not activation memory per step.

  • training.max_seq_len — Unchanged from your standard configuration. Sequence length still determines activation memory, which streaming does not affect.

Advanced Tuning

  • training.streaming_buffer_size — Default 1 layer is optimal for tight VRAM budgets. Raising to 2 can smooth I/O stalls but increases resident memory proportionally. Monitor the runtime log and adjust if estimates appear optimistic.

  • Mixed precision — Combine streaming with torch.float16 or bfloat16. Streaming saves weight memory; activations still benefit from reduced precision.

Practical Configuration Examples

YAML Configuration


# my_config.yaml

training:
  device: cuda
  batch_size: 32
  max_seq_len: 2048
  stream_layers: true          # Enable layer streaming

  streaming_buffer_size: 1     # Prefetch 1 layer (default)

  gradient_checkpointing: false  # Automatically disabled when streaming

CLI Invocation

soup train --config my_config.yaml

Python API

from soup_cli.config.schema import Config

cfg = Config.load("my_config.yaml")
cfg.training.stream_layers = True
cfg.training.streaming_buffer_size = 1
cfg.validate()  # Sanity-check configuration

When to Enable Stream Layers

Enable streaming only when you encounter VRAM limitations. If a run fits comfortably without streaming, disabling it removes minor I/O overhead.

Monitor the runtime log for estimated peak VRAM and buffer size reports from StreamingSetupMixin. Adjust streaming_buffer_size if actual consumption exceeds estimates.

For NF4-quantized models, additional streamed weight handling is validated in tests/test_v07300.py.

Summary

  • Enable streaming with training.stream_layers: true to reduce peak VRAM by layer weight size
  • Keep buffer size low (1 layer) for maximum memory savings; increase to 2 only if I/O bottlenecks appear
  • Disable gradient checkpointing when streaming—Soup handles this automatically
  • Estimate VRAM upfront using estimate_stream_peak_vram() to validate configurations
  • Preserve batch size and sequence length settings; streaming does not affect activation memory

Frequently Asked Questions

Does stream_layers affect training speed?

Layer streaming introduces modest I/O overhead from disk-to-GPU weight transfers. The streaming_buffer_size setting mitigates this by prefetching layers. For NVMe storage with good sequential read performance, overhead is typically 5-15%. Slower storage or very large models may see more significant impacts.

Can I use stream_layers with CPU training?

Yes. The StreamRuntime works on both CUDA and CPU devices. On CUDA, the runtime prefers float32 for intermediate buffers. CPU streaming benefits from reduced system RAM pressure rather than VRAM constraints.

How does streaming interact with model quantization?

Stream layers operate independently of weight quantization. The tests/test_v07300.py suite validates streamed NF4 weight handling for large quantized models. Quantization reduces per-layer size, allowing smaller streaming_buffer_size values or enabling streaming on hardware that couldn't otherwise fit the model.

Why is gradient_checkpointing disabled when streaming is enabled?

Both mechanisms trade computation for memory. Gradient checkpointing saves activation memory by recomputing forward passes during backpropagation. Layer streaming saves weight memory by loading layers on-demand. Enabling both creates redundant computation without additive memory benefits, so Soup automatically disables HF gradient checkpointing when stream_layers: true.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →