How to Configure Stream Layers for Optimal VRAM Usage in Soup
Set training.stream_layers: true in your YAML config to enable layer streaming, which loads transformer-layer weights from disk on-demand instead of keeping the full model resident in GPU memory.
Configuring stream layers for optimal VRAM usage is essential when training large transformer models with limited GPU memory. Soup v0.72+ introduces a layer streaming mechanism that dramatically reduces peak VRAM consumption by swapping weights between disk and GPU during the forward pass. This guide explains how to configure stream layers based on the actual implementation in the MakazhanAlpamys/Soup repository.
How Stream Layers Work
The Core Streaming Runtime
The StreamRuntime class in src/soup_cli/utils/layer_stream.py implements the weight-streaming mechanism. It creates a temporary buffer on the device and lazily swaps in weight tensors for each layer as the forward pass reaches it.
- Only the currently-active layer (or a small prefetch buffer) lives in VRAM at any moment
- This cuts peak memory footprint roughly by the size of one-to-several transformer layers
- Training and evaluation semantics remain identical to non-streamed runs
VRAM Estimation Before Training
The estimate_stream_peak_vram() function analyzes model topology and batch size to predict maximum VRAM consumption when streaming is active.
from soup_cli.utils.layer_stream import estimate_stream_peak_vram
peak = estimate_stream_peak_vram(cfg)
print(f"Estimated peak VRAM with streaming: {peak / (1024**3):.2f} GB")
Source: src/soup_cli/utils/layer_stream.py, line 215
Trainer Integration
The StreamingSetupMixin in src/soup_cli/trainer/stream_setup.py injects the streaming runtime into the trainer's lifecycle—handling initialization, training loops, and finalization automatically.
Configuration Schema for Stream Layers
The Pydantic configuration schema defines the streaming options in src/soup_cli/config/schema.py (line 78):
| Parameter | Type | Default | Description |
|---|---|---|---|
training.stream_layers |
bool |
false |
Master switch to enable layer streaming |
training.streaming_buffer_size |
int |
1 |
Number of layers to prefetch for I/O smoothing |
Set stream_layers: true to activate the streaming path. When false, the model uses standard non-streamed behavior.
Recommended Settings for Optimal VRAM
Essential Configuration
-
training.stream_layers— Set totrue(or omit to let the CLI auto-detect). This activates the streaming path and is the single most effective VRAM-saving setting. -
training.gradient_checkpointing— Set tofalseor leave at default. When streaming is enabled, Hugging Face gradient checkpointing is automatically disabled per the test suite intests/test_v07204.py. Both mechanisms trade computation for memory; combining them causes unnecessary overhead. -
training.batch_size— Keep as high as your non-streamed VRAM allows. Streaming reduces weight memory, not activation memory per step. -
training.max_seq_len— Unchanged from your standard configuration. Sequence length still determines activation memory, which streaming does not affect.
Advanced Tuning
-
training.streaming_buffer_size— Default1layer is optimal for tight VRAM budgets. Raising to2can smooth I/O stalls but increases resident memory proportionally. Monitor the runtime log and adjust if estimates appear optimistic. -
Mixed precision — Combine streaming with
torch.float16orbfloat16. Streaming saves weight memory; activations still benefit from reduced precision.
Practical Configuration Examples
YAML Configuration
# my_config.yaml
training:
device: cuda
batch_size: 32
max_seq_len: 2048
stream_layers: true # Enable layer streaming
streaming_buffer_size: 1 # Prefetch 1 layer (default)
gradient_checkpointing: false # Automatically disabled when streaming
CLI Invocation
soup train --config my_config.yaml
Python API
from soup_cli.config.schema import Config
cfg = Config.load("my_config.yaml")
cfg.training.stream_layers = True
cfg.training.streaming_buffer_size = 1
cfg.validate() # Sanity-check configuration
When to Enable Stream Layers
Enable streaming only when you encounter VRAM limitations. If a run fits comfortably without streaming, disabling it removes minor I/O overhead.
Monitor the runtime log for estimated peak VRAM and buffer size reports from StreamingSetupMixin. Adjust streaming_buffer_size if actual consumption exceeds estimates.
For NF4-quantized models, additional streamed weight handling is validated in tests/test_v07300.py.
Summary
- Enable streaming with
training.stream_layers: trueto reduce peak VRAM by layer weight size - Keep buffer size low (
1layer) for maximum memory savings; increase to2only if I/O bottlenecks appear - Disable gradient checkpointing when streaming—Soup handles this automatically
- Estimate VRAM upfront using
estimate_stream_peak_vram()to validate configurations - Preserve batch size and sequence length settings; streaming does not affect activation memory
Frequently Asked Questions
Does stream_layers affect training speed?
Layer streaming introduces modest I/O overhead from disk-to-GPU weight transfers. The streaming_buffer_size setting mitigates this by prefetching layers. For NVMe storage with good sequential read performance, overhead is typically 5-15%. Slower storage or very large models may see more significant impacts.
Can I use stream_layers with CPU training?
Yes. The StreamRuntime works on both CUDA and CPU devices. On CUDA, the runtime prefers float32 for intermediate buffers. CPU streaming benefits from reduced system RAM pressure rather than VRAM constraints.
How does streaming interact with model quantization?
Stream layers operate independently of weight quantization. The tests/test_v07300.py suite validates streamed NF4 weight handling for large quantized models. Quantization reduces per-layer size, allowing smaller streaming_buffer_size values or enabling streaming on hardware that couldn't otherwise fit the model.
Why is gradient_checkpointing disabled when streaming is enabled?
Both mechanisms trade computation for memory. Gradient checkpointing saves activation memory by recomputing forward passes during backpropagation. Layer streaming saves weight memory by loading layers on-demand. Enabling both creates redundant computation without additive memory benefits, so Soup automatically disables HF gradient checkpointing when stream_layers: true.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →