`stream_source: auto` vs `RAM` vs `disk` in Soup: Storage Mode Explained
stream_source controls where Soup materializes streamed datasets during training: ram keeps data in memory for speed, disk spills to filesystem for scalability, and auto automatically chooses based on available free RAM.
The stream_source configuration parameter in the Soup training framework determines how streamed datasets are buffered while a training run is active. This setting directly impacts both training speed and memory usage, making it a critical tuning knob for large-scale fine-tuning jobs. Based on the source implementation in src/soup_cli/trainer/stream_setup.py, Soup supports three explicit storage modes with clear behavioral differences.
What Each stream_source Mode Does
Soup provides three mutually exclusive options for the stream_source parameter, defined in the configuration schema as Literal["ram", "disk", "auto"]:
RAM Mode (stream_source: ram)
ram forces all streamed data into in-process host memory. This provides the lowest possible latency when the training loop requests batches, since no disk I/O occurs.
- Data is buffered in a memory-mapped structure accessible to the GPU data loader
- No temporary files are created on the host filesystem
- Risk of OOM (out-of-memory) errors if the streamed dataset exceeds available RAM
Use this mode when your dataset is small enough to fit comfortably in memory and you need maximum throughput.
Disk Mode (stream_source: disk)
disk writes all streamed data to a temporary cache directory on the host filesystem (default: ~/.soup/stream-cache/). This enables training on arbitrarily large datasets regardless of RAM constraints.
- Data is written sequentially during the streaming phase
- Batches are read back from disk on-demand during training
- Incurs I/O latency but prevents memory exhaustion
Use this mode when working with large datasets on memory-constrained machines, or when you want deterministic, predictable memory usage.
Auto Mode (stream_source: auto)
auto delegates the decision to Soup's heuristic detection logic. This is the default behavior and balances performance with safety.
The selection logic in choose_stream_source() at src/soup_cli/trainer/stream_setup.py:45-78 works as follows:
- Calls
detect_stream_source()(src/soup_cli/trainer/stream_setup.py:80-106) - Reads free memory via
/proc/meminfo(Linux) orpsutil.virtual_memory()(cross-platform) - Compares against
STREAM_AUTO_RAM_LIMIT(src/soup_cli/trainer/stream_setup.py:12-14) - Returns
"ram"if sufficient memory exists; otherwise returns"disk"
The default threshold is approximately < 4 GiB of free RAM, though this constant may be tuned per deployment.
How Auto Mode Decides: Code Walkthrough
The detection logic lives in two key functions within src/soup_cli/trainer/stream_setup.py:
# From src/soup_cli/trainer/stream_setup.py
STREAM_AUTO_RAM_LIMIT = 4 * 1024 * 1024 * 1024 # 4 GiB default threshold
def choose_stream_source(config: TrainerConfig) -> str:
"""
Selects the appropriate stream source based on configuration.
"""
if config.stream_source != "auto":
return config.stream_source # Explicit ram or disk
return detect_stream_source()
def detect_stream_source() -> str:
"""
Detects system memory status and selects ram vs disk.
"""
free_bytes = get_free_memory() # Platform-specific implementation
if free_bytes > STREAM_AUTO_RAM_LIMIT:
return "ram"
else:
return "disk"
This implementation ensures that auto silently adapts to runtime conditions without user intervention.
Practical Configuration Examples
# Example 1: Let Soup decide automatically (recommended default)
trainer = soup.trainer(
model="meta-llama/Meta-Llama-3-8B-Instruct",
data="my_dataset",
stream_source="auto",
)
# Example 2: Force RAM for small, high-speed workloads
trainer = soup.trainer(
model="meta-llama/Meta-Llama-3-8B-Instruct",
data="small_dataset",
stream_source="ram",
)
# Example 3: Force disk for large-scale or memory-constrained training
trainer = soup.trainer(
model="meta-llama/Meta-Llama-3-8B-Instruct",
data="massive_web_corpus",
stream_source="disk",
)
Performance and Capacity Trade-offs
| Mode | Speed | Scalability | Memory Predictability | Use Case |
|---|---|---|---|---|
| ram | Fastest | Limited by RAM | Low (dataset size dependent) | Small datasets, latency-sensitive training |
| disk | Slower (I/O bound) | Unlimited | High (fixed buffer) | Large datasets, memory-constrained systems |
| auto | Adaptive | Automatic fallback | Medium (heuristic-based) | General workloads, unknown dataset sizes |
Key Source Files
| File | Purpose |
|---|---|
src/soup_cli/trainer/stream_setup.py |
Core detection logic: choose_stream_source(), detect_stream_source(), and STREAM_AUTO_RAM_LIMIT |
src/soup_cli/config/schema.py |
Pydantic schema defining stream_source: Literal["ram", "disk", "auto"] |
docs/backends-and-ops.md |
High-level documentation of streaming backends |
Summary
stream_source: rambuffers streamed data entirely in memory for maximum speed, but risks OOM on large datasetsstream_source: diskspills data to temporary filesystem storage, enabling unlimited dataset sizes at the cost of I/O latencystream_source: autoautomatically selectsramwhen free memory exceedsSTREAM_AUTO_RAM_LIMIT(default ~4 GiB), otherwise falls back todisk- The selection heuristic is implemented in
src/soup_cli/trainer/stream_setup.pyand runs transparently at trainer initialization
Frequently Asked Questions
What happens if I set stream_source: ram but the dataset is too large?
Soup will attempt to allocate the full buffer in memory and likely raise a MemoryError or trigger the operating system's OOM killer. No automatic fallback occurs when ram is explicitly requested. Monitor your dataset size relative to available RAM, or use auto for safety.
Where does stream_source: disk write its temporary files?
By default, disk-cached streams are written to ~/.soup/stream-cache/ with unique subdirectories per training run. This path is configurable through the SOUP_CACHE_DIR environment variable or the cache_dir parameter in the trainer configuration.
Can I change the 4 GiB threshold that auto mode uses?
Yes. The threshold is defined as STREAM_AUTO_RAM_LIMIT in src/soup_cli/trainer/stream_setup.py. Advanced users can patch this constant or set the SOUP_STREAM_AUTO_RAM_LIMIT environment variable (bytes as integer) to override the default without modifying source code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →