`stream_source: auto` vs `RAM` vs `disk` in Soup: Storage Mode Explained

stream_source controls where Soup materializes streamed datasets during training: ram keeps data in memory for speed, disk spills to filesystem for scalability, and auto automatically chooses based on available free RAM.

The stream_source configuration parameter in the Soup training framework determines how streamed datasets are buffered while a training run is active. This setting directly impacts both training speed and memory usage, making it a critical tuning knob for large-scale fine-tuning jobs. Based on the source implementation in src/soup_cli/trainer/stream_setup.py, Soup supports three explicit storage modes with clear behavioral differences.

What Each stream_source Mode Does

Soup provides three mutually exclusive options for the stream_source parameter, defined in the configuration schema as Literal["ram", "disk", "auto"]:

RAM Mode (stream_source: ram)

ram forces all streamed data into in-process host memory. This provides the lowest possible latency when the training loop requests batches, since no disk I/O occurs.

  • Data is buffered in a memory-mapped structure accessible to the GPU data loader
  • No temporary files are created on the host filesystem
  • Risk of OOM (out-of-memory) errors if the streamed dataset exceeds available RAM

Use this mode when your dataset is small enough to fit comfortably in memory and you need maximum throughput.

Disk Mode (stream_source: disk)

disk writes all streamed data to a temporary cache directory on the host filesystem (default: ~/.soup/stream-cache/). This enables training on arbitrarily large datasets regardless of RAM constraints.

  • Data is written sequentially during the streaming phase
  • Batches are read back from disk on-demand during training
  • Incurs I/O latency but prevents memory exhaustion

Use this mode when working with large datasets on memory-constrained machines, or when you want deterministic, predictable memory usage.

Auto Mode (stream_source: auto)

auto delegates the decision to Soup's heuristic detection logic. This is the default behavior and balances performance with safety.

The selection logic in choose_stream_source() at src/soup_cli/trainer/stream_setup.py:45-78 works as follows:

  1. Calls detect_stream_source() (src/soup_cli/trainer/stream_setup.py:80-106)
  2. Reads free memory via /proc/meminfo (Linux) or psutil.virtual_memory() (cross-platform)
  3. Compares against STREAM_AUTO_RAM_LIMIT (src/soup_cli/trainer/stream_setup.py:12-14)
  4. Returns "ram" if sufficient memory exists; otherwise returns "disk"

The default threshold is approximately < 4 GiB of free RAM, though this constant may be tuned per deployment.

How Auto Mode Decides: Code Walkthrough

The detection logic lives in two key functions within src/soup_cli/trainer/stream_setup.py:


# From src/soup_cli/trainer/stream_setup.py

STREAM_AUTO_RAM_LIMIT = 4 * 1024 * 1024 * 1024  # 4 GiB default threshold

def choose_stream_source(config: TrainerConfig) -> str:
    """
    Selects the appropriate stream source based on configuration.
    """
    if config.stream_source != "auto":
        return config.stream_source  # Explicit ram or disk

    
    return detect_stream_source()

def detect_stream_source() -> str:
    """
    Detects system memory status and selects ram vs disk.
    """
    free_bytes = get_free_memory()  # Platform-specific implementation

    
    if free_bytes > STREAM_AUTO_RAM_LIMIT:
        return "ram"
    else:
        return "disk"

This implementation ensures that auto silently adapts to runtime conditions without user intervention.

Practical Configuration Examples


# Example 1: Let Soup decide automatically (recommended default)

trainer = soup.trainer(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    data="my_dataset",
    stream_source="auto",
)

# Example 2: Force RAM for small, high-speed workloads

trainer = soup.trainer(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    data="small_dataset",
    stream_source="ram",
)

# Example 3: Force disk for large-scale or memory-constrained training

trainer = soup.trainer(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    data="massive_web_corpus",
    stream_source="disk",
)

Performance and Capacity Trade-offs

Mode Speed Scalability Memory Predictability Use Case
ram Fastest Limited by RAM Low (dataset size dependent) Small datasets, latency-sensitive training
disk Slower (I/O bound) Unlimited High (fixed buffer) Large datasets, memory-constrained systems
auto Adaptive Automatic fallback Medium (heuristic-based) General workloads, unknown dataset sizes

Key Source Files

File Purpose
src/soup_cli/trainer/stream_setup.py Core detection logic: choose_stream_source(), detect_stream_source(), and STREAM_AUTO_RAM_LIMIT
src/soup_cli/config/schema.py Pydantic schema defining stream_source: Literal["ram", "disk", "auto"]
docs/backends-and-ops.md High-level documentation of streaming backends

Summary

  • stream_source: ram buffers streamed data entirely in memory for maximum speed, but risks OOM on large datasets
  • stream_source: disk spills data to temporary filesystem storage, enabling unlimited dataset sizes at the cost of I/O latency
  • stream_source: auto automatically selects ram when free memory exceeds STREAM_AUTO_RAM_LIMIT (default ~4 GiB), otherwise falls back to disk
  • The selection heuristic is implemented in src/soup_cli/trainer/stream_setup.py and runs transparently at trainer initialization

Frequently Asked Questions

What happens if I set stream_source: ram but the dataset is too large?

Soup will attempt to allocate the full buffer in memory and likely raise a MemoryError or trigger the operating system's OOM killer. No automatic fallback occurs when ram is explicitly requested. Monitor your dataset size relative to available RAM, or use auto for safety.

Where does stream_source: disk write its temporary files?

By default, disk-cached streams are written to ~/.soup/stream-cache/ with unique subdirectories per training run. This path is configurable through the SOUP_CACHE_DIR environment variable or the cache_dir parameter in the trainer configuration.

Can I change the 4 GiB threshold that auto mode uses?

Yes. The threshold is defined as STREAM_AUTO_RAM_LIMIT in src/soup_cli/trainer/stream_setup.py. Advanced users can patch this constant or set the SOUP_STREAM_AUTO_RAM_LIMIT environment variable (bytes as integer) to override the default without modifying source code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →