# `stream_source: auto` vs `RAM` vs `disk` in Soup: Storage Mode Explained

> Understand stream_source: auto, RAM, and disk in Soup. Learn how Soup materializes datasets for faster training or greater scalability based on available memory.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: deep-dive
- Published: 2026-08-16

---

**`stream_source` controls where Soup materializes streamed datasets during training: `ram` keeps data in memory for speed, `disk` spills to filesystem for scalability, and `auto` automatically chooses based on available free RAM.**

The `stream_source` configuration parameter in the **Soup** training framework determines how streamed datasets are buffered while a training run is active. This setting directly impacts both training speed and memory usage, making it a critical tuning knob for large-scale fine-tuning jobs. Based on the source implementation in [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py), Soup supports three explicit storage modes with clear behavioral differences.

## What Each `stream_source` Mode Does

Soup provides three mutually exclusive options for the `stream_source` parameter, defined in the configuration schema as `Literal["ram", "disk", "auto"]`:

### RAM Mode (`stream_source: ram`)

**`ram`** forces all streamed data into **in-process host memory**. This provides the lowest possible latency when the training loop requests batches, since no disk I/O occurs.

- Data is buffered in a memory-mapped structure accessible to the GPU data loader
- No temporary files are created on the host filesystem
- Risk of **OOM (out-of-memory) errors** if the streamed dataset exceeds available RAM

Use this mode when your dataset is small enough to fit comfortably in memory and you need maximum throughput.

### Disk Mode (`stream_source: disk`)

**`disk`** writes all streamed data to a **temporary cache directory** on the host filesystem (default: `~/.soup/stream-cache/`). This enables training on arbitrarily large datasets regardless of RAM constraints.

- Data is written sequentially during the streaming phase
- Batches are read back from disk on-demand during training
- Incurs I/O latency but prevents memory exhaustion

Use this mode when working with large datasets on memory-constrained machines, or when you want deterministic, predictable memory usage.

### Auto Mode (`stream_source: auto`)

**`auto`** delegates the decision to Soup's heuristic detection logic. This is the **default behavior** and balances performance with safety.

The selection logic in `choose_stream_source()` at `src/soup_cli/trainer/stream_setup.py:45-78` works as follows:

1. Calls `detect_stream_source()` (`src/soup_cli/trainer/stream_setup.py:80-106`)
2. Reads free memory via `/proc/meminfo` (Linux) or `psutil.virtual_memory()` (cross-platform)
3. Compares against `STREAM_AUTO_RAM_LIMIT` (`src/soup_cli/trainer/stream_setup.py:12-14`)
4. Returns `"ram"` if sufficient memory exists; otherwise returns `"disk"`

The default threshold is approximately **< 4 GiB of free RAM**, though this constant may be tuned per deployment.

## How Auto Mode Decides: Code Walkthrough

The detection logic lives in two key functions within [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py):

```python

# From src/soup_cli/trainer/stream_setup.py

STREAM_AUTO_RAM_LIMIT = 4 * 1024 * 1024 * 1024  # 4 GiB default threshold

def choose_stream_source(config: TrainerConfig) -> str:
    """
    Selects the appropriate stream source based on configuration.
    """
    if config.stream_source != "auto":
        return config.stream_source  # Explicit ram or disk

    
    return detect_stream_source()

def detect_stream_source() -> str:
    """
    Detects system memory status and selects ram vs disk.
    """
    free_bytes = get_free_memory()  # Platform-specific implementation

    
    if free_bytes > STREAM_AUTO_RAM_LIMIT:
        return "ram"
    else:
        return "disk"

```

This implementation ensures that `auto` silently adapts to runtime conditions without user intervention.

## Practical Configuration Examples

```python

# Example 1: Let Soup decide automatically (recommended default)

trainer = soup.trainer(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    data="my_dataset",
    stream_source="auto",
)

# Example 2: Force RAM for small, high-speed workloads

trainer = soup.trainer(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    data="small_dataset",
    stream_source="ram",
)

# Example 3: Force disk for large-scale or memory-constrained training

trainer = soup.trainer(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    data="massive_web_corpus",
    stream_source="disk",
)

```

## Performance and Capacity Trade-offs

| Mode | Speed | Scalability | Memory Predictability | Use Case |
|------|-------|-------------|----------------------|----------|
| **ram** | Fastest | Limited by RAM | Low (dataset size dependent) | Small datasets, latency-sensitive training |
| **disk** | Slower (I/O bound) | Unlimited | High (fixed buffer) | Large datasets, memory-constrained systems |
| **auto** | Adaptive | Automatic fallback | Medium (heuristic-based) | General workloads, unknown dataset sizes |

## Key Source Files

| File | Purpose |
|------|---------|
| [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py) | Core detection logic: `choose_stream_source()`, `detect_stream_source()`, and `STREAM_AUTO_RAM_LIMIT` |
| [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) | Pydantic schema defining `stream_source: Literal["ram", "disk", "auto"]` |
| [`docs/backends-and-ops.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/backends-and-ops.md) | High-level documentation of streaming backends |

## Summary

- **`stream_source: ram`** buffers streamed data entirely in memory for maximum speed, but risks OOM on large datasets
- **`stream_source: disk`** spills data to temporary filesystem storage, enabling unlimited dataset sizes at the cost of I/O latency
- **`stream_source: auto`** automatically selects `ram` when free memory exceeds `STREAM_AUTO_RAM_LIMIT` (default ~4 GiB), otherwise falls back to `disk`
- The selection heuristic is implemented in [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py) and runs transparently at trainer initialization

## Frequently Asked Questions

### What happens if I set `stream_source: ram` but the dataset is too large?

Soup will attempt to allocate the full buffer in memory and likely raise a **MemoryError** or trigger the operating system's OOM killer. No automatic fallback occurs when `ram` is explicitly requested. Monitor your dataset size relative to available RAM, or use `auto` for safety.

### Where does `stream_source: disk` write its temporary files?

By default, disk-cached streams are written to `~/.soup/stream-cache/` with unique subdirectories per training run. This path is configurable through the `SOUP_CACHE_DIR` environment variable or the `cache_dir` parameter in the trainer configuration.

### Can I change the 4 GiB threshold that `auto` mode uses?

Yes. The threshold is defined as `STREAM_AUTO_RAM_LIMIT` in [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py). Advanced users can patch this constant or set the `SOUP_STREAM_AUTO_RAM_LIMIT` environment variable (bytes as integer) to override the default without modifying source code.