# How to Configure SSD Streaming to Run Models Larger Than Available RAM in ds4

> Learn how to configure SSD streaming in ds4 to run models larger than available RAM. ds4 streams model weights from SSD, optimizing performance for Metal CUDA and ROCm.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-07

---

**ds4 streams model weights from SSD instead of loading the entire GGUF into RAM, using a dynamic expert cache to keep frequently used layers in memory while routing active experts from disk on Metal, CUDA, and ROCm backends.**

Running large language models often exceeds available system memory, especially with modern expert architectures. The `antirez/ds4` inference engine solves this through **SSD streaming**, a capacity mode that lets you configure SSD streaming to run models larger than available RAM by maintaining only a working set of expert layers in memory while fetching the rest from high-speed storage.

## Understanding SSD Streaming Capacity Mode

SSD streaming in ds4 works by routing only the expert layers needed for the current token from your SSD, rather than keeping the complete model resident in RAM. The system maintains a **dynamic expert cache** in memory for the most frequently accessed experts, ensuring performance remains acceptable while dramatically reducing memory footprint.

This mode is supported across Metal, CUDA, and ROCm backends. The implementation handles the complexity of expert routing and cache management automatically once configured.

## Cache Configuration Options

The dynamic expert cache size can be controlled through three distinct methods, parsed in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) by the function `ds4_parse_streaming_cache_experts_arg` (see lines 46-70).

### Automatic Cache Sizing

By default, ds4 automatically determines the cache size based on the backend's recommended working-set size. The environment variable `DS4_SSD_AUTO_CACHE_PCT` controls the percentage of this working-set to use, defaulting to 80%.

The logic resides in `ds4_ssd_auto_cache_plan` within [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c). If you do not specify `--ssd-streaming-cache-experts`, this automatic budgeting applies.

### Manual Byte Budget

Specify a concrete memory budget in gigabytes to limit the cache size. This converts your byte budget to an expert count using `ds4_ssd_cache_experts_for_byte_budget`.

```bash
--ssd-streaming-cache-experts 32GB

```

This reserves exactly 32 GiB of RAM for the dynamic expert cache.

### Manual Expert Count

Alternatively, specify an exact number of experts to cache:

```bash
--ssd-streaming-cache-experts 4000

```

This bypasses automatic calculations and maintains exactly 4000 experts in memory.

## Controlling Expert Preloading

By default, ds4 preloads "hot" experts to improve initial performance. You can modify this behavior using flags defined in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c).

### Disable Hot-Expert Preload

Use `--ssd-streaming-cold` to skip the default hot-expert preload. This is useful for benchmarking or when you want to minimize startup memory usage.

```bash
--ssd-streaming-cold

```

### Adjust Preload Count

Control exactly how many experts are preloaded during initialization:

```bash
--ssd-streaming-preload-experts 200

```

## Step-by-Step Configuration Guide

To configure SSD streaming to run models larger than available RAM, follow these steps:

1. **Enable streaming** – Add the `--ssd-streaming` flag (requires Metal, CUDA, or ROCm).
2. **Set cache policy** – Either rely on automatic sizing, or specify a byte budget or expert count using `--ssd-streaming-cache-experts`.
3. **Configure preloading** – Optionally add `--ssd-streaming-cold` to disable preloading, or set `--ssd-streaming-preload-experts N` for a custom preload count.

The CLI option definitions are located in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) around lines 170-174, while the parsing logic and cache planning live in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c).

## Practical Configuration Examples

These examples demonstrate real-world configurations for different hardware constraints and use cases.

### Basic Streaming with Automatic Cache

Simplest configuration using automatic 80% working-set cache:

```bash
./ds4 -m ./ds4flash.gguf --ssd-streaming

```

### Manual Byte-Budget Cache

Specify 32 GiB for routed-expert memory:

```bash
./ds4 -m ./ds4flash.gguf \
  --ssd-streaming \
  --ssd-streaming-cache-experts 32GB

```

### Manual Expert-Count Cache

Cache exactly 4000 dynamic experts:

```bash
./ds4 -m ./ds4flash.gguf \
  --ssd-streaming \
  --ssd-streaming-cache-experts 4000

```

### Disable Hot-Expert Preload

Useful for benchmarking cold-start behavior:

```bash
./ds4 -m ./ds4flash.gguf \
  --ssd-streaming \
  --ssd-streaming-cold

```

### Adjust Preload Count

Preload exactly 200 experts at startup:

```bash
./ds4 -m ./ds4flash.gguf \
  --ssd-streaming \
  --ssd-streaming-preload-experts 200

```

### MacBook with 64 GB RAM (Q2 Flash Model)

For a moderate cache with a 64 GB system:

```bash
./download_model.sh ds4f-q2
./ds4 -m ./ds4flash.gguf \
  --ssd-streaming \
  --ssd-streaming-cache-experts 32GB \
  --ctx 32768 \
  --nothink

```

### MacBook with 128 GB RAM (DeepSeek-V4-Pro)

For larger models on high-memory laptops:

```bash
./download_model.sh pro-q2-imatrix
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
  --ssd-streaming \
  --ctx 32768 \
  --nothink

```

### Strix Halo with ROCm (128 GB)

For AMD ROCm systems with routed Q2_K models:

```bash
./download_model.sh glm-antirez-q2
make strix-halo
./ds4 --rocm -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
  --ssd-streaming --ctx 4096

```

## Key Source Files and Implementation

Understanding the source structure helps when debugging or extending functionality:

| File | Role |
|------|------|
| [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) | Parses cache arguments (`ds4_parse_streaming_cache_experts_arg`), computes automatic cache plans (`ds4_ssd_auto_cache_plan`), and converts byte budgets to expert counts (`ds4_ssd_cache_experts_for_byte_budget`). |
| [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) | Defines CLI flags `--ssd-streaming`, `--ssd-streaming-cache-experts`, `--ssd-streaming-cold`, and `--ssd-streaming-preload-experts` around lines 170-174. |
| [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) | Contains the core inference engine with the SSD-streaming decode path and backend-specific error handling (e.g., "Metal SSD streaming could not build …"). |
| [`tests/ds4_test.c`](https://github.com/antirez/ds4/blob/main/tests/ds4_test.c) | Test harness supporting SSD streaming via the `DS4_TEST_SSD_STREAMING` environment variable. |

## Summary

- **SSD streaming** routes only active expert layers from disk, keeping a dynamic cache in RAM to run models exceeding physical memory.
- **Three cache modes** are available: automatic (default 80% of working-set via `DS4_SSD_AUTO_CACHE_PCT`), manual byte budget (e.g., `32GB`), and manual expert count (e.g., `4000`).
- **Preload control** via `--ssd-streaming-cold` (disable) or `--ssd-streaming-preload-experts N` (custom count).
- **Backend support** requires Metal, CUDA, or ROCm.
- **Core implementation** resides in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) and [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c), with runtime logic in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c).

## Frequently Asked Questions

### What hardware backends support SSD streaming in ds4?

SSD streaming requires Metal, CUDA, or ROCm backends. The feature is not available on CPU-only builds. According to the [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) source, the engine validates backend compatibility before initializing the SSD streaming path.

### How is the automatic cache size calculated?

The automatic mode uses `ds4_ssd_auto_cache_plan` in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) to query the backend's recommended working-set size, then applies the percentage specified by `DS4_SSD_AUTO_CACHE_PCT` (defaulting to 80%). This determines how many experts fit in the calculated memory budget without manual intervention.

### Can I run a model that is twice the size of my RAM?

Yes. SSD streaming is specifically designed for this scenario. As long as your SSD has sufficient space for the GGUF file and you configure an appropriate cache size using `--ssd-streaming-cache-experts`, ds4 will route experts from disk while keeping only the working set in memory.

### What is the performance impact of using --ssd-streaming-cold?

Disabling hot-expert preload with `--ssd-streaming-cold` increases initial latency for the first few tokens because experts must be loaded from SSD on first access rather than being preloaded into cache. This is primarily useful for benchmarking cold-start behavior or conserving memory during initialization, as implemented in the startup sequence of [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c).