How to Configure SSD Streaming to Run Models Larger Than Available RAM in ds4
ds4 streams model weights from SSD instead of loading the entire GGUF into RAM, using a dynamic expert cache to keep frequently used layers in memory while routing active experts from disk on Metal, CUDA, and ROCm backends.
Running large language models often exceeds available system memory, especially with modern expert architectures. The antirez/ds4 inference engine solves this through SSD streaming, a capacity mode that lets you configure SSD streaming to run models larger than available RAM by maintaining only a working set of expert layers in memory while fetching the rest from high-speed storage.
Understanding SSD Streaming Capacity Mode
SSD streaming in ds4 works by routing only the expert layers needed for the current token from your SSD, rather than keeping the complete model resident in RAM. The system maintains a dynamic expert cache in memory for the most frequently accessed experts, ensuring performance remains acceptable while dramatically reducing memory footprint.
This mode is supported across Metal, CUDA, and ROCm backends. The implementation handles the complexity of expert routing and cache management automatically once configured.
Cache Configuration Options
The dynamic expert cache size can be controlled through three distinct methods, parsed in ds4_ssd.c by the function ds4_parse_streaming_cache_experts_arg (see lines 46-70).
Automatic Cache Sizing
By default, ds4 automatically determines the cache size based on the backend's recommended working-set size. The environment variable DS4_SSD_AUTO_CACHE_PCT controls the percentage of this working-set to use, defaulting to 80%.
The logic resides in ds4_ssd_auto_cache_plan within ds4_ssd.c. If you do not specify --ssd-streaming-cache-experts, this automatic budgeting applies.
Manual Byte Budget
Specify a concrete memory budget in gigabytes to limit the cache size. This converts your byte budget to an expert count using ds4_ssd_cache_experts_for_byte_budget.
--ssd-streaming-cache-experts 32GB
This reserves exactly 32 GiB of RAM for the dynamic expert cache.
Manual Expert Count
Alternatively, specify an exact number of experts to cache:
--ssd-streaming-cache-experts 4000
This bypasses automatic calculations and maintains exactly 4000 experts in memory.
Controlling Expert Preloading
By default, ds4 preloads "hot" experts to improve initial performance. You can modify this behavior using flags defined in ds4_help.c.
Disable Hot-Expert Preload
Use --ssd-streaming-cold to skip the default hot-expert preload. This is useful for benchmarking or when you want to minimize startup memory usage.
--ssd-streaming-cold
Adjust Preload Count
Control exactly how many experts are preloaded during initialization:
--ssd-streaming-preload-experts 200
Step-by-Step Configuration Guide
To configure SSD streaming to run models larger than available RAM, follow these steps:
- Enable streaming – Add the
--ssd-streamingflag (requires Metal, CUDA, or ROCm). - Set cache policy – Either rely on automatic sizing, or specify a byte budget or expert count using
--ssd-streaming-cache-experts. - Configure preloading – Optionally add
--ssd-streaming-coldto disable preloading, or set--ssd-streaming-preload-experts Nfor a custom preload count.
The CLI option definitions are located in ds4_help.c around lines 170-174, while the parsing logic and cache planning live in ds4_ssd.c.
Practical Configuration Examples
These examples demonstrate real-world configurations for different hardware constraints and use cases.
Basic Streaming with Automatic Cache
Simplest configuration using automatic 80% working-set cache:
./ds4 -m ./ds4flash.gguf --ssd-streaming
Manual Byte-Budget Cache
Specify 32 GiB for routed-expert memory:
./ds4 -m ./ds4flash.gguf \
--ssd-streaming \
--ssd-streaming-cache-experts 32GB
Manual Expert-Count Cache
Cache exactly 4000 dynamic experts:
./ds4 -m ./ds4flash.gguf \
--ssd-streaming \
--ssd-streaming-cache-experts 4000
Disable Hot-Expert Preload
Useful for benchmarking cold-start behavior:
./ds4 -m ./ds4flash.gguf \
--ssd-streaming \
--ssd-streaming-cold
Adjust Preload Count
Preload exactly 200 experts at startup:
./ds4 -m ./ds4flash.gguf \
--ssd-streaming \
--ssd-streaming-preload-experts 200
MacBook with 64 GB RAM (Q2 Flash Model)
For a moderate cache with a 64 GB system:
./download_model.sh ds4f-q2
./ds4 -m ./ds4flash.gguf \
--ssd-streaming \
--ssd-streaming-cache-experts 32GB \
--ctx 32768 \
--nothink
MacBook with 128 GB RAM (DeepSeek-V4-Pro)
For larger models on high-memory laptops:
./download_model.sh pro-q2-imatrix
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
--ssd-streaming \
--ctx 32768 \
--nothink
Strix Halo with ROCm (128 GB)
For AMD ROCm systems with routed Q2_K models:
./download_model.sh glm-antirez-q2
make strix-halo
./ds4 --rocm -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
--ssd-streaming --ctx 4096
Key Source Files and Implementation
Understanding the source structure helps when debugging or extending functionality:
| File | Role |
|---|---|
ds4_ssd.c |
Parses cache arguments (ds4_parse_streaming_cache_experts_arg), computes automatic cache plans (ds4_ssd_auto_cache_plan), and converts byte budgets to expert counts (ds4_ssd_cache_experts_for_byte_budget). |
ds4_help.c |
Defines CLI flags --ssd-streaming, --ssd-streaming-cache-experts, --ssd-streaming-cold, and --ssd-streaming-preload-experts around lines 170-174. |
ds4.c |
Contains the core inference engine with the SSD-streaming decode path and backend-specific error handling (e.g., "Metal SSD streaming could not build …"). |
tests/ds4_test.c |
Test harness supporting SSD streaming via the DS4_TEST_SSD_STREAMING environment variable. |
Summary
- SSD streaming routes only active expert layers from disk, keeping a dynamic cache in RAM to run models exceeding physical memory.
- Three cache modes are available: automatic (default 80% of working-set via
DS4_SSD_AUTO_CACHE_PCT), manual byte budget (e.g.,32GB), and manual expert count (e.g.,4000). - Preload control via
--ssd-streaming-cold(disable) or--ssd-streaming-preload-experts N(custom count). - Backend support requires Metal, CUDA, or ROCm.
- Core implementation resides in
ds4_ssd.candds4_help.c, with runtime logic inds4.c.
Frequently Asked Questions
What hardware backends support SSD streaming in ds4?
SSD streaming requires Metal, CUDA, or ROCm backends. The feature is not available on CPU-only builds. According to the ds4.c source, the engine validates backend compatibility before initializing the SSD streaming path.
How is the automatic cache size calculated?
The automatic mode uses ds4_ssd_auto_cache_plan in ds4_ssd.c to query the backend's recommended working-set size, then applies the percentage specified by DS4_SSD_AUTO_CACHE_PCT (defaulting to 80%). This determines how many experts fit in the calculated memory budget without manual intervention.
Can I run a model that is twice the size of my RAM?
Yes. SSD streaming is specifically designed for this scenario. As long as your SSD has sufficient space for the GGUF file and you configure an appropriate cache size using --ssd-streaming-cache-experts, ds4 will route experts from disk while keeping only the working set in memory.
What is the performance impact of using --ssd-streaming-cold?
Disabling hot-expert preload with --ssd-streaming-cold increases initial latency for the first few tokens because experts must be loaded from SSD on first access rather than being preloaded into cache. This is primarily useful for benchmarking cold-start behavior or conserving memory during initialization, as implemented in the startup sequence of ds4.c.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →