How SSD Streaming Works in Dwarf Star (ds4) and Recommended Cache Sizes for MacBook Models
SSD streaming in Dwarf Star (ds4) loads routed MoE experts from MacBook internal storage on-demand while keeping dense weights in GPU memory, with automatic cache budgeting at ~80% of available RAM that users can override per model.
Dwarf Star, the ds4 runtime from antirez/ds4, enables inference of large Mixture-of-Experts (MoE) models on Apple Silicon Macs that lack sufficient unified memory to load full models. The SSD streaming architecture separates model components by access pattern: non-routed (dense) layers stay resident, while routed experts fetch from fast internal SSD when cache misses occur.
Why SSD Streaming Makes Large MoE Models Runnable
Modern MacBook SSDs deliver sufficient bandwidth to mask on-demand expert loading during generation. According to the ds4 source, this trade-off succeeds because routed experts dominate model size—often 90%+ of parameters—while modern Mac SSDs are "fast enough to make cache misses tolerable" (see README.md lines 66-78). Dense weights, KV cache, and computational scratch space remain in memory where latency matters most.
How the Cache Budget System Works
The cache allocation in ds4_ssd.c follows a two-phase reservation strategy.
Automatic Budget Calculation
On startup, ds4 computes a working set budget targeting ~80% of available backend memory (controllable via DS4_SSD_AUTO_CACHE_PCT environment variable). The implementation in ds4_ssd.c lines 80-106 performs this calculation before partitioning resources.
Two-Phase Memory Split
- Non-routed memory reservation: Dense weights, KV cache, and scratch buffers claim their required bytes first.
- Routed-expert cache: Remaining bytes become the expert cache, converted to expert slots via
ds4_ssd_cache_experts_for_byte_budget()(lines 71-78 inds4_ssd.c).
Manual Override Options
The CLI provides precise control (documented in ds4_help.c lines 70-73):
| Flag | Purpose |
|---|---|
--ssd-streaming |
Enable streaming mode |
--ssd-streaming-cold |
Skip popularity-based preload of hot experts |
--ssd-streaming-cache-experts N|NGB |
Set exact expert count or byte budget |
If --ssd-streaming-cache-experts is omitted, the automatic ~80% plan applies.
Recommended Cache Sizes by MacBook Model
The ds4 README provides tested configurations:
64 GB MacBook (M2/M2 Pro)
./ds4 -m ./ds4flash.gguf \
--ssd-streaming \
--ssd-streaming-cache-experts 32GB \
--ctx 32768 \
--nothink
For 2-bit Flash GGUF models, 32 GB leaves headroom for system operations and context window (README lines 11-20).
128 GB MacBook Pro M3 Max
Automatic mode (recommended starting point):
./ds4 -m ./ds4flash.gguf --ssd-streaming --ctx 32768 --nothink
The auto-budget typically selects ~59 GB expert cache (README lines 37-41).
Conservative manual tuning:
./ds4 -m ./ds4flash.gguf \
--ssd-streaming \
--ssd-streaming-cache-experts 48GB \
--ctx 32768 \
--think \
--tokens 1500
Start at 48 GB, monitor responsiveness, then increase if system remains stable (README lines 40-43).
Mac Studio M3 Ultra (512 GB)
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
--ssd-streaming \
--ctx 32768 \
--nothink
With ample memory, automatic planning yields 64-75 GB expert cache without manual intervention (README lines 55-58).
AMD Strix Halo (128 GB Unified Memory)
./ds4 -m glm-5.2.gguf --ssd-streaming --ctx 4096
The automatic budget preserves space for graph execution and KV state on non-Apple silicon (README lines 66-71).
Implementation Files and Key Functions
| File | Function/Purpose |
|---|---|
ds4_ssd.c |
ds4_ssd_cache_experts_for_byte_budget() — converts bytes to expert count; auto-budget logic at lines 80-106 |
ds4_help.c |
CLI documentation for --ssd-streaming* flags at lines 70-73 |
README.md |
Architectural rationale, model-specific examples, and 80% default target explanation |
Summary
- SSD streaming keeps dense weights in memory while routing experts load from MacBook SSD on demand.
- Automatic cache budgeting targets ~80% of available RAM, split between non-routed needs and expert cache.
- Override via
--ssd-streaming-cache-expertsusing expert count (256) or byte budget (32GB). - Start with auto-mode, then tune down if startup warns of excessive cache or system pressure appears.
Frequently Asked Questions
How does ds4 decide how many experts to cache automatically?
The function implementing the automatic budget in ds4_ssd.c queries backend memory availability, applies the DS4_SSD_AUTO_CACHE_PCT percentage (default ~80%), reserves non-routed memory first, then converts remaining bytes to expert slots via ds4_ssd_cache_experts_for_byte_budget().
What happens if I set a cache size larger than available memory?
Ds4 validates against physical constraints during startup and emits a warning if the requested expert cache exceeds viable space. The runtime may fall back to smaller allocation or refuse initialization depending on backend implementation details in the Metal/MLX path.
Is --ssd-streaming-cold slower than the default preload?
Yes. The cold flag skips popularity-based preloading of frequently accessed experts, causing initial generations to incur more SSD fetch latency. Use this when you prioritize faster startup over initial prompt response time.
Can SSD streaming work with external Thunderbolt SSDs?
The ds4 implementation targets internal MacBook SSDs where latency and bandwidth characteristics are known. External SSDs vary in performance; while the code does not explicitly block them, cache miss latency may become intolerable on slower Thunderbolt or USB storage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →