How SSD Streaming Works in Dwarf Star (ds4) and Recommended Cache Sizes for MacBook Models

SSD streaming in Dwarf Star (ds4) loads routed MoE experts from MacBook internal storage on-demand while keeping dense weights in GPU memory, with automatic cache budgeting at ~80% of available RAM that users can override per model.

Dwarf Star, the ds4 runtime from antirez/ds4, enables inference of large Mixture-of-Experts (MoE) models on Apple Silicon Macs that lack sufficient unified memory to load full models. The SSD streaming architecture separates model components by access pattern: non-routed (dense) layers stay resident, while routed experts fetch from fast internal SSD when cache misses occur.

Why SSD Streaming Makes Large MoE Models Runnable

Modern MacBook SSDs deliver sufficient bandwidth to mask on-demand expert loading during generation. According to the ds4 source, this trade-off succeeds because routed experts dominate model size—often 90%+ of parameters—while modern Mac SSDs are "fast enough to make cache misses tolerable" (see README.md lines 66-78). Dense weights, KV cache, and computational scratch space remain in memory where latency matters most.

How the Cache Budget System Works

The cache allocation in ds4_ssd.c follows a two-phase reservation strategy.

Automatic Budget Calculation

On startup, ds4 computes a working set budget targeting ~80% of available backend memory (controllable via DS4_SSD_AUTO_CACHE_PCT environment variable). The implementation in ds4_ssd.c lines 80-106 performs this calculation before partitioning resources.

Two-Phase Memory Split

  1. Non-routed memory reservation: Dense weights, KV cache, and scratch buffers claim their required bytes first.
  2. Routed-expert cache: Remaining bytes become the expert cache, converted to expert slots via ds4_ssd_cache_experts_for_byte_budget() (lines 71-78 in ds4_ssd.c).

Manual Override Options

The CLI provides precise control (documented in ds4_help.c lines 70-73):

Flag Purpose
--ssd-streaming Enable streaming mode
--ssd-streaming-cold Skip popularity-based preload of hot experts
--ssd-streaming-cache-experts N|NGB Set exact expert count or byte budget

If --ssd-streaming-cache-experts is omitted, the automatic ~80% plan applies.

The ds4 README provides tested configurations:

64 GB MacBook (M2/M2 Pro)

./ds4 -m ./ds4flash.gguf \
      --ssd-streaming \
      --ssd-streaming-cache-experts 32GB \
      --ctx 32768 \
      --nothink

For 2-bit Flash GGUF models, 32 GB leaves headroom for system operations and context window (README lines 11-20).

128 GB MacBook Pro M3 Max

Automatic mode (recommended starting point):

./ds4 -m ./ds4flash.gguf --ssd-streaming --ctx 32768 --nothink

The auto-budget typically selects ~59 GB expert cache (README lines 37-41).

Conservative manual tuning:

./ds4 -m ./ds4flash.gguf \
      --ssd-streaming \
      --ssd-streaming-cache-experts 48GB \
      --ctx 32768 \
      --think \
      --tokens 1500

Start at 48 GB, monitor responsiveness, then increase if system remains stable (README lines 40-43).

Mac Studio M3 Ultra (512 GB)

./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
      --ssd-streaming \
      --ctx 32768 \
      --nothink

With ample memory, automatic planning yields 64-75 GB expert cache without manual intervention (README lines 55-58).

AMD Strix Halo (128 GB Unified Memory)

./ds4 -m glm-5.2.gguf --ssd-streaming --ctx 4096

The automatic budget preserves space for graph execution and KV state on non-Apple silicon (README lines 66-71).

Implementation Files and Key Functions

File Function/Purpose
ds4_ssd.c ds4_ssd_cache_experts_for_byte_budget() — converts bytes to expert count; auto-budget logic at lines 80-106
ds4_help.c CLI documentation for --ssd-streaming* flags at lines 70-73
README.md Architectural rationale, model-specific examples, and 80% default target explanation

Summary

  • SSD streaming keeps dense weights in memory while routing experts load from MacBook SSD on demand.
  • Automatic cache budgeting targets ~80% of available RAM, split between non-routed needs and expert cache.
  • Override via --ssd-streaming-cache-experts using expert count (256) or byte budget (32GB).
  • Start with auto-mode, then tune down if startup warns of excessive cache or system pressure appears.

Frequently Asked Questions

How does ds4 decide how many experts to cache automatically?

The function implementing the automatic budget in ds4_ssd.c queries backend memory availability, applies the DS4_SSD_AUTO_CACHE_PCT percentage (default ~80%), reserves non-routed memory first, then converts remaining bytes to expert slots via ds4_ssd_cache_experts_for_byte_budget().

What happens if I set a cache size larger than available memory?

Ds4 validates against physical constraints during startup and emits a warning if the requested expert cache exceeds viable space. The runtime may fall back to smaller allocation or refuse initialization depending on backend implementation details in the Metal/MLX path.

Is --ssd-streaming-cold slower than the default preload?

Yes. The cold flag skips popularity-based preloading of frequently accessed experts, causing initial generations to incur more SSD fetch latency. Use this when you prioritize faster startup over initial prompt response time.

Can SSD streaming work with external Thunderbolt SSDs?

The ds4 implementation targets internal MacBook SSDs where latency and bandwidth characteristics are known. External SSDs vary in performance; while the code does not explicitly block them, cache miss latency may become intolerable on slower Thunderbolt or USB storage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →