# How SSD Streaming Works in Dwarf Star (ds4) and Recommended Cache Sizes for MacBook Models

> Discover how SSD streaming in Dwarf Star ds4 loads expert MoE models on-demand. Learn recommended cache sizes for MacBooks to optimize performance.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: performance
- Published: 2026-08-04

---

**SSD streaming in Dwarf Star (ds4) loads routed MoE experts from MacBook internal storage on-demand while keeping dense weights in GPU memory, with automatic cache budgeting at ~80% of available RAM that users can override per model.**

Dwarf Star, the `ds4` runtime from antirez/ds4, enables inference of large Mixture-of-Experts (MoE) models on Apple Silicon Macs that lack sufficient unified memory to load full models. The **SSD streaming** architecture separates model components by access pattern: non-routed (dense) layers stay resident, while routed experts fetch from fast internal SSD when cache misses occur.

## Why SSD Streaming Makes Large MoE Models Runnable

Modern MacBook SSDs deliver sufficient bandwidth to mask on-demand expert loading during generation. According to the ds4 source, this trade-off succeeds because **routed experts dominate model size**—often 90%+ of parameters—while modern Mac SSDs are "fast enough to make cache misses tolerable" (see [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) lines 66-78). Dense weights, KV cache, and computational scratch space remain in memory where latency matters most.

## How the Cache Budget System Works

The cache allocation in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) follows a two-phase reservation strategy.

### Automatic Budget Calculation

On startup, ds4 computes a working set budget targeting **~80% of available backend memory** (controllable via `DS4_SSD_AUTO_CACHE_PCT` environment variable). The implementation in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) lines 80-106 performs this calculation before partitioning resources.

### Two-Phase Memory Split

1. **Non-routed memory reservation**: Dense weights, KV cache, and scratch buffers claim their required bytes first.
2. **Routed-expert cache**: Remaining bytes become the expert cache, converted to expert slots via `ds4_ssd_cache_experts_for_byte_budget()` (lines 71-78 in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c)).

### Manual Override Options

The CLI provides precise control (documented in [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) lines 70-73):

| Flag | Purpose |
|------|---------|
| `--ssd-streaming` | Enable streaming mode |
| `--ssd-streaming-cold` | Skip popularity-based preload of hot experts |
| `--ssd-streaming-cache-experts N\|NGB` | Set exact expert count or byte budget |

If `--ssd-streaming-cache-experts` is omitted, the automatic ~80% plan applies.

## Recommended Cache Sizes by MacBook Model

The ds4 README provides tested configurations:

### 64 GB MacBook (M2/M2 Pro)

```bash
./ds4 -m ./ds4flash.gguf \
      --ssd-streaming \
      --ssd-streaming-cache-experts 32GB \
      --ctx 32768 \
      --nothink

```

For 2-bit Flash GGUF models, 32 GB leaves headroom for system operations and context window (README lines 11-20).

### 128 GB MacBook Pro M3 Max

**Automatic mode** (recommended starting point):

```bash
./ds4 -m ./ds4flash.gguf --ssd-streaming --ctx 32768 --nothink

```

The auto-budget typically selects **~59 GB expert cache** (README lines 37-41).

**Conservative manual tuning**:

```bash
./ds4 -m ./ds4flash.gguf \
      --ssd-streaming \
      --ssd-streaming-cache-experts 48GB \
      --ctx 32768 \
      --think \
      --tokens 1500

```

Start at 48 GB, monitor responsiveness, then increase if system remains stable (README lines 40-43).

### Mac Studio M3 Ultra (512 GB)

```bash
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
      --ssd-streaming \
      --ctx 32768 \
      --nothink

```

With ample memory, automatic planning yields **64-75 GB expert cache** without manual intervention (README lines 55-58).

### AMD Strix Halo (128 GB Unified Memory)

```bash
./ds4 -m glm-5.2.gguf --ssd-streaming --ctx 4096

```

The automatic budget preserves space for graph execution and KV state on non-Apple silicon (README lines 66-71).

## Implementation Files and Key Functions

| File | Function/Purpose |
|------|----------------|
| [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) | `ds4_ssd_cache_experts_for_byte_budget()` — converts bytes to expert count; auto-budget logic at lines 80-106 |
| [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) | CLI documentation for `--ssd-streaming*` flags at lines 70-73 |
| [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) | Architectural rationale, model-specific examples, and 80% default target explanation |

## Summary

- **SSD streaming** keeps dense weights in memory while routing experts load from MacBook SSD on demand.
- **Automatic cache budgeting** targets ~80% of available RAM, split between non-routed needs and expert cache.
- **Override via `--ssd-streaming-cache-experts`** using expert count (`256`) or byte budget (`32GB`).
- **Start with auto-mode, then tune down** if startup warns of excessive cache or system pressure appears.

## Frequently Asked Questions

### How does ds4 decide how many experts to cache automatically?

The function implementing the automatic budget in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) queries backend memory availability, applies the `DS4_SSD_AUTO_CACHE_PCT` percentage (default ~80%), reserves non-routed memory first, then converts remaining bytes to expert slots via `ds4_ssd_cache_experts_for_byte_budget()`.

### What happens if I set a cache size larger than available memory?

Ds4 validates against physical constraints during startup and emits a warning if the requested expert cache exceeds viable space. The runtime may fall back to smaller allocation or refuse initialization depending on backend implementation details in the Metal/MLX path.

### Is `--ssd-streaming-cold` slower than the default preload?

Yes. The cold flag skips popularity-based preloading of frequently accessed experts, causing initial generations to incur more SSD fetch latency. Use this when you prioritize faster startup over initial prompt response time.

### Can SSD streaming work with external Thunderbolt SSDs?

The ds4 implementation targets internal MacBook SSDs where latency and bandwidth characteristics are known. External SSDs vary in performance; while the code does not explicitly block them, cache miss latency may become intolerable on slower Thunderbolt or USB storage.