# DeepSeek V4 PRO Memory Requirements with SSD Streaming: Complete Guide

> Discover DeepSeek V4 PRO memory needs with SSD streaming. Learn about RAM requirements, expert cache allocation, and non-routed weights for optimal performance on your 128GB system.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-04

---

**Running DeepSeek V4 PRO with SSD streaming requires approximately 80 GB of total RAM on a 128 GB system, with the engine automatically allocating ~59 GB for the routed-expert cache and ~20 GB for non-routed weights.**

DeepSeek V4 PRO is one of the largest open-weight language models available, with routed-expert weights totaling roughly 90 GiB. The **SSD streaming** feature in the `antirez/ds4` inference engine makes running this model feasible on consumer hardware by keeping only non-routed tensors resident in RAM while loading MoE experts on-demand from the GGUF file. Understanding these memory requirements helps you determine whether your hardware can run PRO and how to optimize cache allocation for your workload.

## How SSD Streaming Reduces Memory Requirements

Without SSD streaming, DeepSeek V4 PRO would require loading the full 90+ GiB of weights into RAM—impossible on most consumer machines. The **SSD-streaming architecture** in `ds4` solves this by treating the GGUF file as a memory-mapped backing store.

The engine maintains an in-memory **expert cache** for frequently accessed routed experts. This cache size is calculated automatically based on your system's available RAM, with the core logic implemented in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c).

### Automatic Cache Calculation

The cache-planning algorithm follows this sequence in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 80-106):

1. Receives a **recommended working-set size** from the backend
2. Applies the **automatic cache percentage** — default 80%, overrideable via `DS4_SSD_AUTO_CACHE_PCT`
3. Subtracts **non-routed bytes** (weights that must stay resident)
4. Divides remaining bytes by **per-expert size** to get cached expert count

The final calculation caps at `max_model_experts` and `UINT32_MAX` for safety.

## Memory Requirements by Hardware Class

| Machine Configuration | Viability | Approximate RAM Breakdown |
|-----------------------|-----------|---------------------------|
| **64 GB** (low-end Mac) | ❌ Not supported | Non-routed weights alone exceed budget |
| **128 GB** (M5 Max, M3 Ultra) | ✅ Recommended | ~20 GB non-routed + ~59 GB cache + ~5 GB KV/scratch = **~80 GB total** |
| **256 GB+** (workstation) | ✅ Full-resident possible | SSD streaming still works with same automatic cache plan |

The 128 GB configuration is the practical minimum documented in the README (lines 37-42). On this hardware, the automatic plan yields roughly **59 GiB of routed-expert cache**—sufficient for most inference patterns while leaving headroom for context windows and activation memory.

## Controlling Cache Size Manually

While automatic cache sizing works well for most deployments, you can override it using the `--ssd-streaming-cache-experts` flag. This accepts either a **byte budget** with suffix (`32GB`, `48GB`) or an explicit expert count.

```bash

# Automatic cache sizing (recommended for most users)

./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
      --ssd-streaming

# Force 32 GiB cache for routed experts

./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
      --ssd-streaming --ssd-streaming-cache-experts 32GB

```

The argument parser in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 46-70) handles suffix parsing—`GB` multiplies by 1024³—then converts the byte budget to concrete expert slots via `ds4_ssd_cache_experts_for_byte_budget`.

## Practical Configuration Examples

### Standard 128 GB Deployment

```bash
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
      --ssd-streaming --ctx 32768 --nothink

```

Uses the automatic ~59 GiB cache. Suitable for most production inference with 32K context.

### Constrained Cache for Large Context Windows

```bash
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
      --ssd-streaming --ssd-streaming-cache-experts 48GB --ctx 65536

```

Reduces expert cache to 48 GiB, freeing ~11 GiB for doubled context size. Expect slightly higher SSD read amplification.

### Memory-Constrained Testing

```bash
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
      --ssd-streaming --ssd-streaming-cache-experts 16GB --ctx 8192

```

Minimal viable configuration for functional testing. Significant performance degradation due to expert thrashing.

## Key Source Files for Memory Management

| File | Purpose |
|------|---------|
| [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) | Cache budget logic: `ds4_ssd_auto_cache_percent()`, `ds4_ssd_auto_cache_plan()` |
| [`ds4_ssd.h`](https://github.com/antirez/ds4/blob/main/ds4_ssd.h) | Public API declarations: `ds4_ssd_memory_lock_acquire()`, cache planning structs |
| [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) / [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) | CLI flag parsing for `--ssd-streaming` and `--ssd-streaming-cache-experts` |
| [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) | Empirical guidance: 59 GiB cache target for 128 GB Mac systems |

These files collectively define how `ds4` determines memory requirements and exposes control to users.

## Performance Implications of Cache Sizing

**Larger cache** → fewer SSD reads, lower latency, higher RAM pressure  
**Smaller cache** → more SSD reads, higher latency, room for context/activations

The 80% default in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) balances these factors for typical interactive use. For batch processing with redundant expert access patterns, you may benefit from manual expansion. For single-turn queries with diverse routing, contraction may improve throughput by allowing larger batch sizes.

## Summary

- **Minimum viable RAM**: 128 GB for DeepSeek V4 PRO with SSD streaming; 64 GB systems cannot load non-routed weights
- **Automatic allocation**: ~59 GiB expert cache on 128 GB hardware, computed as 80% of backend recommendation minus non-routed bytes
- **Override mechanism**: `--ssd-streaming-cache-experts` with `GB` suffix for byte budgets
- **Implementation core**: [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) lines 46-70 (parsing), 80-106 (planning)
- **Total footprint estimate**: ~80 GB on 128 GB system (non-routed + cache + overhead)

## Frequently Asked Questions

### Can I run DeepSeek V4 PRO with SSD streaming on a 64 GB Mac?

No. The non-routed weights alone exceed available RAM, making the model unloadable regardless of cache configuration. The README explicitly documents 128 GB as the practical minimum for PRO.

### How does the engine decide how many experts to cache?

The engine calls `ds4_ssd_auto_cache_plan()` in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c), which takes 80% of the backend's recommended working set, subtracts non-routed weight bytes, and divides by per-expert size. The result is capped to valid expert indices. Override via `DS4_SSD_AUTO_CACHE_PCT` or `--ssd-streaming-cache-experts`.

### What happens if I set a cache larger than available RAM?

The allocation will likely fail during `ds4_ssd_memory_lock_acquire()` or cause system swapping. The automatic planner prevents this by deriving cache size from reported RAM. Manual overrides bypass this protection—verify your system's capacity before forcing large caches.

### Is SSD streaming slower than full RAM residence?

For cache-friendly access patterns (repeated experts, common tokens), latency approaches full residence. For adversarial patterns (random expert selection), SSD latency dominates. The 59 GiB automatic cache captures the working set for most practical prompts.