DeepSeek V4 PRO Memory Requirements with SSD Streaming: Complete Guide
Running DeepSeek V4 PRO with SSD streaming requires approximately 80 GB of total RAM on a 128 GB system, with the engine automatically allocating ~59 GB for the routed-expert cache and ~20 GB for non-routed weights.
DeepSeek V4 PRO is one of the largest open-weight language models available, with routed-expert weights totaling roughly 90 GiB. The SSD streaming feature in the antirez/ds4 inference engine makes running this model feasible on consumer hardware by keeping only non-routed tensors resident in RAM while loading MoE experts on-demand from the GGUF file. Understanding these memory requirements helps you determine whether your hardware can run PRO and how to optimize cache allocation for your workload.
How SSD Streaming Reduces Memory Requirements
Without SSD streaming, DeepSeek V4 PRO would require loading the full 90+ GiB of weights into RAM—impossible on most consumer machines. The SSD-streaming architecture in ds4 solves this by treating the GGUF file as a memory-mapped backing store.
The engine maintains an in-memory expert cache for frequently accessed routed experts. This cache size is calculated automatically based on your system's available RAM, with the core logic implemented in ds4_ssd.c.
Automatic Cache Calculation
The cache-planning algorithm follows this sequence in ds4_ssd.c (lines 80-106):
- Receives a recommended working-set size from the backend
- Applies the automatic cache percentage — default 80%, overrideable via
DS4_SSD_AUTO_CACHE_PCT - Subtracts non-routed bytes (weights that must stay resident)
- Divides remaining bytes by per-expert size to get cached expert count
The final calculation caps at max_model_experts and UINT32_MAX for safety.
Memory Requirements by Hardware Class
| Machine Configuration | Viability | Approximate RAM Breakdown |
|---|---|---|
| 64 GB (low-end Mac) | ❌ Not supported | Non-routed weights alone exceed budget |
| 128 GB (M5 Max, M3 Ultra) | ✅ Recommended | ~20 GB non-routed + ~59 GB cache + ~5 GB KV/scratch = ~80 GB total |
| 256 GB+ (workstation) | ✅ Full-resident possible | SSD streaming still works with same automatic cache plan |
The 128 GB configuration is the practical minimum documented in the README (lines 37-42). On this hardware, the automatic plan yields roughly 59 GiB of routed-expert cache—sufficient for most inference patterns while leaving headroom for context windows and activation memory.
Controlling Cache Size Manually
While automatic cache sizing works well for most deployments, you can override it using the --ssd-streaming-cache-experts flag. This accepts either a byte budget with suffix (32GB, 48GB) or an explicit expert count.
# Automatic cache sizing (recommended for most users)
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
--ssd-streaming
# Force 32 GiB cache for routed experts
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
--ssd-streaming --ssd-streaming-cache-experts 32GB
The argument parser in ds4_ssd.c (lines 46-70) handles suffix parsing—GB multiplies by 1024³—then converts the byte budget to concrete expert slots via ds4_ssd_cache_experts_for_byte_budget.
Practical Configuration Examples
Standard 128 GB Deployment
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
--ssd-streaming --ctx 32768 --nothink
Uses the automatic ~59 GiB cache. Suitable for most production inference with 32K context.
Constrained Cache for Large Context Windows
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
--ssd-streaming --ssd-streaming-cache-experts 48GB --ctx 65536
Reduces expert cache to 48 GiB, freeing ~11 GiB for doubled context size. Expect slightly higher SSD read amplification.
Memory-Constrained Testing
./ds4 -m gguf/DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf \
--ssd-streaming --ssd-streaming-cache-experts 16GB --ctx 8192
Minimal viable configuration for functional testing. Significant performance degradation due to expert thrashing.
Key Source Files for Memory Management
| File | Purpose |
|---|---|
ds4_ssd.c |
Cache budget logic: ds4_ssd_auto_cache_percent(), ds4_ssd_auto_cache_plan() |
ds4_ssd.h |
Public API declarations: ds4_ssd_memory_lock_acquire(), cache planning structs |
ds4.c / ds4_cli.c |
CLI flag parsing for --ssd-streaming and --ssd-streaming-cache-experts |
README.md |
Empirical guidance: 59 GiB cache target for 128 GB Mac systems |
These files collectively define how ds4 determines memory requirements and exposes control to users.
Performance Implications of Cache Sizing
Larger cache → fewer SSD reads, lower latency, higher RAM pressure
Smaller cache → more SSD reads, higher latency, room for context/activations
The 80% default in ds4_ssd.c balances these factors for typical interactive use. For batch processing with redundant expert access patterns, you may benefit from manual expansion. For single-turn queries with diverse routing, contraction may improve throughput by allowing larger batch sizes.
Summary
- Minimum viable RAM: 128 GB for DeepSeek V4 PRO with SSD streaming; 64 GB systems cannot load non-routed weights
- Automatic allocation: ~59 GiB expert cache on 128 GB hardware, computed as 80% of backend recommendation minus non-routed bytes
- Override mechanism:
--ssd-streaming-cache-expertswithGBsuffix for byte budgets - Implementation core:
ds4_ssd.clines 46-70 (parsing), 80-106 (planning) - Total footprint estimate: ~80 GB on 128 GB system (non-routed + cache + overhead)
Frequently Asked Questions
Can I run DeepSeek V4 PRO with SSD streaming on a 64 GB Mac?
No. The non-routed weights alone exceed available RAM, making the model unloadable regardless of cache configuration. The README explicitly documents 128 GB as the practical minimum for PRO.
How does the engine decide how many experts to cache?
The engine calls ds4_ssd_auto_cache_plan() in ds4_ssd.c, which takes 80% of the backend's recommended working set, subtracts non-routed weight bytes, and divides by per-expert size. The result is capped to valid expert indices. Override via DS4_SSD_AUTO_CACHE_PCT or --ssd-streaming-cache-experts.
What happens if I set a cache larger than available RAM?
The allocation will likely fail during ds4_ssd_memory_lock_acquire() or cause system swapping. The automatic planner prevents this by deriving cache size from reported RAM. Manual overrides bypass this protection—verify your system's capacity before forcing large caches.
Is SSD streaming slower than full RAM residence?
For cache-friendly access patterns (repeated experts, common tokens), latency approaches full residence. For adversarial patterns (random expert selection), SSD latency dominates. The 59 GiB automatic cache captures the working set for most practical prompts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →