How ds4 SSD Streaming Runs Models Larger Than Available RAM: Architecture and Implementation

The ds4 engine executes models exceeding system memory by keeping non-routed weights resident in RAM while streaming routed Mixture-of-Experts (MoE) weights from an SSD on demand, using a dynamic expert cache to balance speed and memory usage.

The ds4 inference engine by antirez solves the fundamental problem of running massive language models—often 150 GB or larger—on consumer hardware with limited RAM. Through a technique called ds4 SSD streaming, the system loads only essential weights into memory while fetching expert layers from fast local storage as needed. This article examines the implementation details in the antirez/ds4 repository to explain how the cache budget is calculated, how expert fetching works at runtime, and how to configure the streaming behavior.

How ds4 SSD Streaming Works

Separating Resident and Streaming Weights

The architecture divides model weights into two distinct categories based on access patterns. Non-routed weights—including dense layers, embeddings, and shared experts—are loaded into RAM at startup and remain resident throughout inference. Routed MoE experts, which comprise the bulk of large models, are stored in the GGUF file on an SSD. Only a subset of these experts, determined by the expert cache size, resides in memory at any given time. When the decoder encounters a token requiring an expert not present in the cache, the system performs a random read from the SSD to fetch the weights.

Automatic Cache Budget Calculation

On startup, ds4 automatically determines the optimal cache size through ds4_ssd_auto_cache_plan() in ds4_ssd.c (lines 108-135). This function accepts the backend-recommended working-set size (recommended_bytes) and applies a configurable percentage defined by DS4_SSD_AUTO_CACHE_PCT, which defaults to 80%. It subtracts the bytes required for non-routed weights (non_routed_bytes), then converts the remaining memory budget into a number of dynamic expert slots (cache_experts) based on per_expert_bytes. The resulting value determines exactly how many experts can remain resident before eviction occurs.

Manual Cache Configuration

Users can override the automatic calculation using the --ssd-streaming-cache-experts flag, parsed by ds4_parse_streaming_cache_experts_arg() in ds4_ssd.c (lines 46-70). This option accepts either an integer count of expert slots or a byte budget with the GB suffix. For example, specifying 32GB directs ds4 to allocate approximately that much memory to the expert cache, while a bare integer like 4000 reserves exactly that many slots regardless of available system memory.

Runtime Expert Fetching and Eviction

During generation, each token may route to multiple experts. When a required expert is not in the cache, the decoder issues an SSD read operation, copies the expert weights into an available cache slot, and evicts the least-recently-used expert if necessary. Because this process occurs per-token, generation speed becomes bandwidth-bound rather than compute-bound, resulting in slower throughput compared to fully resident models but dramatically reduced RAM requirements.

Prefill Optimizations and Hot-Expert Preloading

To mitigate SSD latency during the initial prefill phase, ds4 preloads a small set of "hot" experts—the most statistically likely candidates for the current context—before generation begins. These hot-expert lists are pre-computed from the GGUF routing tables and stored in ds4_streaming_hotlist.inc and ds4_streaming_hotlist_glm52.inc. Users can disable this optimization using the --ssd-streaming-cold flag to measure raw SSD performance or reduce startup time.

Memory-Lock Simulation for Testing

For debugging out-of-memory conditions, ds4 can simulate memory pressure via ds4_ssd_memory_lock_acquire() and ds4_ssd_memory_lock_release() in ds4_ssd.c (lines 37-102). This functionality mmaps a region of the specified size, touches it in chunks to ensure physical page allocation, and applies mlock to prevent the kernel from swapping it out. This effectively reserves memory to simulate a system with less available RAM than physically installed.

Configuring ds4 SSD Streaming

Enable streaming mode and specify cache sizes using command-line flags:


# Automatic cache sizing (uses 80% of recommended working set)

./ds4 -m ./ds4flash.gguf --ssd-streaming

# Specify cache size in gigabytes

./ds4 -m ./ds4flash.gguf \
      --ssd-streaming \
      --ssd-streaming-cache-experts 32GB \
      --ctx 32768

# Specify exact number of expert slots

./ds4 -m ./ds4flash.gguf \
      --ssd-streaming \
      --ssd-streaming-cache-experts 4000

# Disable hot-expert preloading for latency testing

./ds4 -m ./ds4flash.gguf \
      --ssd-streaming \
      --ssd-streaming-cold

# Simulate 8GB of already-used memory to test OOM behavior

./ds4 -m ./ds4flash.gguf \
      --ssd-streaming \
      --simulate-used-memory 8GB

Key Implementation Files

The ds4 SSD streaming mechanism spans several core files:

  • ds4_ssd.h — Defines the public API structures and parsing helpers for SSD streaming configuration.
  • ds4_ssd.c — Implements ds4_ssd_auto_cache_plan(), ds4_parse_streaming_cache_experts_arg(), and the memory-lock simulation functions.
  • ds4.c — Contains the core engine logic that invokes the cache planner and manages the expert cache during inference.
  • ds4_server.c — Handles CLI argument parsing for --ssd-streaming, --ssd-streaming-cache-experts, and related flags.
  • ds4_streaming_hotlist.inc / ds4_streaming_hotlist_glm52.inc — Pre-computed hot-expert lists for Flash and GLM 5.2 models to optimize prefill performance.

Summary

  • ds4 SSD streaming separates non-routed weights (resident in RAM) from routed MoE experts (stored on SSD) to run models larger than available memory.
  • The ds4_ssd_auto_cache_plan() function automatically calculates the expert cache size using 80% of the recommended working set minus non-routed weight requirements.
  • Users override automatic sizing via --ssd-streaming-cache-experts, accepting either slot counts or byte budgets (e.g., 32GB).
  • Cache misses trigger random SSD reads during token generation, making the system bandwidth-bound but memory-efficient.
  • Hot-expert preloading during the prefill phase reduces latency by pre-fetching likely experts from the routing tables.
  • The memory-lock simulation functions allow testing OOM scenarios by reserving physical RAM with mlock.

Frequently Asked Questions

What hardware is required for ds4 SSD streaming?

You need a fast NVMe SSD to maintain acceptable token generation speeds, as the system relies on random read performance for cache misses. While ds4 can run 150 GB models on machines with as little as 64 GB of RAM, the SSD must have sufficient capacity to store the GGUF file and sufficient IOPS to handle multiple expert fetches per token.

How does ds4 calculate the automatic cache size?

The ds4_ssd_auto_cache_plan() function in ds4_ssd.c takes the backend's recommended working-set size, multiplies it by DS4_SSD_AUTO_CACHE_PCT (default 80%), subtracts the bytes consumed by non-routed weights, and divides the remainder by the bytes per expert to determine the number of cache slots. This ensures the operating system retains sufficient free memory for buffers and system processes.

Can I disable expert preloading to measure raw SSD performance?

Yes. Pass the --ssd-streaming-cold flag to prevent ds4 from loading hot experts during the prefill phase. This forces every expert lookup to hit the SSD, useful for benchmarking storage latency or testing worst-case generation speeds.

Why doesn't ds4 stream all weights from the SSD?

Non-routed weights—dense layers, embeddings, and shared experts—are accessed for every token and are latency-sensitive. Streaming these would create a severe bottleneck. Routed experts are sparse (only a few are activated per token), making them ideal for on-demand streaming. This hybrid approach minimizes RAM usage while maintaining acceptable inference latency for the critical path.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →