# How SSD Streaming Works in DwarfStar: Architecture and Memory Management

> Learn how DwarfStar implements SSD streaming for efficient LLM expert weight management. Discover its architecture and memory management strategies for optimal performance.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: architecture
- Published: 2026-08-09

---

**DwarfStar uses SSD streaming to keep large language model expert weights on local NVMe storage while caching only active experts in RAM, using a configurable percentage-based cache planner to prevent memory overcommit.**

The **ds4** repository implements DwarfStar, a small native inference engine for DeepSeek V4 Flash that enables running massive models on consumer hardware through intelligent SSD streaming. Instead of loading entire models into RAM, the system streams expert weights from NVMe on demand, employing sophisticated memory management to balance inference performance against system stability.

## SSD Streaming Architecture

The streaming subsystem lives in the **ds4_ssd** module and operates on a *cache-plan* that dictates how many experts can remain resident based on safe RAM allocation limits.

### Cache Planning Algorithm

The **SSD cache planner** (`ds4_ssd_auto_cache_plan`) dynamically computes the memory budget for resident experts. Implemented in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 108-134), the planner reads the optional `DS4_SSD_AUTO_CACHE_PCT` environment variable (valid range 50-95%) and falls back to **80%** if unset or invalid.

The algorithm calculates `model_target_bytes` as the percentage of the model's recommended RAM, then derives the actual cache budget by subtracting the "non-routed" memory already allocated:

```c
/* ds4_ssd.c – auto-cache planning */
out->model_target_bytes =
    recommended_bytes > UINT64_MAX / pct ?
        UINT64_MAX : (recommended_bytes * pct) / 100ull;
if (out->model_target_bytes > non_routed_bytes) {
    out->cache_bytes = out->model_target_bytes - non_routed_bytes;
}

```

The planner then determines the number of cacheable experts by dividing `cache_bytes` by `per_expert_bytes`, clamping the result to the model's maximum expert count and `UINT32_MAX`:

```c
uint64_t cache_experts = out->cache_bytes / per_expert_bytes;
if (cache_experts == 0) cache_experts = 1;
out->cache_experts = (uint32_t)cache_experts;
out->effective_cache_bytes = cache_experts * per_expert_bytes;

```

### Memory Lock Simulation

To guarantee that the OS does not swap out critical cache pages, DwarfStar optionally employs a **memory-lock simulator** (`ds4_ssd_memory_lock_acquire` and `ds4_ssd_memory_lock_release`). As implemented in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 37-102), this utility maps the target RAM region and touches it in 256 MiB chunks, then locks each chunk with `mlock`. This ensures resident memory for low-latency SSD reads and prevents overcommitment.

### Streaming Loader and LRU Eviction

When the inference engine requires an expert not currently resident, the **streaming loader** reads the weight slice from SSD into a pre-allocated buffer. If the cache is full, an LRU-style eviction policy frees space for the incoming expert, strictly maintaining the cache within the computed `cache_bytes` limit.

### Unified Memory APU Support

On APU platforms where CPU and GPU share the same physical RAM, the cache lives in the same pool as the operating system. The default 80% limit provides empirical safety against the Linux OOM killer, though users can trade safety for performance by adjusting `DS4_SSD_AUTO_CACHE_PCT`.

## Memory Management Implications

SSD streaming in DwarfStar introduces specific constraints and guarantees that affect system stability:

- **Predictable RAM Footprint**: The cache size is strictly capped to a configurable percentage of the model's recommended memory, preventing the uncontrolled growth typical of greedy allocation strategies.

- **OOM Protection on Unified Systems**: The 80% default provides a safety buffer on APU systems where GPU and CPU memory are shared, though raising this value via environment variables may trigger the OOM killer under memory pressure.

- **Memory Lock Guarantees**: When using the `--simulate-used-memory` flag, the `mlock` implementation ensures cache pages remain resident, which is critical for maintaining consistent SSD streaming latency.

- **Dynamic Cache Resizing**: The planner recalculates when model parameters change, automatically clamping `cache_experts` to `UINT32_MAX` or the model's theoretical maximum expert count.

- **Graceful Error Handling**: If `mlock` fails due to system limits (such as `RLIMIT_MEMLOCK`), the code releases already-locked portions and reports a clear diagnostic rather than proceeding with unstable memory.

## Implementation Examples

### Selecting a Cache Size

This example demonstrates computing a cache plan for a 64 GiB model with 8 GiB already allocated to non-routed components and 512 MiB per expert:

```c
uint64_t model_bytes = 64ULL * 1024 * 1024 * 1024;   // 64 GiB model recommendation
uint64_t non_routed = 8ULL * 1024 * 1024 * 1024;    // 8 GiB already allocated
uint64_t per_expert = 512ULL * 1024 * 1024;          // 512 MiB per expert

ds4_ssd_cache_plan plan;
bool ok = ds4_ssd_auto_cache_plan(
    model_bytes,               // recommended total RAM
    non_routed,                // RAM already used by non-routed parts
    per_expert,                // bytes per expert
    0,                         // no explicit max (0 = unlimited)
    &plan);
assert(ok);
printf("Cache bytes: %llu, cached experts: %u\n",
       (unsigned long long)plan.cache_bytes,
       plan.cache_experts);

```

*Reference: `ds4_ssd_auto_cache_plan` implementation in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 108-134).*

### Simulating Memory Reservation

Before loading the model, you can lock the cache memory to prevent swapping:

```c
ds4_ssd_memory_lock lock;
if (!ds4_ssd_memory_lock_acquire(&lock, plan.cache_bytes)) {
    fprintf(stderr, "Failed to reserve cache memory\n");
    exit(1);
}

/* … load the model, stream experts into the reserved region … */

ds4_ssd_memory_lock_release(&lock);

```

*Reference: Memory-lock implementation in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 37-102).*

### Runtime Cache Percentage Override

Adjust the cache size without recompiling by setting the environment variable:

```bash
export DS4_SSD_AUTO_CACHE_PCT=90   # Use 90% of recommended RAM for cache

./ds4 -m my_deepseek_v4_flash.gguf --stream

```

If the value is outside the 50-95% range, the system falls back to 80% and prints a warning (see [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) lines 81-94).

## Summary

- DwarfStar streams expert weights from NVMe SSD using the **ds4_ssd** module with LRU eviction to minimize RAM usage.
- The **cache planner** (`ds4_ssd_auto_cache_plan`) defaults to caching 80% of recommended model RAM, configurable via `DS4_SSD_AUTO_CACHE_PCT`.
- **Memory locking** via `ds4_ssd_memory_lock_acquire` prevents the OS from swapping critical cache pages, ensuring low-latency access.
- The architecture includes specific safeguards for **unified memory APU** systems to prevent OOM kills.
- Core logic resides in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) and [`ds4_ssd.h`](https://github.com/antirez/ds4/blob/main/ds4_ssd.h), integrated through [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) and [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c).

## Frequently Asked Questions

### What is the default cache percentage for SSD streaming in DwarfStar?

The default cache percentage is **80%** of the model's recommended RAM, as implemented in `ds4_ssd_auto_cache_percent()` within [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 80-106). You can override this by setting the `DS4_SSD_AUTO_CACHE_PCT` environment variable to any integer between 50 and 95.

### How does DwarfStar prevent the Linux OOM killer from terminating the process?

DwarfStar prevents OOM kills by limiting the SSD cache to a safe fraction of total system memory (default 80%) and optionally using `mlock` via `ds4_ssd_memory_lock_acquire()` to pin cache pages in physical RAM. This prevents the OS from swapping out critical expert weights during inference, which is especially important on APU systems where CPU and GPU share the same memory pool.

### Can I adjust the number of experts cached in RAM without reloading the model?

While the cache plan is typically computed at model load time via `ds4_ssd_auto_cache_plan()`, you can trigger recomputation by reloading the model with different parameters or adjusting the `DS4_SSD_AUTO_CACHE_PCT` environment variable before startup. The planner dynamically calculates `cache_experts` by dividing the cache budget by `per_expert_bytes`, automatically clamping to valid ranges.

### What happens if the system doesn't have enough memory for the locked cache?

If `ds4_ssd_memory_lock_acquire()` fails to lock memory (for example, due to `ulimit` restrictions on locked pages), the function gracefully releases any already-locked portions, returns `false`, and the calling code aborts with a clear diagnostic message. This fail-safe prevents the process from proceeding with unstable memory that could be swapped out during critical SSD streaming operations.