How SSD Streaming Works in DwarfStar: Architecture and Memory Management
DwarfStar uses SSD streaming to keep large language model expert weights on local NVMe storage while caching only active experts in RAM, using a configurable percentage-based cache planner to prevent memory overcommit.
The ds4 repository implements DwarfStar, a small native inference engine for DeepSeek V4 Flash that enables running massive models on consumer hardware through intelligent SSD streaming. Instead of loading entire models into RAM, the system streams expert weights from NVMe on demand, employing sophisticated memory management to balance inference performance against system stability.
SSD Streaming Architecture
The streaming subsystem lives in the ds4_ssd module and operates on a cache-plan that dictates how many experts can remain resident based on safe RAM allocation limits.
Cache Planning Algorithm
The SSD cache planner (ds4_ssd_auto_cache_plan) dynamically computes the memory budget for resident experts. Implemented in ds4_ssd.c (lines 108-134), the planner reads the optional DS4_SSD_AUTO_CACHE_PCT environment variable (valid range 50-95%) and falls back to 80% if unset or invalid.
The algorithm calculates model_target_bytes as the percentage of the model's recommended RAM, then derives the actual cache budget by subtracting the "non-routed" memory already allocated:
/* ds4_ssd.c – auto-cache planning */
out->model_target_bytes =
recommended_bytes > UINT64_MAX / pct ?
UINT64_MAX : (recommended_bytes * pct) / 100ull;
if (out->model_target_bytes > non_routed_bytes) {
out->cache_bytes = out->model_target_bytes - non_routed_bytes;
}
The planner then determines the number of cacheable experts by dividing cache_bytes by per_expert_bytes, clamping the result to the model's maximum expert count and UINT32_MAX:
uint64_t cache_experts = out->cache_bytes / per_expert_bytes;
if (cache_experts == 0) cache_experts = 1;
out->cache_experts = (uint32_t)cache_experts;
out->effective_cache_bytes = cache_experts * per_expert_bytes;
Memory Lock Simulation
To guarantee that the OS does not swap out critical cache pages, DwarfStar optionally employs a memory-lock simulator (ds4_ssd_memory_lock_acquire and ds4_ssd_memory_lock_release). As implemented in ds4_ssd.c (lines 37-102), this utility maps the target RAM region and touches it in 256 MiB chunks, then locks each chunk with mlock. This ensures resident memory for low-latency SSD reads and prevents overcommitment.
Streaming Loader and LRU Eviction
When the inference engine requires an expert not currently resident, the streaming loader reads the weight slice from SSD into a pre-allocated buffer. If the cache is full, an LRU-style eviction policy frees space for the incoming expert, strictly maintaining the cache within the computed cache_bytes limit.
Unified Memory APU Support
On APU platforms where CPU and GPU share the same physical RAM, the cache lives in the same pool as the operating system. The default 80% limit provides empirical safety against the Linux OOM killer, though users can trade safety for performance by adjusting DS4_SSD_AUTO_CACHE_PCT.
Memory Management Implications
SSD streaming in DwarfStar introduces specific constraints and guarantees that affect system stability:
-
Predictable RAM Footprint: The cache size is strictly capped to a configurable percentage of the model's recommended memory, preventing the uncontrolled growth typical of greedy allocation strategies.
-
OOM Protection on Unified Systems: The 80% default provides a safety buffer on APU systems where GPU and CPU memory are shared, though raising this value via environment variables may trigger the OOM killer under memory pressure.
-
Memory Lock Guarantees: When using the
--simulate-used-memoryflag, themlockimplementation ensures cache pages remain resident, which is critical for maintaining consistent SSD streaming latency. -
Dynamic Cache Resizing: The planner recalculates when model parameters change, automatically clamping
cache_expertstoUINT32_MAXor the model's theoretical maximum expert count. -
Graceful Error Handling: If
mlockfails due to system limits (such asRLIMIT_MEMLOCK), the code releases already-locked portions and reports a clear diagnostic rather than proceeding with unstable memory.
Implementation Examples
Selecting a Cache Size
This example demonstrates computing a cache plan for a 64 GiB model with 8 GiB already allocated to non-routed components and 512 MiB per expert:
uint64_t model_bytes = 64ULL * 1024 * 1024 * 1024; // 64 GiB model recommendation
uint64_t non_routed = 8ULL * 1024 * 1024 * 1024; // 8 GiB already allocated
uint64_t per_expert = 512ULL * 1024 * 1024; // 512 MiB per expert
ds4_ssd_cache_plan plan;
bool ok = ds4_ssd_auto_cache_plan(
model_bytes, // recommended total RAM
non_routed, // RAM already used by non-routed parts
per_expert, // bytes per expert
0, // no explicit max (0 = unlimited)
&plan);
assert(ok);
printf("Cache bytes: %llu, cached experts: %u\n",
(unsigned long long)plan.cache_bytes,
plan.cache_experts);
Reference: ds4_ssd_auto_cache_plan implementation in ds4_ssd.c (lines 108-134).
Simulating Memory Reservation
Before loading the model, you can lock the cache memory to prevent swapping:
ds4_ssd_memory_lock lock;
if (!ds4_ssd_memory_lock_acquire(&lock, plan.cache_bytes)) {
fprintf(stderr, "Failed to reserve cache memory\n");
exit(1);
}
/* … load the model, stream experts into the reserved region … */
ds4_ssd_memory_lock_release(&lock);
Reference: Memory-lock implementation in ds4_ssd.c (lines 37-102).
Runtime Cache Percentage Override
Adjust the cache size without recompiling by setting the environment variable:
export DS4_SSD_AUTO_CACHE_PCT=90 # Use 90% of recommended RAM for cache
./ds4 -m my_deepseek_v4_flash.gguf --stream
If the value is outside the 50-95% range, the system falls back to 80% and prints a warning (see ds4_ssd.c lines 81-94).
Summary
- DwarfStar streams expert weights from NVMe SSD using the ds4_ssd module with LRU eviction to minimize RAM usage.
- The cache planner (
ds4_ssd_auto_cache_plan) defaults to caching 80% of recommended model RAM, configurable viaDS4_SSD_AUTO_CACHE_PCT. - Memory locking via
ds4_ssd_memory_lock_acquireprevents the OS from swapping critical cache pages, ensuring low-latency access. - The architecture includes specific safeguards for unified memory APU systems to prevent OOM kills.
- Core logic resides in
ds4_ssd.candds4_ssd.h, integrated throughds4.candds4_server.c.
Frequently Asked Questions
What is the default cache percentage for SSD streaming in DwarfStar?
The default cache percentage is 80% of the model's recommended RAM, as implemented in ds4_ssd_auto_cache_percent() within ds4_ssd.c (lines 80-106). You can override this by setting the DS4_SSD_AUTO_CACHE_PCT environment variable to any integer between 50 and 95.
How does DwarfStar prevent the Linux OOM killer from terminating the process?
DwarfStar prevents OOM kills by limiting the SSD cache to a safe fraction of total system memory (default 80%) and optionally using mlock via ds4_ssd_memory_lock_acquire() to pin cache pages in physical RAM. This prevents the OS from swapping out critical expert weights during inference, which is especially important on APU systems where CPU and GPU share the same memory pool.
Can I adjust the number of experts cached in RAM without reloading the model?
While the cache plan is typically computed at model load time via ds4_ssd_auto_cache_plan(), you can trigger recomputation by reloading the model with different parameters or adjusting the DS4_SSD_AUTO_CACHE_PCT environment variable before startup. The planner dynamically calculates cache_experts by dividing the cache budget by per_expert_bytes, automatically clamping to valid ranges.
What happens if the system doesn't have enough memory for the locked cache?
If ds4_ssd_memory_lock_acquire() fails to lock memory (for example, due to ulimit restrictions on locked pages), the function gracefully releases any already-locked portions, returns false, and the calling code aborts with a clear diagnostic message. This fail-safe prevents the process from proceeding with unstable memory that could be swapped out during critical SSD streaming operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →