How the 80% Backend Working Set Calculation Works for Automatic SSD Streaming Cache Sizing in ds4
The ds4 inference engine automatically sizes its SSD streaming expert cache by taking 80% of the GPU backend's recommended working set, subtracting non-routed memory overhead, and converting the remainder into a resident expert count.
The antirez/ds4 repository implements an intelligent cache sizing mechanism that balances GPU memory pressure against SSD streaming performance. Rather than requiring manual tuning, the system queries each backend (CUDA, ROCm, Metal) for its optimal working set, then derives a cache budget through a deterministic percentage-based calculation.
Backend Working Set Discovery
Before any percentage can be applied, ds4 must determine how much GPU memory the backend recommends for model execution. This value varies by platform and is surfaced through ds4_gpu_recommended_working_set_size():
- ROCm: Queries
hipMemGetInfofor total device memory viacudaMemGetInfocompatibility layer inrocm/ds4_rocm_current_api_compat.cuh(lines 44-53) - Metal: Reads the
recommendedMaxWorkingSetSizeproperty inds4_metal.m(lines 95-99) - CUDA: Returns aggregate VRAM across all available GPUs (
total_b * n) inds4_cuda.cu(lines 54-63)
Each implementation returns a uint64_t representing the total device memory the backend considers usable for the model.
The 80% Default and DS4_SSD_AUTO_CACHE_PCT Override
The percentage applied to the working set recommendation is controlled by ds4_ssd_auto_cache_percent() in ds4_ssd.c (lines 80-106). The logic follows this priority:
- Check for
DS4_SSD_AUTO_CACHE_PCTenvironment variable - Validate the value falls within 50-95 (inclusive)
- Fall back to 80 if unset or invalid, issuing a warning for out-of-range values
# Override the default before launching ds4
export DS4_SSD_AUTO_CACHE_PCT=90
./ds4-cli --model model.gguf --ssd-streaming
The 80% default represents a balance between cache hit rate and headroom for dynamic allocations during inference.
The Four-Step Cache Budget Calculation
ds4_ssd_auto_cache_plan() in ds4_ssd.c (lines 108-135) executes the complete calculation:
// 1. Apply percentage to working set recommendation
uint64_t model_target_bytes = recommended_bytes * pct / 100;
// 2. Subtract non-routed memory (context buffers, KV cache, pre-fill)
uint64_t cache_bytes = (model_target_bytes > non_routed_bytes) ?
model_target_bytes - non_routed_bytes : 0;
// 3. Convert bytes to expert count
uint64_t cache_experts = cache_bytes / per_expert_bytes;
// 4. Apply safety clamps: ≥1, ≤ max_model_experts, ≤ UINT32_MAX
if (cache_experts == 0) cache_experts = 1;
The non-routed bytes parameter is critical—it accounts for all GPU memory not occupied by routed expert weights: context buffers, KV cache storage, pre-fill scratch space, and alignment padding. Subtracting this overhead prevents out-of-memory errors when the cache is populated.
Computing Effective Cache Size
The final output, ds4_ssd_cache_plan, contains two key fields:
| Field | Meaning |
|---|---|
cache_experts |
Number of experts that can reside in GPU memory |
effective_cache_bytes |
cache_experts * per_expert_bytes (actual allocated bytes) |
Note that effective_cache_bytes may be slightly less than cache_bytes due to integer division in the expert count calculation.
Complete C Implementation Example
#include "ds4_ssd.h"
#include "ds4_gpu.h"
// Gather inputs from model loading and context setup
uint64_t recommended = ds4_gpu_recommended_working_set_size();
uint64_t non_routed = ds4_context_memory_estimate(ctx, batch_size).total_bytes;
uint64_t per_expert = model->expert_size_bytes; // from GGUF metadata
ds4_ssd_cache_plan plan;
bool success = ds4_ssd_auto_cache_plan(
recommended,
non_routed,
per_expert,
model->max_total_experts, // 0 for no model-level cap
&plan
);
if (success) {
printf("SSD cache: %u experts, %.2f GiB effective\n",
plan.cache_experts,
plan.effective_cache_bytes / (1024.0 * 1024.0 * 1024.0));
}
This pattern matches the initialization sequence in the main ds4 inference path.
Debugging the Calculation Manually
To verify the engine's math, reproduce the arithmetic directly:
uint64_t pct = ds4_ssd_auto_cache_percent(); // 80 or env override
uint64_t target = recommended * pct / 100ULL;
uint64_t available = (target > non_routed) ? target - non_routed : 0;
uint64_t experts = available / per_expert;
if (experts == 0) experts = 1; // matches ds4_ssd_auto_cache_plan clamp
printf("Manual calc: %llu experts\n", (unsigned long long)experts);
Discrepancies between manual and engine output typically indicate incorrect non_routed_bytes estimation.
Key Source Files
| File | Lines | Purpose |
|---|---|---|
ds4_ssd.c |
80-106 | Environment variable parsing and percentage validation |
ds4_ssd.c |
108-135 | Full cache plan computation |
ds4_gpu.h |
— | Declaration of working set query interface |
rocm/ds4_rocm_current_api_compat.cuh |
44-53 | ROCm backend implementation |
ds4_metal.m |
95-99 | Metal backend implementation |
ds4_cuda.cu |
54-63 | CUDA multi-GPU aggregation |
Summary
- Backend query: Each platform reports recommended working set via
ds4_gpu_recommended_working_set_size() - Percentage control: 80% default, overridable via
DS4_SSD_AUTO_CACHE_PCT(50-95 range) - Budget formula:
(recommended * pct / 100) - non_routed_bytes = cache_bytes - Expert conversion:
cache_bytes / per_expert_bytes, clamped to valid range - Output:
ds4_ssd_cache_planwith resident expert count and effective byte allocation
Frequently Asked Questions
What happens if non-routed bytes exceeds the target model size?
The calculation yields cache_bytes = 0, which ds4_ssd_auto_cache_plan() promotes to 1 expert minimum. This guarantees at least one expert can stream, though performance degrades. The engine logs a warning when this clamp activates.
Why 80% instead of 100% of the working set recommendation?
The 20% headroom accommodates dynamic allocations during inference: temporary activation buffers,CUDA/ROCm driver overhead, and memory fragmentation. Lower percentages increase safety margins; higher percentages (up to 95%) trade stability for cache hit rate.
How does multi-GPU CUDA handle working set calculations?
The CUDA backend in ds4_cuda.cu aggregates total VRAM across all GPUs as total_b * n, then reports this sum as the working set. The cache planning logic treats this as a single pool—expert distribution across GPUs occurs later in the routing layer.
Can I disable automatic cache sizing and specify expert count manually?
Yes. The automatic planner in ds4_ssd.c only activates when SSD streaming is enabled without explicit cache configuration. Set DS4_SSD_AUTO_CACHE_PCT outside 50-95 to force the 80% default, or bypass entirely by providing explicit cache parameters through the ds4 CLI or API.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →