How the 80% Backend Working Set Calculation Works for Automatic SSD Streaming Cache Sizing in ds4

The ds4 inference engine automatically sizes its SSD streaming expert cache by taking 80% of the GPU backend's recommended working set, subtracting non-routed memory overhead, and converting the remainder into a resident expert count.

The antirez/ds4 repository implements an intelligent cache sizing mechanism that balances GPU memory pressure against SSD streaming performance. Rather than requiring manual tuning, the system queries each backend (CUDA, ROCm, Metal) for its optimal working set, then derives a cache budget through a deterministic percentage-based calculation.

Backend Working Set Discovery

Before any percentage can be applied, ds4 must determine how much GPU memory the backend recommends for model execution. This value varies by platform and is surfaced through ds4_gpu_recommended_working_set_size():

  • ROCm: Queries hipMemGetInfo for total device memory via cudaMemGetInfo compatibility layer in rocm/ds4_rocm_current_api_compat.cuh (lines 44-53)
  • Metal: Reads the recommendedMaxWorkingSetSize property in ds4_metal.m (lines 95-99)
  • CUDA: Returns aggregate VRAM across all available GPUs (total_b * n) in ds4_cuda.cu (lines 54-63)

Each implementation returns a uint64_t representing the total device memory the backend considers usable for the model.

The 80% Default and DS4_SSD_AUTO_CACHE_PCT Override

The percentage applied to the working set recommendation is controlled by ds4_ssd_auto_cache_percent() in ds4_ssd.c (lines 80-106). The logic follows this priority:

  1. Check for DS4_SSD_AUTO_CACHE_PCT environment variable
  2. Validate the value falls within 50-95 (inclusive)
  3. Fall back to 80 if unset or invalid, issuing a warning for out-of-range values

# Override the default before launching ds4

export DS4_SSD_AUTO_CACHE_PCT=90
./ds4-cli --model model.gguf --ssd-streaming

The 80% default represents a balance between cache hit rate and headroom for dynamic allocations during inference.

The Four-Step Cache Budget Calculation

ds4_ssd_auto_cache_plan() in ds4_ssd.c (lines 108-135) executes the complete calculation:

// 1. Apply percentage to working set recommendation
uint64_t model_target_bytes = recommended_bytes * pct / 100;

// 2. Subtract non-routed memory (context buffers, KV cache, pre-fill)
uint64_t cache_bytes = (model_target_bytes > non_routed_bytes) ?
                       model_target_bytes - non_routed_bytes : 0;

// 3. Convert bytes to expert count
uint64_t cache_experts = cache_bytes / per_expert_bytes;

// 4. Apply safety clamps: ≥1, ≤ max_model_experts, ≤ UINT32_MAX
if (cache_experts == 0) cache_experts = 1;

The non-routed bytes parameter is critical—it accounts for all GPU memory not occupied by routed expert weights: context buffers, KV cache storage, pre-fill scratch space, and alignment padding. Subtracting this overhead prevents out-of-memory errors when the cache is populated.

Computing Effective Cache Size

The final output, ds4_ssd_cache_plan, contains two key fields:

Field Meaning
cache_experts Number of experts that can reside in GPU memory
effective_cache_bytes cache_experts * per_expert_bytes (actual allocated bytes)

Note that effective_cache_bytes may be slightly less than cache_bytes due to integer division in the expert count calculation.

Complete C Implementation Example

#include "ds4_ssd.h"
#include "ds4_gpu.h"

// Gather inputs from model loading and context setup
uint64_t recommended = ds4_gpu_recommended_working_set_size();
uint64_t non_routed = ds4_context_memory_estimate(ctx, batch_size).total_bytes;
uint64_t per_expert = model->expert_size_bytes;  // from GGUF metadata

ds4_ssd_cache_plan plan;
bool success = ds4_ssd_auto_cache_plan(
    recommended,
    non_routed,
    per_expert,
    model->max_total_experts,  // 0 for no model-level cap
    &plan
);

if (success) {
    printf("SSD cache: %u experts, %.2f GiB effective\n",
           plan.cache_experts,
           plan.effective_cache_bytes / (1024.0 * 1024.0 * 1024.0));
}

This pattern matches the initialization sequence in the main ds4 inference path.

Debugging the Calculation Manually

To verify the engine's math, reproduce the arithmetic directly:

uint64_t pct = ds4_ssd_auto_cache_percent();        // 80 or env override
uint64_t target = recommended * pct / 100ULL;
uint64_t available = (target > non_routed) ? target - non_routed : 0;
uint64_t experts = available / per_expert;
if (experts == 0) experts = 1;  // matches ds4_ssd_auto_cache_plan clamp

printf("Manual calc: %llu experts\n", (unsigned long long)experts);

Discrepancies between manual and engine output typically indicate incorrect non_routed_bytes estimation.

Key Source Files

File Lines Purpose
ds4_ssd.c 80-106 Environment variable parsing and percentage validation
ds4_ssd.c 108-135 Full cache plan computation
ds4_gpu.h — Declaration of working set query interface
rocm/ds4_rocm_current_api_compat.cuh 44-53 ROCm backend implementation
ds4_metal.m 95-99 Metal backend implementation
ds4_cuda.cu 54-63 CUDA multi-GPU aggregation

Summary

  • Backend query: Each platform reports recommended working set via ds4_gpu_recommended_working_set_size()
  • Percentage control: 80% default, overridable via DS4_SSD_AUTO_CACHE_PCT (50-95 range)
  • Budget formula: (recommended * pct / 100) - non_routed_bytes = cache_bytes
  • Expert conversion: cache_bytes / per_expert_bytes, clamped to valid range
  • Output: ds4_ssd_cache_plan with resident expert count and effective byte allocation

Frequently Asked Questions

What happens if non-routed bytes exceeds the target model size?

The calculation yields cache_bytes = 0, which ds4_ssd_auto_cache_plan() promotes to 1 expert minimum. This guarantees at least one expert can stream, though performance degrades. The engine logs a warning when this clamp activates.

Why 80% instead of 100% of the working set recommendation?

The 20% headroom accommodates dynamic allocations during inference: temporary activation buffers,CUDA/ROCm driver overhead, and memory fragmentation. Lower percentages increase safety margins; higher percentages (up to 95%) trade stability for cache hit rate.

How does multi-GPU CUDA handle working set calculations?

The CUDA backend in ds4_cuda.cu aggregates total VRAM across all GPUs as total_b * n, then reports this sum as the working set. The cache planning logic treats this as a single pool—expert distribution across GPUs occurs later in the routing layer.

Can I disable automatic cache sizing and specify expert count manually?

Yes. The automatic planner in ds4_ssd.c only activates when SSD streaming is enabled without explicit cache configuration. Set DS4_SSD_AUTO_CACHE_PCT outside 50-95 to force the 80% default, or bypass entirely by providing explicit cache parameters through the ds4 CLI or API.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →