# How the 80% Backend Working Set Calculation Works for Automatic SSD Streaming Cache Sizing in ds4

> Understand the 80% backend working set calculation for automatic SSD streaming cache sizing in ds4. Learn how ds4 optimizes cache allocation for performance.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: internals
- Published: 2026-08-04

---

**The ds4 inference engine automatically sizes its SSD streaming expert cache by taking 80% of the GPU backend's recommended working set, subtracting non-routed memory overhead, and converting the remainder into a resident expert count.**

The `antirez/ds4` repository implements an intelligent cache sizing mechanism that balances GPU memory pressure against SSD streaming performance. Rather than requiring manual tuning, the system queries each backend (CUDA, ROCm, Metal) for its optimal working set, then derives a cache budget through a deterministic percentage-based calculation.

## Backend Working Set Discovery

Before any percentage can be applied, ds4 must determine how much GPU memory the backend recommends for model execution. This value varies by platform and is surfaced through `ds4_gpu_recommended_working_set_size()`:

- **ROCm**: Queries `hipMemGetInfo` for total device memory via `cudaMemGetInfo` compatibility layer in `rocm/ds4_rocm_current_api_compat.cuh` (lines 44-53)
- **Metal**: Reads the `recommendedMaxWorkingSetSize` property in `ds4_metal.m` (lines 95-99)
- **CUDA**: Returns aggregate VRAM across all available GPUs (`total_b * n`) in `ds4_cuda.cu` (lines 54-63)

Each implementation returns a `uint64_t` representing the total device memory the backend considers usable for the model.

## The 80% Default and `DS4_SSD_AUTO_CACHE_PCT` Override

The percentage applied to the working set recommendation is controlled by `ds4_ssd_auto_cache_percent()` in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 80-106). The logic follows this priority:

1. Check for `DS4_SSD_AUTO_CACHE_PCT` environment variable
2. Validate the value falls within 50-95 (inclusive)
3. Fall back to **80** if unset or invalid, issuing a warning for out-of-range values

```bash

# Override the default before launching ds4

export DS4_SSD_AUTO_CACHE_PCT=90
./ds4-cli --model model.gguf --ssd-streaming

```

The 80% default represents a balance between cache hit rate and headroom for dynamic allocations during inference.

## The Four-Step Cache Budget Calculation

`ds4_ssd_auto_cache_plan()` in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) (lines 108-135) executes the complete calculation:

```c
// 1. Apply percentage to working set recommendation
uint64_t model_target_bytes = recommended_bytes * pct / 100;

// 2. Subtract non-routed memory (context buffers, KV cache, pre-fill)
uint64_t cache_bytes = (model_target_bytes > non_routed_bytes) ?
                       model_target_bytes - non_routed_bytes : 0;

// 3. Convert bytes to expert count
uint64_t cache_experts = cache_bytes / per_expert_bytes;

// 4. Apply safety clamps: ≥1, ≤ max_model_experts, ≤ UINT32_MAX
if (cache_experts == 0) cache_experts = 1;

```

The **non-routed bytes** parameter is critical—it accounts for all GPU memory not occupied by routed expert weights: context buffers, KV cache storage, pre-fill scratch space, and alignment padding. Subtracting this overhead prevents out-of-memory errors when the cache is populated.

## Computing Effective Cache Size

The final output, `ds4_ssd_cache_plan`, contains two key fields:

| Field | Meaning |
|-------|---------|
| `cache_experts` | Number of experts that can reside in GPU memory |
| `effective_cache_bytes` | `cache_experts * per_expert_bytes` (actual allocated bytes) |

Note that `effective_cache_bytes` may be slightly less than `cache_bytes` due to integer division in the expert count calculation.

## Complete C Implementation Example

```c
#include "ds4_ssd.h"
#include "ds4_gpu.h"

// Gather inputs from model loading and context setup
uint64_t recommended = ds4_gpu_recommended_working_set_size();
uint64_t non_routed = ds4_context_memory_estimate(ctx, batch_size).total_bytes;
uint64_t per_expert = model->expert_size_bytes;  // from GGUF metadata

ds4_ssd_cache_plan plan;
bool success = ds4_ssd_auto_cache_plan(
    recommended,
    non_routed,
    per_expert,
    model->max_total_experts,  // 0 for no model-level cap
    &plan
);

if (success) {
    printf("SSD cache: %u experts, %.2f GiB effective\n",
           plan.cache_experts,
           plan.effective_cache_bytes / (1024.0 * 1024.0 * 1024.0));
}

```

This pattern matches the initialization sequence in the main ds4 inference path.

## Debugging the Calculation Manually

To verify the engine's math, reproduce the arithmetic directly:

```c
uint64_t pct = ds4_ssd_auto_cache_percent();        // 80 or env override
uint64_t target = recommended * pct / 100ULL;
uint64_t available = (target > non_routed) ? target - non_routed : 0;
uint64_t experts = available / per_expert;
if (experts == 0) experts = 1;  // matches ds4_ssd_auto_cache_plan clamp

printf("Manual calc: %llu experts\n", (unsigned long long)experts);

```

Discrepancies between manual and engine output typically indicate incorrect `non_routed_bytes` estimation.

## Key Source Files

| File | Lines | Purpose |
|------|-------|---------|
| [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) | 80-106 | Environment variable parsing and percentage validation |
| [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) | 108-135 | Full cache plan computation |
| [`ds4_gpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu.h) | — | Declaration of working set query interface |
| `rocm/ds4_rocm_current_api_compat.cuh` | 44-53 | ROCm backend implementation |
| `ds4_metal.m` | 95-99 | Metal backend implementation |
| `ds4_cuda.cu` | 54-63 | CUDA multi-GPU aggregation |

## Summary

- **Backend query**: Each platform reports recommended working set via `ds4_gpu_recommended_working_set_size()`
- **Percentage control**: 80% default, overridable via `DS4_SSD_AUTO_CACHE_PCT` (50-95 range)
- **Budget formula**: `(recommended * pct / 100) - non_routed_bytes = cache_bytes`
- **Expert conversion**: `cache_bytes / per_expert_bytes`, clamped to valid range
- **Output**: `ds4_ssd_cache_plan` with resident expert count and effective byte allocation

## Frequently Asked Questions

### What happens if non-routed bytes exceeds the target model size?

The calculation yields `cache_bytes = 0`, which `ds4_ssd_auto_cache_plan()` promotes to 1 expert minimum. This guarantees at least one expert can stream, though performance degrades. The engine logs a warning when this clamp activates.

### Why 80% instead of 100% of the working set recommendation?

The 20% headroom accommodates dynamic allocations during inference: temporary activation buffers,CUDA/ROCm driver overhead, and memory fragmentation. Lower percentages increase safety margins; higher percentages (up to 95%) trade stability for cache hit rate.

### How does multi-GPU CUDA handle working set calculations?

The CUDA backend in `ds4_cuda.cu` aggregates total VRAM across all GPUs as `total_b * n`, then reports this sum as the working set. The cache planning logic treats this as a single pool—expert distribution across GPUs occurs later in the routing layer.

### Can I disable automatic cache sizing and specify expert count manually?

Yes. The automatic planner in [`ds4_ssd.c`](https://github.com/antirez/ds4/blob/main/ds4_ssd.c) only activates when SSD streaming is enabled without explicit cache configuration. Set `DS4_SSD_AUTO_CACHE_PCT` outside 50-95 to force the 80% default, or bypass entirely by providing explicit cache parameters through the ds4 CLI or API.