# How KV Cache Management Works in ds4 for Long Context: Disk-Based Checkpointing Explained

> Discover how ds4 manages KV cache for long context using disk-based checkpointing. Resume long conversations efficiently without recomputing. Learn about sparse checkpoints and prefix matching.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-07

---

**ds4 implements a disk-based KV cache that stores attention state checkpoints to disk, enabling resumption of very long conversations without recomputing the entire prompt history through sparse "continued" checkpoints and intelligent prefix matching.**

The `ds4` inference engine solves the computational challenge of long-context language models by persisting **key-value (KV) cache** states to disk. This sophisticated **KV cache management** system creates retrievable checkpoints that allow the model to resume generation from any point in a conversation history, eliminating the need to reprocess thousands of tokens on every turn.

## Core Architecture and Initialization

The cache management system centers on the `ds4_kvstore` structure implemented in [`ds4_kvstore.c`](https://github.com/antirez/ds4/blob/main/ds4_kvstore.c). When initializing the system, `ds4_kvstore_open` (lines 00909-00912) establishes a dedicated directory for cache files, enforces a configurable byte budget, and immediately triggers eviction to ensure the cache complies with storage limits from startup.

Every checkpoint file begins with a 20-byte header containing the magic bytes `KVC`, a version identifier, and metadata describing the checkpoint state. The function `ds4_kvstore_fill_header` (lines 00393-00415) populates these fields, including **quantization bits**, **model ID**, **token count**, **hit counter**, and **timestamps**, ensuring each file is self-describing and traceable.

## Storing Checkpoints for Long Contexts

### Boundary Alignment and Trimming

To optimize cache efficiency, ds4 does not store arbitrary token positions. Instead, it applies **boundary trimming** and **alignment** using the constants `KV_CACHE_DEFAULT_BOUNDARY_TRIM_TOKENS` and `KV_CACHE_DEFAULT_BOUNDARY_ALIGN_TOKENS` (lines 00333-00341). This ensures checkpoints begin on clean token boundaries that align with the backend’s pre-fill chunk size, maximizing computational reuse when loading the cache.

### Continued Checkpoints and Sparse Waypoints

For very long contexts, ds4 implements **continued checkpoints** that create sparse "waypoints" throughout the conversation. The function `ds4_kvstore_continued_store_target` (lines 01241-01248) determines whether to emit a new checkpoint based on configurable step intervals (defaulting to every 10,000 tokens). A checkpoint is written only when the live token count exceeds the previous stored length and aligns with the step multiple (lines 00730-00738).

The main storage routine `ds4_kvstore_store_live_prefix_text` (lines 01135-01155) hashes the stable text prefix of the prompt and atomically writes the KV state payload. To prevent corruption, the system writes to a temporary file (`*.tmp.*`) before renaming it to the final name (lines 01055-01071), ensuring that partially written files never appear in the cache directory.

## Loading and Prefix Matching

### Text Prefix Hash Matching

When resuming a conversation, `ds4_kvstore_try_load_text` (lines 01223-01241) searches for a stored entry whose **text prefix** matches the beginning of the incoming prompt. The helper `ds4_kvstore_find_text_prefix` (lines 01290-01212) uses `ds4_kvstore_sha1_bytes_hex` (lines 00321-00328) to compute hashes and locate the longest matching prefix that satisfies model, quantization, and context size constraints.

### Prompt Reconstruction

Upon finding a match, the system deserializes the payload using `ds4_session_load_payload` (line 01277) and reconstructs the full prompt by concatenating the exact token history from the cache with the remaining suffix of the new prompt. This reconstruction happens in `ds4_kvstore_build_prompt_from_exact_prefix_and_text_suffix` (lines 00885-00897), enabling seamless continuation from the cached state.

## Eviction Policy and Cache Management

### Budget Enforcement and Eviction Scoring

The cache operates within a strict **megabyte budget** (`budget_mb`). When new files would exceed this limit, `ds4_kvstore_evict` (lines 00612-00671) selects victims using a sophisticated scoring algorithm implemented in `ds4_kvstore_entry_eviction_score` (lines 00632-00658). The score incorporates:

- **Exponentially decaying hit counts** based on `DS4_KVSTORE_HIT_HALF_LIFE_SECONDS`
- **Anchor checkpoint protection** via `KV_CACHE_ANCHOR_REASON_SCORE_FACTOR` for cold starts, evictions, and shutdowns
- **Continued checkpoint supersession logic** in `kv_cache_incoming_supersedes_continued` (lines 00504-00524), which protects checkpoints that serve as prefixes to incoming requests

### Hit Tracking and Decay

Each successful cache load invokes `ds4_kvstore_touch_file` (lines 00885-00902) to increment the hit counter and update the `last_used` timestamp. This tracking ensures frequently accessed checkpoints resist eviction, while stale entries naturally age out through the half-life decay mechanism.

## Practical Implementation Examples

### Initializing the KV Store

```c
ds4_kvstore kv;
ds4_kvstore_options opt = ds4_kvstore_default_options();

bool ok = ds4_kvstore_open(&kv,
                           "/var/ds4/kvcache",
                           1024,              // 1 GiB budget
                           false,             // strict quantization matching
                           opt,
                           "ds4-server",
                           my_log_fn,
                           NULL);

```

This creates the cache directory, sets a 1 GiB budget, and runs initial eviction to clean stale files.

### Conditionally Storing Continued Checkpoints

```c
int target = ds4_kvstore_continued_store_target(&kv, current_token_count);
if (target > 0) {
    ds4_kvstore_store_live_prefix(&kv,
                                  engine,
                                  session,
                                  tokens,
                                  target,
                                  "continued",
                                  NULL,
                                  NULL, 0);
}

```

The store operation only executes when the token count justifies a new waypoint, preventing excessive disk writes during short generations.

### Loading Cached State

```c
int loaded = ds4_kvstore_try_load_text(&kv,
                                        engine,
                                        session,
                                        prompt_text,
                                        &effective_prompt,
                                        NULL,
                                        NULL,
                                        false);
if (loaded > 0) {
    // KV cache restored; effective_prompt contains full token history
}

```

## Summary

- **ds4** maintains a **disk-based KV cache** in [`ds4_kvstore.c`](https://github.com/antirez/ds4/blob/main/ds4_kvstore.c) that persists attention states between sessions.
- **Sparse continued checkpoints** create resumable waypoints at configurable token intervals, minimizing storage while supporting arbitrary context lengths.
- **Prefix hash matching** enables rapid retrieval of compatible cache entries based on text content rather than just metadata.
- **Atomic file operations** ensure cache integrity by preventing partially written files from entering the store.
- **Budget-aware eviction** with decaying hit scores automatically manages disk usage while preserving high-value anchor checkpoints.

## Frequently Asked Questions

### How does ds4 prevent cache corruption during unexpected shutdowns?

The system writes all checkpoint data to temporary files with the extension `*.tmp.*` and performs an atomic rename only after the write completes successfully. This pattern, implemented in lines 01055-01071 of [`ds4_kvstore.c`](https://github.com/antirez/ds4/blob/main/ds4_kvstore.c), ensures that incomplete writes never result in valid cache entries that could corrupt the model state on load.

### What distinguishes anchor checkpoints from continued checkpoints?

**Anchor checkpoints** (marked with reasons `cold`, `evict`, or `shutdown`) receive a score multiplier that makes them harder to evict, as they represent stable session boundaries. **Continued checkpoints** serve as intermediate waypoints within long contexts and receive protection only when they act as strict prefixes to incoming prompts, as determined by the supersession logic in `kv_cache_incoming_supersedes_continued`.

### How does the eviction score balance hit frequency against recency?

The scoring algorithm in `ds4_kvstore_entry_eviction_score` calculates **effective hits** by exponentially decaying raw hit counts based on `DS4_KVSTORE_HIT_HALF_LIFE_SECONDS`. This decay mechanism ensures that frequently used but recently inactive checkpoints can be evicted, while currently hot entries remain protected regardless of absolute hit count.

### Can the checkpoint interval be adjusted for different context lengths?

Yes. The continued checkpoint step size is configurable through the `kv_cache_continued_step` parameter (referenced in lines 00730-00738). By adjusting this interval, operators can trade off between storage consumption and recomputation granularity—smaller steps reduce recomputation at the cost of more disk writes, while larger steps conserve space but may require processing more tokens on cache misses.