How KV Cache Management Works in ds4 for Long Context: Disk-Based Checkpointing Explained

ds4 implements a disk-based KV cache that stores attention state checkpoints to disk, enabling resumption of very long conversations without recomputing the entire prompt history through sparse "continued" checkpoints and intelligent prefix matching.

The ds4 inference engine solves the computational challenge of long-context language models by persisting key-value (KV) cache states to disk. This sophisticated KV cache management system creates retrievable checkpoints that allow the model to resume generation from any point in a conversation history, eliminating the need to reprocess thousands of tokens on every turn.

Core Architecture and Initialization

The cache management system centers on the ds4_kvstore structure implemented in ds4_kvstore.c. When initializing the system, ds4_kvstore_open (lines 00909-00912) establishes a dedicated directory for cache files, enforces a configurable byte budget, and immediately triggers eviction to ensure the cache complies with storage limits from startup.

Every checkpoint file begins with a 20-byte header containing the magic bytes KVC, a version identifier, and metadata describing the checkpoint state. The function ds4_kvstore_fill_header (lines 00393-00415) populates these fields, including quantization bits, model ID, token count, hit counter, and timestamps, ensuring each file is self-describing and traceable.

Storing Checkpoints for Long Contexts

Boundary Alignment and Trimming

To optimize cache efficiency, ds4 does not store arbitrary token positions. Instead, it applies boundary trimming and alignment using the constants KV_CACHE_DEFAULT_BOUNDARY_TRIM_TOKENS and KV_CACHE_DEFAULT_BOUNDARY_ALIGN_TOKENS (lines 00333-00341). This ensures checkpoints begin on clean token boundaries that align with the backend’s pre-fill chunk size, maximizing computational reuse when loading the cache.

Continued Checkpoints and Sparse Waypoints

For very long contexts, ds4 implements continued checkpoints that create sparse "waypoints" throughout the conversation. The function ds4_kvstore_continued_store_target (lines 01241-01248) determines whether to emit a new checkpoint based on configurable step intervals (defaulting to every 10,000 tokens). A checkpoint is written only when the live token count exceeds the previous stored length and aligns with the step multiple (lines 00730-00738).

The main storage routine ds4_kvstore_store_live_prefix_text (lines 01135-01155) hashes the stable text prefix of the prompt and atomically writes the KV state payload. To prevent corruption, the system writes to a temporary file (*.tmp.*) before renaming it to the final name (lines 01055-01071), ensuring that partially written files never appear in the cache directory.

Loading and Prefix Matching

Text Prefix Hash Matching

When resuming a conversation, ds4_kvstore_try_load_text (lines 01223-01241) searches for a stored entry whose text prefix matches the beginning of the incoming prompt. The helper ds4_kvstore_find_text_prefix (lines 01290-01212) uses ds4_kvstore_sha1_bytes_hex (lines 00321-00328) to compute hashes and locate the longest matching prefix that satisfies model, quantization, and context size constraints.

Prompt Reconstruction

Upon finding a match, the system deserializes the payload using ds4_session_load_payload (line 01277) and reconstructs the full prompt by concatenating the exact token history from the cache with the remaining suffix of the new prompt. This reconstruction happens in ds4_kvstore_build_prompt_from_exact_prefix_and_text_suffix (lines 00885-00897), enabling seamless continuation from the cached state.

Eviction Policy and Cache Management

Budget Enforcement and Eviction Scoring

The cache operates within a strict megabyte budget (budget_mb). When new files would exceed this limit, ds4_kvstore_evict (lines 00612-00671) selects victims using a sophisticated scoring algorithm implemented in ds4_kvstore_entry_eviction_score (lines 00632-00658). The score incorporates:

  • Exponentially decaying hit counts based on DS4_KVSTORE_HIT_HALF_LIFE_SECONDS
  • Anchor checkpoint protection via KV_CACHE_ANCHOR_REASON_SCORE_FACTOR for cold starts, evictions, and shutdowns
  • Continued checkpoint supersession logic in kv_cache_incoming_supersedes_continued (lines 00504-00524), which protects checkpoints that serve as prefixes to incoming requests

Hit Tracking and Decay

Each successful cache load invokes ds4_kvstore_touch_file (lines 00885-00902) to increment the hit counter and update the last_used timestamp. This tracking ensures frequently accessed checkpoints resist eviction, while stale entries naturally age out through the half-life decay mechanism.

Practical Implementation Examples

Initializing the KV Store

ds4_kvstore kv;
ds4_kvstore_options opt = ds4_kvstore_default_options();

bool ok = ds4_kvstore_open(&kv,
                           "/var/ds4/kvcache",
                           1024,              // 1 GiB budget
                           false,             // strict quantization matching
                           opt,
                           "ds4-server",
                           my_log_fn,
                           NULL);

This creates the cache directory, sets a 1 GiB budget, and runs initial eviction to clean stale files.

Conditionally Storing Continued Checkpoints

int target = ds4_kvstore_continued_store_target(&kv, current_token_count);
if (target > 0) {
    ds4_kvstore_store_live_prefix(&kv,
                                  engine,
                                  session,
                                  tokens,
                                  target,
                                  "continued",
                                  NULL,
                                  NULL, 0);
}

The store operation only executes when the token count justifies a new waypoint, preventing excessive disk writes during short generations.

Loading Cached State

int loaded = ds4_kvstore_try_load_text(&kv,
                                        engine,
                                        session,
                                        prompt_text,
                                        &effective_prompt,
                                        NULL,
                                        NULL,
                                        false);
if (loaded > 0) {
    // KV cache restored; effective_prompt contains full token history
}

Summary

  • ds4 maintains a disk-based KV cache in ds4_kvstore.c that persists attention states between sessions.
  • Sparse continued checkpoints create resumable waypoints at configurable token intervals, minimizing storage while supporting arbitrary context lengths.
  • Prefix hash matching enables rapid retrieval of compatible cache entries based on text content rather than just metadata.
  • Atomic file operations ensure cache integrity by preventing partially written files from entering the store.
  • Budget-aware eviction with decaying hit scores automatically manages disk usage while preserving high-value anchor checkpoints.

Frequently Asked Questions

How does ds4 prevent cache corruption during unexpected shutdowns?

The system writes all checkpoint data to temporary files with the extension *.tmp.* and performs an atomic rename only after the write completes successfully. This pattern, implemented in lines 01055-01071 of ds4_kvstore.c, ensures that incomplete writes never result in valid cache entries that could corrupt the model state on load.

What distinguishes anchor checkpoints from continued checkpoints?

Anchor checkpoints (marked with reasons cold, evict, or shutdown) receive a score multiplier that makes them harder to evict, as they represent stable session boundaries. Continued checkpoints serve as intermediate waypoints within long contexts and receive protection only when they act as strict prefixes to incoming prompts, as determined by the supersession logic in kv_cache_incoming_supersedes_continued.

How does the eviction score balance hit frequency against recency?

The scoring algorithm in ds4_kvstore_entry_eviction_score calculates effective hits by exponentially decaying raw hit counts based on DS4_KVSTORE_HIT_HALF_LIFE_SECONDS. This decay mechanism ensures that frequently used but recently inactive checkpoints can be evicted, while currently hot entries remain protected regardless of absolute hit count.

Can the checkpoint interval be adjusted for different context lengths?

Yes. The continued checkpoint step size is configurable through the kv_cache_continued_step parameter (referenced in lines 00730-00738). By adjusting this interval, operators can trade off between storage consumption and recomputation granularity—smaller steps reduce recomputation at the cost of more disk writes, while larger steps conserve space but may require processing more tokens on cache misses.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →