# How to Persist and Restore KV Cache Sessions to Disk in ds4

> Learn to persist and restore KV cache sessions to disk in ds4. Save and load raw binary data to resume generation without recomputing attention history.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**To persist and restore KV cache sessions to disk in ds4, serialize the GPU-resident KV tensors using `ds4_gpu_store_raw_kv_tensor()` or `ds4_gpu_store_raw_kv_batch_tensor()`, write the raw binary data to a file, and later reload them using `ds4_gpu_load_raw_kv_tensor()` or `ds4_gpu_load_raw_kv_batch_tensor()` to resume generation without recomputing attention history.**

The ds4 inference engine by antirez stores transformer key-value (KV) caches in GPU memory to accelerate autoregressive decoding. When building long-running applications or resumable chat sessions, you must persist and restore KV cache sessions to disk in ds4 to maintain context across process restarts or system recoveries.

## Core APIs for KV Cache Serialization

The ds4 library exposes two complementary function pairs for serializing KV data. The `*_tensor()` variants handle single-token checkpointing, while the `*_batch_tensor()` variants optimize for contiguous token blocks.

### Storing Single Tokens with `ds4_gpu_store_raw_kv_tensor()`

Use `ds4_gpu_store_raw_kv_tensor()` when persisting after individual generation steps. This function, implemented in [`ds4_kvstore.c`](https://github.com/antirez/ds4/blob/main/ds4_kvstore.c), copies raw float values from a specific layer's GPU KV tensor into a pre-allocated host buffer. It accepts the target tensor pointer, host buffer, byte size, starting row offset, and head dimension size (`DS4_N_HEAD_DIM`).

### Storing Token Batches with `ds4_gpu_store_raw_kv_batch_tensor()`

For checkpointing after processing multiple tokens simultaneously, use `ds4_gpu_store_raw_kv_batch_tensor()`. This batch variant minimizes API overhead when persisting long contexts by handling contiguous memory blocks in a single operation.

### Loading Functions for Restoration

The restoration process uses `ds4_gpu_load_raw_kv_tensor()` and `ds4_gpu_load_raw_kv_batch_tensor()`, which read the binary blobs back into host memory and upload them to the GPU tensors. Both functions return `0` on success or a non-zero error code if the transfer fails.

## Low-Level GPU Memory Primitives

Under the hood, the storage helpers rely on `ds4_gpu_tensor_read()` and `ds4_gpu_tensor_write()` defined in [`ds4_gpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu.h) and [`ds4_gpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu.c). These primitives perform the actual DMA transfers between GPU tensors and host buffers, returning the number of bytes transferred.

## Implementation Workflow

The KV cache in ds4 is layer-wise; you must iterate over all layers via `g->layer_kv[layer]` where `g` represents the model state containing `g->n_layers` layers.

### Step 1: Serialize and Save to Disk

Allocate host buffers sized to `kv->size` bytes for each layer, copy data using `ds4_gpu_tensor_read()`, and write to disk:

```c
FILE *fp = fopen("kv_state.bin", "wb");
if (!fp) { perror("open"); exit(1); }

for (int layer = 0; layer < g->n_layers; ++layer) {
    ds4_gpu_tensor *kv = g->layer_kv[layer];
    size_t kv_bytes = kv->size;
    void *host_buf = malloc(kv_bytes);
    if (!host_buf) abort();

    // Pull raw data from GPU
    if (ds4_gpu_tensor_read(kv, 0, host_buf, kv_bytes) != 0) {
        fprintf(stderr, "GPU-to-CPU copy failed\n");
        abort();
    }

    // Write raw blob
    if (fwrite(host_buf, 1, kv_bytes, fp) != kv_bytes) {
        perror("write");
        abort();
    }
    free(host_buf);
}
fclose(fp);

```

### Step 2: Restore from Disk

Read the binary data back and upload to GPU using `ds4_gpu_tensor_write()`:

```c
FILE *rp = fopen("kv_state.bin", "rb");
if (!rp) { perror("open"); exit(1); }

for (int layer = 0; layer < g->n_layers; ++layer) {
    ds4_gpu_tensor *kv = g->layer_kv[layer];
    size_t kv_bytes = kv->size;
    void *host_buf = malloc(kv_bytes);
    if (!host_buf) abort();

    // Read raw blob
    if (fread(host_buf, 1, kv_bytes, rp) != kv_bytes) {
        perror("read");
        abort();
    }

    // Push data back to GPU
    if (ds4_gpu_tensor_write(kv, 0, host_buf, kv_bytes) != 0) {
        fprintf(stderr, "CPU-to-GPU copy failed\n");
        abort();
    }
    free(host_buf);
}
fclose(rp);

```

## Key Considerations for Checkpointing

**Layer-wise Structure:** You must persist all layers in `g->layer_kv` or maintain explicit layer indices, as the restoration process expects tensors to map exactly to their original layer positions.

**Raw Binary Format:** The storage format contains raw `float` or `half` values (depending on quantization) with no headers or metadata. The file is simply a concatenation of layer blobs, making it compact but requiring external tracking of model configuration.

**Model Configuration Consistency:** When restoring, the engine must initialize with identical parameters—same number of layers, head dimension (`DS4_N_HEAD_DIM`), and KV dimensions—because the tensor layout depends on these values. Mismatched configurations will result in memory corruption or generation errors.

**Unit Test Reference:** See `test_metal_store_raw_kv_batch_wrap()` in [`tests/ds4_test.c`](https://github.com/antirez/ds4/blob/main/tests/ds4_test.c) (lines 626-658) for a working demonstration of the batch storage API.

## Summary

- Use `ds4_gpu_store_raw_kv_tensor()` for single-token checkpointing and `ds4_gpu_store_raw_kv_batch_tensor()` for batch operations when you need to persist and restore KV cache sessions to disk in ds4.
- Implement the actual disk I/O using `ds4_gpu_tensor_read()` to serialize from GPU to host, then `fwrite()` to disk, reversing the process with `fread()` and `ds4_gpu_tensor_write()` for restoration.
- Process all layers in `g->layer_kv` iteratively, allocating host buffers sized to each tensor's `kv->size` property.
- Maintain identical model configuration between save and restore operations, as the raw binary format contains no metadata headers.
- Reference the implementation in [`ds4_kvstore.c`](https://github.com/antirez/ds4/blob/main/ds4_kvstore.c) for storage logic and [`ds4_gpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu.c) for low-level memory primitives.

## Frequently Asked Questions

### What file format does ds4 use for KV cache persistence?

ds4 uses a raw binary format containing concatenated `float` or `half` values with no headers or metadata. Each layer's KV tensor is written sequentially to the file, so you must track layer count, head dimensions, and data types externally.

### Can I persist only specific layers of the KV cache?

Yes, you can selectively persist layers by iterating only over the specific indices in `g->layer_kv` that you wish to checkpoint. However, restoration requires writing to the same corresponding layer indices, and partial checkpoints cannot be restored to a full model state without careful index management.

### What happens if I try to restore a KV cache to a different model configuration?

Restoring to a mismatched configuration—such as different layer counts, head dimensions (`DS4_N_HEAD_DIM`), or quantization settings—will cause memory corruption or silent generation errors. The raw binary layout depends entirely on the original tensor shapes, so the engine must initialize with identical parameters before calling `ds4_gpu_load_raw_kv_tensor()`.

### Is there a performance penalty when persisting KV cache to disk?

The primary overhead comes from PCIe transfers via `ds4_gpu_tensor_read()` and disk I/O. For large models, the KV cache can consume gigabytes of memory, so expect save/resume operations to take several seconds depending on storage speed. Use `ds4_gpu_store_raw_kv_batch_tensor()` to minimize API call overhead when persisting long contexts.