How to Persist and Restore KV Cache Sessions to Disk in ds4
To persist and restore KV cache sessions to disk in ds4, serialize the GPU-resident KV tensors using ds4_gpu_store_raw_kv_tensor() or ds4_gpu_store_raw_kv_batch_tensor(), write the raw binary data to a file, and later reload them using ds4_gpu_load_raw_kv_tensor() or ds4_gpu_load_raw_kv_batch_tensor() to resume generation without recomputing attention history.
The ds4 inference engine by antirez stores transformer key-value (KV) caches in GPU memory to accelerate autoregressive decoding. When building long-running applications or resumable chat sessions, you must persist and restore KV cache sessions to disk in ds4 to maintain context across process restarts or system recoveries.
Core APIs for KV Cache Serialization
The ds4 library exposes two complementary function pairs for serializing KV data. The *_tensor() variants handle single-token checkpointing, while the *_batch_tensor() variants optimize for contiguous token blocks.
Storing Single Tokens with ds4_gpu_store_raw_kv_tensor()
Use ds4_gpu_store_raw_kv_tensor() when persisting after individual generation steps. This function, implemented in ds4_kvstore.c, copies raw float values from a specific layer's GPU KV tensor into a pre-allocated host buffer. It accepts the target tensor pointer, host buffer, byte size, starting row offset, and head dimension size (DS4_N_HEAD_DIM).
Storing Token Batches with ds4_gpu_store_raw_kv_batch_tensor()
For checkpointing after processing multiple tokens simultaneously, use ds4_gpu_store_raw_kv_batch_tensor(). This batch variant minimizes API overhead when persisting long contexts by handling contiguous memory blocks in a single operation.
Loading Functions for Restoration
The restoration process uses ds4_gpu_load_raw_kv_tensor() and ds4_gpu_load_raw_kv_batch_tensor(), which read the binary blobs back into host memory and upload them to the GPU tensors. Both functions return 0 on success or a non-zero error code if the transfer fails.
Low-Level GPU Memory Primitives
Under the hood, the storage helpers rely on ds4_gpu_tensor_read() and ds4_gpu_tensor_write() defined in ds4_gpu.h and ds4_gpu.c. These primitives perform the actual DMA transfers between GPU tensors and host buffers, returning the number of bytes transferred.
Implementation Workflow
The KV cache in ds4 is layer-wise; you must iterate over all layers via g->layer_kv[layer] where g represents the model state containing g->n_layers layers.
Step 1: Serialize and Save to Disk
Allocate host buffers sized to kv->size bytes for each layer, copy data using ds4_gpu_tensor_read(), and write to disk:
FILE *fp = fopen("kv_state.bin", "wb");
if (!fp) { perror("open"); exit(1); }
for (int layer = 0; layer < g->n_layers; ++layer) {
ds4_gpu_tensor *kv = g->layer_kv[layer];
size_t kv_bytes = kv->size;
void *host_buf = malloc(kv_bytes);
if (!host_buf) abort();
// Pull raw data from GPU
if (ds4_gpu_tensor_read(kv, 0, host_buf, kv_bytes) != 0) {
fprintf(stderr, "GPU-to-CPU copy failed\n");
abort();
}
// Write raw blob
if (fwrite(host_buf, 1, kv_bytes, fp) != kv_bytes) {
perror("write");
abort();
}
free(host_buf);
}
fclose(fp);
Step 2: Restore from Disk
Read the binary data back and upload to GPU using ds4_gpu_tensor_write():
FILE *rp = fopen("kv_state.bin", "rb");
if (!rp) { perror("open"); exit(1); }
for (int layer = 0; layer < g->n_layers; ++layer) {
ds4_gpu_tensor *kv = g->layer_kv[layer];
size_t kv_bytes = kv->size;
void *host_buf = malloc(kv_bytes);
if (!host_buf) abort();
// Read raw blob
if (fread(host_buf, 1, kv_bytes, rp) != kv_bytes) {
perror("read");
abort();
}
// Push data back to GPU
if (ds4_gpu_tensor_write(kv, 0, host_buf, kv_bytes) != 0) {
fprintf(stderr, "CPU-to-GPU copy failed\n");
abort();
}
free(host_buf);
}
fclose(rp);
Key Considerations for Checkpointing
Layer-wise Structure: You must persist all layers in g->layer_kv or maintain explicit layer indices, as the restoration process expects tensors to map exactly to their original layer positions.
Raw Binary Format: The storage format contains raw float or half values (depending on quantization) with no headers or metadata. The file is simply a concatenation of layer blobs, making it compact but requiring external tracking of model configuration.
Model Configuration Consistency: When restoring, the engine must initialize with identical parameters—same number of layers, head dimension (DS4_N_HEAD_DIM), and KV dimensions—because the tensor layout depends on these values. Mismatched configurations will result in memory corruption or generation errors.
Unit Test Reference: See test_metal_store_raw_kv_batch_wrap() in tests/ds4_test.c (lines 626-658) for a working demonstration of the batch storage API.
Summary
- Use
ds4_gpu_store_raw_kv_tensor()for single-token checkpointing andds4_gpu_store_raw_kv_batch_tensor()for batch operations when you need to persist and restore KV cache sessions to disk in ds4. - Implement the actual disk I/O using
ds4_gpu_tensor_read()to serialize from GPU to host, thenfwrite()to disk, reversing the process withfread()andds4_gpu_tensor_write()for restoration. - Process all layers in
g->layer_kviteratively, allocating host buffers sized to each tensor'skv->sizeproperty. - Maintain identical model configuration between save and restore operations, as the raw binary format contains no metadata headers.
- Reference the implementation in
ds4_kvstore.cfor storage logic andds4_gpu.cfor low-level memory primitives.
Frequently Asked Questions
What file format does ds4 use for KV cache persistence?
ds4 uses a raw binary format containing concatenated float or half values with no headers or metadata. Each layer's KV tensor is written sequentially to the file, so you must track layer count, head dimensions, and data types externally.
Can I persist only specific layers of the KV cache?
Yes, you can selectively persist layers by iterating only over the specific indices in g->layer_kv that you wish to checkpoint. However, restoration requires writing to the same corresponding layer indices, and partial checkpoints cannot be restored to a full model state without careful index management.
What happens if I try to restore a KV cache to a different model configuration?
Restoring to a mismatched configuration—such as different layer counts, head dimensions (DS4_N_HEAD_DIM), or quantization settings—will cause memory corruption or silent generation errors. The raw binary layout depends entirely on the original tensor shapes, so the engine must initialize with identical parameters before calling ds4_gpu_load_raw_kv_tensor().
Is there a performance penalty when persisting KV cache to disk?
The primary overhead comes from PCIe transfers via ds4_gpu_tensor_read() and disk I/O. For large models, the KV cache can consume gigabytes of memory, so expect save/resume operations to take several seconds depending on storage speed. Use ds4_gpu_store_raw_kv_batch_tensor() to minimize API call overhead when persisting long contexts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →