# Understanding KTransformers' 3-Layer GPU-CPU-Disk Prefix Cache Architecture

> Explore KTransformers' 3-layer GPU-CPU-Disk prefix cache architecture. Accelerate LLM inference with sub-millisecond hot prefix access and efficient cold data management.

- Repository: [kvcache.ai/ktransformers](https://github.com/kvcache-ai/ktransformers)
- Tags: architecture
- Published: 2026-07-26

---

**KTransformers accelerates LLM inference by tiering the KV cache across GPU VRAM, CPU RAM, and persistent disk storage, enabling sub-millisecond access to hot prefixes while keeping cold data ready for rapid reloading.**

The `kvcache-ai/ktransformers` repository implements a hierarchical memory system that solves the memory bottleneck in large language model inference. This 3-layer GPU-CPU-Disk prefix cache architecture stores key-value tensors across three distinct hardware tiers, allowing models to reuse computed attention states across multiple requests without recomputation.

## The Three-Layer Memory Hierarchy

The architecture divides the KV cache into three distinct temperature tiers based on access frequency. This design enables **prefix-cache reuse**, where previously computed attention states persist across independent inference calls.

### GPU Layer (Hot Cache)

The GPU layer resides in on-device VRAM and holds the most frequently accessed portion of the prefix KV cache. This provides the lowest latency access during token generation. According to the source code in [`archive/ktransformers/models/custom_cache.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/models/custom_cache.py), the implementation uses `torch._dynamo.mark_static_address` to pin these tensors in GPU memory, eliminating allocation overhead during decoding steps.

### CPU Layer (Warm Buffer)

When the KV cache exceeds GPU capacity, the CPU layer acts as a mid-tier buffer in main system RAM. The size is controlled via the `cpu_memory_size_GB` configuration parameter defined in [`archive/ktransformers/server/config/config.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/server/config/config.py). This layer bridges the speed gap between GPU VRAM and disk storage, holding KV entries that are likely to be needed soon but do not fit in the hot cache.

### Disk Layer (Cold Storage)

The disk layer persists KV cache pages to SSD or other block storage via the `disk_path` configuration option. This cold storage preserves computed prefix states across process restarts and allows multiple independent inference calls to share the same pre-computed attention matrices. When a cached prefix is requested, the system reads these files back into the CPU buffer before paging to GPU.

## Core Implementation in the KVC2 Subsystem

The architecture centers on the **KVC2** subsystem, which manages the page table translation between logical token positions and physical storage locations.

The `KVC2StaticCache` class in [`archive/ktransformers/models/custom_cache.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/models/custom_cache.py) serves as the primary interface. It implements a static cache object that maps a unified page-table across all three layers. Concrete implementations like `KVC2Qwen3Cache` and `KGQACache` extend this base for specific model families such as DeepSeek-V3 and Qwen-3.

Each cache operates on a **page size** (default 256 tokens), translating logical positions into physical pages distributed across GPU, CPU, or Disk. The page table tracks which layer holds each segment, enabling efficient data movement through `get_page_table` lookups.

## Configuration and Setup

To enable the three-layer cache, configure the YAML file with disk path and memory limits:

```yaml

# ./ktransformers/configs/config.yaml

attn:
  page_size: 16                     # KV page size

  chunk_size: 256
kvc2:
  gpu_only: false                   # False → enable CPU+Disk layers

  utilization_percentage: 1.0
  cpu_memory_size_GB: 500           # Reserve ~500 GB RAM for CPU KV cache

  disk_path: /mnt/data/kvc          # Directory where KV pages are persisted

```

The `gpu_only: false` setting is critical to activate the full hierarchy. The system loads these parameters at runtime through the configuration module.

## Runtime Workflow and Python API

Instantiate the static cache in Python using the model configuration:

```python
from transformers import PretrainedConfig
from ktransformers.models.custom_cache import KVC2StaticCache

cfg = PretrainedConfig.from_pretrained("deepseek-v3")
max_batch = 8
static_cache = KVC2StaticCache(config=cfg,
                               max_batch_size=max_batch,
                               page_size=256,
                               dtype=torch.bfloat16,
                               device="cuda:0")

```

During inference, the runtime executes three phases:

1. **Load**: Persisted KV files are read from `disk_path` into the CPU buffer
2. **Page**: The `get_page_table` method translates token positions to physical pages, moving required slices into GPU memory
3. **Update**: The `update` method modifies GPU cache in-place using page indices and offsets

```python

# Generate page table for current batch

page_idx, page_offset = static_cache.get_page_table(mini_batch,
                                                    bsz_tensors=None,
                                                    is_prefill=True)

# Update cache with new KV tensors

cache_kwargs = {"page_idx": page_idx, "page_offset": page_offset}
static_cache.update(combined_tensor, layer_idx=3, cache_kwargs=cache_kwargs)

# Reset between sessions

static_cache.reset()

```

## Summary

- The **3-layer GPU-CPU-Disk prefix cache architecture** tiers KV storage by access frequency: GPU for hot data, CPU for warm buffers, and Disk for cold persistence
- **KVC2StaticCache** in [`archive/ktransformers/models/custom_cache.py`](https://github.com/kvcache-ai/ktransformers/blob/main/archive/ktransformers/models/custom_cache.py) implements the unified page table managing all three layers with a default page size of 256 tokens
- Configuration through `cpu_memory_size_GB` and `disk_path` controls resource allocation, while `gpu_only: false` activates the full hierarchy
- Prefix reuse eliminates redundant computation of attention states, drastically reducing latency for long prompts and repeated requests

## Frequently Asked Questions

### How does the page table map tokens across GPU, CPU, and Disk?

The page table translates logical token positions into physical pages using a default size of 256 tokens per page. Each entry tracks whether the page resides in GPU VRAM, CPU RAM, or on Disk. When `get_page_table` is called, it returns `page_idx` and `page_offset` values that the `update` method uses to locate and modify the correct memory region without searching the entire hierarchy.

### What is the default page size and can it be changed?

The default page size is 256 tokens, defined in the concrete cache implementations like `KVC2Qwen3Cache` and `KGQACache`. You can override this via the `page_size` parameter in the YAML configuration under the `attn` section or when instantiating `KVC2StaticCache` directly in Python.

### How do I verify my disk cache directory has sufficient space?

Use the diagnostic tool in [`kt-kernel/python/cli/commands/doctor.py`](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/python/cli/commands/doctor.py) to check available disk space against your configured `disk_path`. This utility reports whether the storage volume can accommodate the expected KV cache size based on your model configuration and sequence lengths.

### Why does the GPU layer use torch._dynamo.mark_static_address?

This PyTorch decorator pins the GPU cache tensors to fixed memory addresses, preventing PyTorch's allocator from moving them during the forward pass. This optimization eliminates allocation overhead and cache invalidation during the decoding phase, ensuring consistent low-latency access to hot prefix data.