Understanding KTransformers' 3-Layer GPU-CPU-Disk Prefix Cache Architecture
KTransformers accelerates LLM inference by tiering the KV cache across GPU VRAM, CPU RAM, and persistent disk storage, enabling sub-millisecond access to hot prefixes while keeping cold data ready for rapid reloading.
The kvcache-ai/ktransformers repository implements a hierarchical memory system that solves the memory bottleneck in large language model inference. This 3-layer GPU-CPU-Disk prefix cache architecture stores key-value tensors across three distinct hardware tiers, allowing models to reuse computed attention states across multiple requests without recomputation.
The Three-Layer Memory Hierarchy
The architecture divides the KV cache into three distinct temperature tiers based on access frequency. This design enables prefix-cache reuse, where previously computed attention states persist across independent inference calls.
GPU Layer (Hot Cache)
The GPU layer resides in on-device VRAM and holds the most frequently accessed portion of the prefix KV cache. This provides the lowest latency access during token generation. According to the source code in archive/ktransformers/models/custom_cache.py, the implementation uses torch._dynamo.mark_static_address to pin these tensors in GPU memory, eliminating allocation overhead during decoding steps.
CPU Layer (Warm Buffer)
When the KV cache exceeds GPU capacity, the CPU layer acts as a mid-tier buffer in main system RAM. The size is controlled via the cpu_memory_size_GB configuration parameter defined in archive/ktransformers/server/config/config.py. This layer bridges the speed gap between GPU VRAM and disk storage, holding KV entries that are likely to be needed soon but do not fit in the hot cache.
Disk Layer (Cold Storage)
The disk layer persists KV cache pages to SSD or other block storage via the disk_path configuration option. This cold storage preserves computed prefix states across process restarts and allows multiple independent inference calls to share the same pre-computed attention matrices. When a cached prefix is requested, the system reads these files back into the CPU buffer before paging to GPU.
Core Implementation in the KVC2 Subsystem
The architecture centers on the KVC2 subsystem, which manages the page table translation between logical token positions and physical storage locations.
The KVC2StaticCache class in archive/ktransformers/models/custom_cache.py serves as the primary interface. It implements a static cache object that maps a unified page-table across all three layers. Concrete implementations like KVC2Qwen3Cache and KGQACache extend this base for specific model families such as DeepSeek-V3 and Qwen-3.
Each cache operates on a page size (default 256 tokens), translating logical positions into physical pages distributed across GPU, CPU, or Disk. The page table tracks which layer holds each segment, enabling efficient data movement through get_page_table lookups.
Configuration and Setup
To enable the three-layer cache, configure the YAML file with disk path and memory limits:
# ./ktransformers/configs/config.yaml
attn:
page_size: 16 # KV page size
chunk_size: 256
kvc2:
gpu_only: false # False → enable CPU+Disk layers
utilization_percentage: 1.0
cpu_memory_size_GB: 500 # Reserve ~500 GB RAM for CPU KV cache
disk_path: /mnt/data/kvc # Directory where KV pages are persisted
The gpu_only: false setting is critical to activate the full hierarchy. The system loads these parameters at runtime through the configuration module.
Runtime Workflow and Python API
Instantiate the static cache in Python using the model configuration:
from transformers import PretrainedConfig
from ktransformers.models.custom_cache import KVC2StaticCache
cfg = PretrainedConfig.from_pretrained("deepseek-v3")
max_batch = 8
static_cache = KVC2StaticCache(config=cfg,
max_batch_size=max_batch,
page_size=256,
dtype=torch.bfloat16,
device="cuda:0")
During inference, the runtime executes three phases:
- Load: Persisted KV files are read from
disk_pathinto the CPU buffer - Page: The
get_page_tablemethod translates token positions to physical pages, moving required slices into GPU memory - Update: The
updatemethod modifies GPU cache in-place using page indices and offsets
# Generate page table for current batch
page_idx, page_offset = static_cache.get_page_table(mini_batch,
bsz_tensors=None,
is_prefill=True)
# Update cache with new KV tensors
cache_kwargs = {"page_idx": page_idx, "page_offset": page_offset}
static_cache.update(combined_tensor, layer_idx=3, cache_kwargs=cache_kwargs)
# Reset between sessions
static_cache.reset()
Summary
- The 3-layer GPU-CPU-Disk prefix cache architecture tiers KV storage by access frequency: GPU for hot data, CPU for warm buffers, and Disk for cold persistence
- KVC2StaticCache in
archive/ktransformers/models/custom_cache.pyimplements the unified page table managing all three layers with a default page size of 256 tokens - Configuration through
cpu_memory_size_GBanddisk_pathcontrols resource allocation, whilegpu_only: falseactivates the full hierarchy - Prefix reuse eliminates redundant computation of attention states, drastically reducing latency for long prompts and repeated requests
Frequently Asked Questions
How does the page table map tokens across GPU, CPU, and Disk?
The page table translates logical token positions into physical pages using a default size of 256 tokens per page. Each entry tracks whether the page resides in GPU VRAM, CPU RAM, or on Disk. When get_page_table is called, it returns page_idx and page_offset values that the update method uses to locate and modify the correct memory region without searching the entire hierarchy.
What is the default page size and can it be changed?
The default page size is 256 tokens, defined in the concrete cache implementations like KVC2Qwen3Cache and KGQACache. You can override this via the page_size parameter in the YAML configuration under the attn section or when instantiating KVC2StaticCache directly in Python.
How do I verify my disk cache directory has sufficient space?
Use the diagnostic tool in kt-kernel/python/cli/commands/doctor.py to check available disk space against your configured disk_path. This utility reports whether the storage volume can accommodate the expected KV cache size based on your model configuration and sequence lengths.
Why does the GPU layer use torch._dynamo.mark_static_address?
This PyTorch decorator pins the GPU cache tensors to fixed memory addresses, preventing PyTorch's allocator from moving them during the forward pass. This optimization eliminates allocation overhead and cache invalidation during the decoding phase, ensuring consistent low-latency access to hot prefix data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →