How to Configure KV Cache Allocation and Management in vLLM
To configure KV cache allocation and management in vLLM, set the CacheConfig parameters—such as block_size, gpu_memory_utilization, and cache_dtype—via CLI arguments or the Python LLM constructor, which the system converts into a KVCacheConfig to drive tensor allocation in the worker.
The vLLM inference engine stores intermediate key-value tensors in a high-performance KV cache to accelerate transformer generation. Understanding how to configure KV cache allocation and management in vLLM allows you to tune memory utilization, enable KV-sharing optimizations, and prevent out-of-memory errors during long-context inference. The configuration flows from user-facing settings in vllm/config/cache.py down to low-level allocation routines in vllm/v1/worker/.
Core KV-Cache Data Structures
The geometry and grouping of the cache are defined by specification classes in vllm/v1/kv_cache_interface.py, while user-facing controls live in vllm/config/cache.py.
| File | Class | Purpose |
|---|---|---|
vllm/v1/kv_cache_interface.py |
KVCacheSpec |
Base class describing the geometry of a single KV cache block (block size, page size). |
vllm/v1/kv_cache_interface.py |
AttentionSpec |
Concrete specification for standard attention (num KV heads, head size, dtype). |
vllm/v1/kv_cache_interface.py |
FullAttentionSpec, MLAAttentionSpec, ChunkedLocalAttentionSpec, SlidingWindowSpec, MambaSpec |
Variants for different attention types (full, sliding-window, Mamba). |
vllm/v1/kv_cache_interface.py |
KVCacheTensor |
Records byte size of a KV tensor and which layers share it. |
vllm/v1/kv_cache_interface.py |
KVCacheGroupSpec |
Groups model layers that share the same block table for KV-sharing. |
vllm/v1/kv_cache_interface.py |
KVCacheConfig |
Top-level config containing number of blocks and lists of KVCacheTensor and KVCacheGroupSpec objects. |
vllm/config/cache.py |
CacheConfig |
Pydantic model exposing user-facing settings parsed from CLI or environment. |
Key configuration fields in CacheConfig include:
block_size– Tokens per contiguous cache block (must be a power of two ≤ 32 on CUDA).gpu_memory_utilization– Fraction of total GPU memory vLLM may reserve for the KV cache.cache_dtype– Storage dtype for KV tensors (bfloat16,fp8, etc.).kv_sharing_fast_prefill– Boolean flag enabling a prefill optimization when KV-sharing is active.kv_cache_memory_bytes– Optional manual override of the total KV cache size in bytes.
Allocation Workflow
The system translates high-level memory settings into physical GPU buffers through a multi-stage pipeline orchestrated by the worker and model runner.
From User Configuration to KVCacheConfig
CacheConfig is parsed from command-line flags or the LLM constructor arguments. Platform.check_and_update_config() validates and finalizes the concrete block_size. In vllm/v1/worker/worker_base.py, the resulting CacheConfig is stored in self.cache_config and passed to the model runner to initiate allocation.
Profiling and Block Count Calculation
In vllm/v1/worker/gpu_worker.py, the method initialize_from_config calls ensure_kv_transfer_initialized followed by self.model_runner.initialize_kv_cache(kv_cache_config). The model runner computes the number of blocks required for each group by dividing the requested total bytes (derived from kv_cache_memory_bytes or the gpu_memory_utilization calculation) by KVCacheSpec.page_size_bytes.
Uniform vs. Non-Uniform Layout
The system detects whether all layers share an identical layout via KVConnectorModelRunnerMixin.use_uniform_kv_cache(), which returns True for uniform specs consolidated by UniformTypeKVCacheSpecs.
- Uniform: All layers share the same layout (single attention group). Allocation is handled by
allocate_uniform_kv_caches. - Non-Uniform: Each group may have a different layout; the runner falls back to per-group allocation paths such as
allocate_kv_cache_tensorsingpu_model_runner.py.
Uniform Allocation via allocate_uniform_kv_caches
For uniform layouts, the mixin class in vllm/v1/worker/kv_connector_model_runner_mixin.py performs a contiguous allocation:
# vllm/v1/worker/kv_connector_model_runner_mixin.py
kv_caches, cross_layers_kv_cache, attn_backend = \
KVConnectorModelRunnerMixin.allocate_uniform_kv_caches(
kv_cache_config, attn_groups, cache_dtype, device,
kernel_block_sizes)
The routine executes the following steps:
- Verifies all
KVCacheTensorobjects have identical sizes. - Derives
num_blocks = tensor_size // page_size. - Computes a kernel-aligned block count (
kernel_num_blocks). - Requests the raw KV shape from the attention backend via
attn_backend.get_kv_cache_shape. - Prepends a
num_layersdimension, permutes according to the backend’s stride order, and allocates a single contiguous buffer (cross_layers_kv_cache). - Slices the buffer per-layer and populates the
kv_cachesdictionary for layer-wise access.
Binding Tensors to the Forward Context
After allocation, bind_kv_cache in vllm/v1/worker/utils.py registers the tensors with the model’s attention modules:
# vllm/v1/worker/utils.py
bind_kv_cache(kv_caches, forward_context, runner_kv_caches, num_attn_module)
This function builds a layer-ordered list runner_kv_caches and associates each tensor with its corresponding Attention object in the forward context, ensuring the model can read and write KV data during generation.
KV-Sharing and Fast-Prefill
The utility add_kv_sharing_layers_to_kv_cache_groups (also in vllm/v1/worker/utils.py) mutates KVCacheGroupSpec objects so that multiple logical layers reference the same physical block table. When kv_sharing_fast_prefill=True, the prefill code path skips allocating KV slots for shared layers, reducing memory pressure during the initial prompt processing phase.
Sliding-Window and Mamba Special Cases
Specialized attention types alter the allocation math:
- Sliding-Window:
SlidingWindowSpecinkv_cache_interface.pyallocates additional blocks (+1block overhead) to cover the sliding window length, calculated inmax_memory_usage_bytes. - Mamba:
MambaSpecusespage_size_bytesderived from convolution shapes and optionally adds speculative blocks vianum_speculative_blocks. Themamba_cache_modeparameter (e.g.,"align") further adjusts the multiplier used in memory calculations (see lines 94‑101 ofkv_cache_interface.py).
Practical Configuration Examples
Minimal CLI Configuration
Allocate 90 % of available GPU memory and allow vLLM to select an optimal block size:
vllm serve model_path --gpu-memory-utilization 0.9
Explicit Block Size and KV-Sharing
Force a 16-token block size and enable fast-prefill optimization for encoder-decoder architectures:
from vllm import SamplingParams, LLM
llm = LLM(
model="meta-llama/Llama-2-13b-hf",
block_size=16,
kv_sharing_fast_prefill=True,
kv_cache_memory_bytes=30 * 1024**3, # 30 GiB manual override
)
sampling_params = SamplingParams(temperature=0.7, max_tokens=128)
outputs = llm.generate(prompt="Explain quantum tunneling.", sampling_params=sampling_params)
Sliding-Window for Long-Context Models
Configure a sliding-window attention cache for models like Longformer:
from vllm import LLM
llm = LLM(
model="mosaicml/longformer-base-4096",
block_size=32,
sliding_window=4096, # window size in tokens
gpu_memory_utilization=0.8,
)
Mamba Cache Customization
Optimize cache layout for State Space Models by aligning cache updates to block boundaries:
from vllm import LLM
llm = LLM(
model="state-spaces/mamba-2.8b",
block_size=64,
mamba_cache_mode="align", # cache only at block boundaries
mamba_page_size_padded=8192, # optional manual page-size override
)
Key Source Files
| Aspect | File | Significance |
|---|---|---|
| KV cache spec definitions | vllm/v1/kv_cache_interface.py |
Central data model for all KV layouts including SlidingWindowSpec and MambaSpec. |
| User-facing cache settings | vllm/config/cache.py |
Pydantic CacheConfig parsed from CLI and environment variables. |
| Uniform allocation routine | vllm/v1/worker/kv_connector_model_runner_mixin.py |
Implements allocate_uniform_kv_caches for contiguous buffer layout. |
| KV binding and sharing | vllm/v1/worker/utils.py |
Contains bind_kv_cache and add_kv_sharing_layers_to_kv_cache_groups. |
| GPU worker initialization | vllm/v1/worker/gpu_worker.py |
Drives memory-request logic and calls initialize_kv_cache. |
| Model runner orchestration | vllm/v1/worker/gpu_model_runner.py |
Invokes allocation utilities and stores resulting tensors. |
Summary
- Entry point: Configure KV cache behavior through
CacheConfigfields (block_size,gpu_memory_utilization,kv_cache_memory_bytes) in the CLI or Python API. - Internal representation: The system converts
CacheConfigintoKVCacheConfigobjects that specify tensor geometry and layer grouping. - Allocation path: Uniform layouts use
allocate_uniform_kv_cachesinkv_connector_model_runner_mixin.pyto create a single contiguous buffer sliced per layer; non-uniform layouts use per-group allocation. - Binding:
bind_kv_cacheinutils.pyattaches allocated tensors to the model’s forward context. - Optimizations: Enable
kv_sharing_fast_prefillto reduce memory during prefill when layers share KV tensors, and use specialized specs (SlidingWindowSpec,MambaSpec) for non-standard attention mechanisms.
Frequently Asked Questions
How do I manually set the total KV cache size instead of using gpu_memory_utilization?
Pass the kv_cache_memory_bytes argument to the LLM constructor or set it in your configuration. According to the source code in vllm/config/cache.py, this value overrides the automatic calculation derived from gpu_memory_utilization, allowing you to specify an exact byte count (e.g., 30 * 1024**3 for 30 GiB).
What is the difference between uniform and non-uniform KV cache allocation in vLLM?
Uniform allocation occurs when all model layers share identical KV cache specifications, detected by use_uniform_kv_cache() returning True. In this case, allocate_uniform_kv_caches creates one large contiguous buffer sliced per layer. Non-uniform allocation handles heterogeneous layer configurations (e.g., mixed attention types) by allocating separate tensors per KVCacheGroupSpec in gpu_model_runner.py.
How does KV-sharing reduce memory usage, and how do I enable it?
KV-sharing allows multiple transformer layers to reference the same physical block table, reducing the total number of allocated blocks. Set kv_sharing_fast_prefill=True in CacheConfig to enable the optimization; the prefill stage then skips slot allocation for shared layers, as implemented in add_kv_sharing_layers_to_kv_cache_groups within vllm/v1/worker/utils.py.
When should I use SlidingWindowSpec versus FullAttentionSpec?
Use SlidingWindowSpec (automatically selected when sliding_window is configured) for long-context models that employ sliding-window attention to limit memory growth; it allocates a fixed window size plus one additional block. Use FullAttentionSpec for standard transformers requiring global attention over the entire sequence, which is the default when no window constraints are specified.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →