How to Configure KV Cache Allocation and Management in vLLM

To configure KV cache allocation and management in vLLM, set the CacheConfig parameters—such as block_size, gpu_memory_utilization, and cache_dtype—via CLI arguments or the Python LLM constructor, which the system converts into a KVCacheConfig to drive tensor allocation in the worker.

The vLLM inference engine stores intermediate key-value tensors in a high-performance KV cache to accelerate transformer generation. Understanding how to configure KV cache allocation and management in vLLM allows you to tune memory utilization, enable KV-sharing optimizations, and prevent out-of-memory errors during long-context inference. The configuration flows from user-facing settings in vllm/config/cache.py down to low-level allocation routines in vllm/v1/worker/.

Core KV-Cache Data Structures

The geometry and grouping of the cache are defined by specification classes in vllm/v1/kv_cache_interface.py, while user-facing controls live in vllm/config/cache.py.

File Class Purpose
vllm/v1/kv_cache_interface.py KVCacheSpec Base class describing the geometry of a single KV cache block (block size, page size).
vllm/v1/kv_cache_interface.py AttentionSpec Concrete specification for standard attention (num KV heads, head size, dtype).
vllm/v1/kv_cache_interface.py FullAttentionSpec, MLAAttentionSpec, ChunkedLocalAttentionSpec, SlidingWindowSpec, MambaSpec Variants for different attention types (full, sliding-window, Mamba).
vllm/v1/kv_cache_interface.py KVCacheTensor Records byte size of a KV tensor and which layers share it.
vllm/v1/kv_cache_interface.py KVCacheGroupSpec Groups model layers that share the same block table for KV-sharing.
vllm/v1/kv_cache_interface.py KVCacheConfig Top-level config containing number of blocks and lists of KVCacheTensor and KVCacheGroupSpec objects.
vllm/config/cache.py CacheConfig Pydantic model exposing user-facing settings parsed from CLI or environment.

Key configuration fields in CacheConfig include:

  • block_size – Tokens per contiguous cache block (must be a power of two ≤ 32 on CUDA).
  • gpu_memory_utilization – Fraction of total GPU memory vLLM may reserve for the KV cache.
  • cache_dtype – Storage dtype for KV tensors (bfloat16, fp8, etc.).
  • kv_sharing_fast_prefill – Boolean flag enabling a prefill optimization when KV-sharing is active.
  • kv_cache_memory_bytes – Optional manual override of the total KV cache size in bytes.

Allocation Workflow

The system translates high-level memory settings into physical GPU buffers through a multi-stage pipeline orchestrated by the worker and model runner.

From User Configuration to KVCacheConfig

CacheConfig is parsed from command-line flags or the LLM constructor arguments. Platform.check_and_update_config() validates and finalizes the concrete block_size. In vllm/v1/worker/worker_base.py, the resulting CacheConfig is stored in self.cache_config and passed to the model runner to initiate allocation.

Profiling and Block Count Calculation

In vllm/v1/worker/gpu_worker.py, the method initialize_from_config calls ensure_kv_transfer_initialized followed by self.model_runner.initialize_kv_cache(kv_cache_config). The model runner computes the number of blocks required for each group by dividing the requested total bytes (derived from kv_cache_memory_bytes or the gpu_memory_utilization calculation) by KVCacheSpec.page_size_bytes.

Uniform vs. Non-Uniform Layout

The system detects whether all layers share an identical layout via KVConnectorModelRunnerMixin.use_uniform_kv_cache(), which returns True for uniform specs consolidated by UniformTypeKVCacheSpecs.

  • Uniform: All layers share the same layout (single attention group). Allocation is handled by allocate_uniform_kv_caches.
  • Non-Uniform: Each group may have a different layout; the runner falls back to per-group allocation paths such as allocate_kv_cache_tensors in gpu_model_runner.py.

Uniform Allocation via allocate_uniform_kv_caches

For uniform layouts, the mixin class in vllm/v1/worker/kv_connector_model_runner_mixin.py performs a contiguous allocation:


# vllm/v1/worker/kv_connector_model_runner_mixin.py

kv_caches, cross_layers_kv_cache, attn_backend = \
    KVConnectorModelRunnerMixin.allocate_uniform_kv_caches(
        kv_cache_config, attn_groups, cache_dtype, device,
        kernel_block_sizes)

The routine executes the following steps:

  1. Verifies all KVCacheTensor objects have identical sizes.
  2. Derives num_blocks = tensor_size // page_size.
  3. Computes a kernel-aligned block count (kernel_num_blocks).
  4. Requests the raw KV shape from the attention backend via attn_backend.get_kv_cache_shape.
  5. Prepends a num_layers dimension, permutes according to the backend’s stride order, and allocates a single contiguous buffer (cross_layers_kv_cache).
  6. Slices the buffer per-layer and populates the kv_caches dictionary for layer-wise access.

Binding Tensors to the Forward Context

After allocation, bind_kv_cache in vllm/v1/worker/utils.py registers the tensors with the model’s attention modules:


# vllm/v1/worker/utils.py

bind_kv_cache(kv_caches, forward_context, runner_kv_caches, num_attn_module)

This function builds a layer-ordered list runner_kv_caches and associates each tensor with its corresponding Attention object in the forward context, ensuring the model can read and write KV data during generation.

KV-Sharing and Fast-Prefill

The utility add_kv_sharing_layers_to_kv_cache_groups (also in vllm/v1/worker/utils.py) mutates KVCacheGroupSpec objects so that multiple logical layers reference the same physical block table. When kv_sharing_fast_prefill=True, the prefill code path skips allocating KV slots for shared layers, reducing memory pressure during the initial prompt processing phase.

Sliding-Window and Mamba Special Cases

Specialized attention types alter the allocation math:

  • Sliding-Window: SlidingWindowSpec in kv_cache_interface.py allocates additional blocks (+1 block overhead) to cover the sliding window length, calculated in max_memory_usage_bytes.
  • Mamba: MambaSpec uses page_size_bytes derived from convolution shapes and optionally adds speculative blocks via num_speculative_blocks. The mamba_cache_mode parameter (e.g., "align") further adjusts the multiplier used in memory calculations (see lines 94‑101 of kv_cache_interface.py).

Practical Configuration Examples

Minimal CLI Configuration

Allocate 90 % of available GPU memory and allow vLLM to select an optimal block size:

vllm serve model_path --gpu-memory-utilization 0.9

Explicit Block Size and KV-Sharing

Force a 16-token block size and enable fast-prefill optimization for encoder-decoder architectures:

from vllm import SamplingParams, LLM

llm = LLM(
    model="meta-llama/Llama-2-13b-hf",
    block_size=16,
    kv_sharing_fast_prefill=True,
    kv_cache_memory_bytes=30 * 1024**3,  # 30 GiB manual override

)

sampling_params = SamplingParams(temperature=0.7, max_tokens=128)
outputs = llm.generate(prompt="Explain quantum tunneling.", sampling_params=sampling_params)

Sliding-Window for Long-Context Models

Configure a sliding-window attention cache for models like Longformer:

from vllm import LLM

llm = LLM(
    model="mosaicml/longformer-base-4096",
    block_size=32,
    sliding_window=4096,          # window size in tokens

    gpu_memory_utilization=0.8,
)

Mamba Cache Customization

Optimize cache layout for State Space Models by aligning cache updates to block boundaries:

from vllm import LLM

llm = LLM(
    model="state-spaces/mamba-2.8b",
    block_size=64,
    mamba_cache_mode="align",     # cache only at block boundaries

    mamba_page_size_padded=8192,  # optional manual page-size override

)

Key Source Files

Aspect File Significance
KV cache spec definitions vllm/v1/kv_cache_interface.py Central data model for all KV layouts including SlidingWindowSpec and MambaSpec.
User-facing cache settings vllm/config/cache.py Pydantic CacheConfig parsed from CLI and environment variables.
Uniform allocation routine vllm/v1/worker/kv_connector_model_runner_mixin.py Implements allocate_uniform_kv_caches for contiguous buffer layout.
KV binding and sharing vllm/v1/worker/utils.py Contains bind_kv_cache and add_kv_sharing_layers_to_kv_cache_groups.
GPU worker initialization vllm/v1/worker/gpu_worker.py Drives memory-request logic and calls initialize_kv_cache.
Model runner orchestration vllm/v1/worker/gpu_model_runner.py Invokes allocation utilities and stores resulting tensors.

Summary

  • Entry point: Configure KV cache behavior through CacheConfig fields (block_size, gpu_memory_utilization, kv_cache_memory_bytes) in the CLI or Python API.
  • Internal representation: The system converts CacheConfig into KVCacheConfig objects that specify tensor geometry and layer grouping.
  • Allocation path: Uniform layouts use allocate_uniform_kv_caches in kv_connector_model_runner_mixin.py to create a single contiguous buffer sliced per layer; non-uniform layouts use per-group allocation.
  • Binding: bind_kv_cache in utils.py attaches allocated tensors to the model’s forward context.
  • Optimizations: Enable kv_sharing_fast_prefill to reduce memory during prefill when layers share KV tensors, and use specialized specs (SlidingWindowSpec, MambaSpec) for non-standard attention mechanisms.

Frequently Asked Questions

How do I manually set the total KV cache size instead of using gpu_memory_utilization?

Pass the kv_cache_memory_bytes argument to the LLM constructor or set it in your configuration. According to the source code in vllm/config/cache.py, this value overrides the automatic calculation derived from gpu_memory_utilization, allowing you to specify an exact byte count (e.g., 30 * 1024**3 for 30 GiB).

What is the difference between uniform and non-uniform KV cache allocation in vLLM?

Uniform allocation occurs when all model layers share identical KV cache specifications, detected by use_uniform_kv_cache() returning True. In this case, allocate_uniform_kv_caches creates one large contiguous buffer sliced per layer. Non-uniform allocation handles heterogeneous layer configurations (e.g., mixed attention types) by allocating separate tensors per KVCacheGroupSpec in gpu_model_runner.py.

How does KV-sharing reduce memory usage, and how do I enable it?

KV-sharing allows multiple transformer layers to reference the same physical block table, reducing the total number of allocated blocks. Set kv_sharing_fast_prefill=True in CacheConfig to enable the optimization; the prefill stage then skips slot allocation for shared layers, as implemented in add_kv_sharing_layers_to_kv_cache_groups within vllm/v1/worker/utils.py.

When should I use SlidingWindowSpec versus FullAttentionSpec?

Use SlidingWindowSpec (automatically selected when sliding_window is configured) for long-context models that employ sliding-window attention to limit memory growth; it allocates a fixed window size plus one additional block. Use FullAttentionSpec for standard transformers requiring global attention over the entire sequence, which is the default when no window constraints are specified.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →