# How to Configure KV Cache Allocation and Management in vLLM

> Master KV cache allocation and management in vLLM. Learn to set CacheConfig parameters like block_size and gpu_memory_utilization for optimized tensor allocation via CLI or Python.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**To configure KV cache allocation and management in vLLM, set the `CacheConfig` parameters—such as `block_size`, `gpu_memory_utilization`, and `cache_dtype`—via CLI arguments or the Python `LLM` constructor, which the system converts into a `KVCacheConfig` to drive tensor allocation in the worker.**

The vLLM inference engine stores intermediate key-value tensors in a high-performance KV cache to accelerate transformer generation. Understanding how to configure KV cache allocation and management in vLLM allows you to tune memory utilization, enable KV-sharing optimizations, and prevent out-of-memory errors during long-context inference. The configuration flows from user-facing settings in [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py) down to low-level allocation routines in `vllm/v1/worker/`.

## Core KV-Cache Data Structures

The geometry and grouping of the cache are defined by specification classes in [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py), while user-facing controls live in [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py).

| File | Class | Purpose |
|------|-------|---------|
| [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py) | `KVCacheSpec` | Base class describing the geometry of a single KV cache block (block size, page size). |
| [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py) | `AttentionSpec` | Concrete specification for standard attention (num KV heads, head size, dtype). |
| [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py) | `FullAttentionSpec`, `MLAAttentionSpec`, `ChunkedLocalAttentionSpec`, `SlidingWindowSpec`, `MambaSpec` | Variants for different attention types (full, sliding-window, Mamba). |
| [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py) | `KVCacheTensor` | Records byte size of a KV tensor and which layers share it. |
| [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py) | `KVCacheGroupSpec` | Groups model layers that share the same block table for KV-sharing. |
| [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py) | `KVCacheConfig` | Top-level config containing number of blocks and lists of `KVCacheTensor` and `KVCacheGroupSpec` objects. |
| [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py) | `CacheConfig` | Pydantic model exposing user-facing settings parsed from CLI or environment. |

**Key configuration fields** in `CacheConfig` include:

- `block_size` – Tokens per contiguous cache block (must be a power of two ≤ 32 on CUDA).
- `gpu_memory_utilization` – Fraction of total GPU memory vLLM may reserve for the KV cache.
- `cache_dtype` – Storage dtype for KV tensors (`bfloat16`, `fp8`, etc.).
- `kv_sharing_fast_prefill` – Boolean flag enabling a prefill optimization when KV-sharing is active.
- `kv_cache_memory_bytes` – Optional manual override of the total KV cache size in bytes.

## Allocation Workflow

The system translates high-level memory settings into physical GPU buffers through a multi-stage pipeline orchestrated by the worker and model runner.

### From User Configuration to KVCacheConfig

`CacheConfig` is parsed from command-line flags or the `LLM` constructor arguments. `Platform.check_and_update_config()` validates and finalizes the concrete `block_size`. In [`vllm/v1/worker/worker_base.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/worker_base.py), the resulting `CacheConfig` is stored in `self.cache_config` and passed to the model runner to initiate allocation.

### Profiling and Block Count Calculation

In [`vllm/v1/worker/gpu_worker.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py), the method `initialize_from_config` calls `ensure_kv_transfer_initialized` followed by `self.model_runner.initialize_kv_cache(kv_cache_config)`. The model runner computes the number of blocks required for each group by dividing the requested total bytes (derived from `kv_cache_memory_bytes` or the `gpu_memory_utilization` calculation) by `KVCacheSpec.page_size_bytes`.

### Uniform vs. Non-Uniform Layout

The system detects whether all layers share an identical layout via `KVConnectorModelRunnerMixin.use_uniform_kv_cache()`, which returns `True` for uniform specs consolidated by `UniformTypeKVCacheSpecs`.

- **Uniform**: All layers share the same layout (single attention group). Allocation is handled by `allocate_uniform_kv_caches`.
- **Non-Uniform**: Each group may have a different layout; the runner falls back to per-group allocation paths such as `allocate_kv_cache_tensors` in [`gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/gpu_model_runner.py).

### Uniform Allocation via `allocate_uniform_kv_caches`

For uniform layouts, the mixin class in [`vllm/v1/worker/kv_connector_model_runner_mixin.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/kv_connector_model_runner_mixin.py) performs a contiguous allocation:

```python

# vllm/v1/worker/kv_connector_model_runner_mixin.py

kv_caches, cross_layers_kv_cache, attn_backend = \
    KVConnectorModelRunnerMixin.allocate_uniform_kv_caches(
        kv_cache_config, attn_groups, cache_dtype, device,
        kernel_block_sizes)

```

The routine executes the following steps:

1. Verifies all `KVCacheTensor` objects have identical sizes.
2. Derives `num_blocks = tensor_size // page_size`.
3. Computes a kernel-aligned block count (`kernel_num_blocks`).
4. Requests the raw KV shape from the attention backend via `attn_backend.get_kv_cache_shape`.
5. Prepends a `num_layers` dimension, permutes according to the backend’s stride order, and allocates a single contiguous buffer (`cross_layers_kv_cache`).
6. Slices the buffer per-layer and populates the `kv_caches` dictionary for layer-wise access.

### Binding Tensors to the Forward Context

After allocation, `bind_kv_cache` in [`vllm/v1/worker/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/utils.py) registers the tensors with the model’s attention modules:

```python

# vllm/v1/worker/utils.py

bind_kv_cache(kv_caches, forward_context, runner_kv_caches, num_attn_module)

```

This function builds a layer-ordered list `runner_kv_caches` and associates each tensor with its corresponding `Attention` object in the forward context, ensuring the model can read and write KV data during generation.

### KV-Sharing and Fast-Prefill

The utility `add_kv_sharing_layers_to_kv_cache_groups` (also in [`vllm/v1/worker/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/utils.py)) mutates `KVCacheGroupSpec` objects so that multiple logical layers reference the same physical block table. When `kv_sharing_fast_prefill=True`, the prefill code path skips allocating KV slots for shared layers, reducing memory pressure during the initial prompt processing phase.

### Sliding-Window and Mamba Special Cases

Specialized attention types alter the allocation math:

- **Sliding-Window**: `SlidingWindowSpec` in [`kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/kv_cache_interface.py) allocates additional blocks (`+1` block overhead) to cover the sliding window length, calculated in `max_memory_usage_bytes`.
- **Mamba**: `MambaSpec` uses `page_size_bytes` derived from convolution shapes and optionally adds speculative blocks via `num_speculative_blocks`. The `mamba_cache_mode` parameter (e.g., `"align"`) further adjusts the multiplier used in memory calculations (see lines 94‑101 of [`kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/kv_cache_interface.py)).

## Practical Configuration Examples

### Minimal CLI Configuration

Allocate 90 % of available GPU memory and allow vLLM to select an optimal block size:

```bash
vllm serve model_path --gpu-memory-utilization 0.9

```

### Explicit Block Size and KV-Sharing

Force a 16-token block size and enable fast-prefill optimization for encoder-decoder architectures:

```python
from vllm import SamplingParams, LLM

llm = LLM(
    model="meta-llama/Llama-2-13b-hf",
    block_size=16,
    kv_sharing_fast_prefill=True,
    kv_cache_memory_bytes=30 * 1024**3,  # 30 GiB manual override

)

sampling_params = SamplingParams(temperature=0.7, max_tokens=128)
outputs = llm.generate(prompt="Explain quantum tunneling.", sampling_params=sampling_params)

```

### Sliding-Window for Long-Context Models

Configure a sliding-window attention cache for models like Longformer:

```python
from vllm import LLM

llm = LLM(
    model="mosaicml/longformer-base-4096",
    block_size=32,
    sliding_window=4096,          # window size in tokens

    gpu_memory_utilization=0.8,
)

```

### Mamba Cache Customization

Optimize cache layout for State Space Models by aligning cache updates to block boundaries:

```python
from vllm import LLM

llm = LLM(
    model="state-spaces/mamba-2.8b",
    block_size=64,
    mamba_cache_mode="align",     # cache only at block boundaries

    mamba_page_size_padded=8192,  # optional manual page-size override

)

```

## Key Source Files

| Aspect | File | Significance |
|--------|------|------------|
| KV cache spec definitions | [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py) | Central data model for all KV layouts including `SlidingWindowSpec` and `MambaSpec`. |
| User-facing cache settings | [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py) | Pydantic `CacheConfig` parsed from CLI and environment variables. |
| Uniform allocation routine | [`vllm/v1/worker/kv_connector_model_runner_mixin.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/kv_connector_model_runner_mixin.py) | Implements `allocate_uniform_kv_caches` for contiguous buffer layout. |
| KV binding and sharing | [`vllm/v1/worker/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/utils.py) | Contains `bind_kv_cache` and `add_kv_sharing_layers_to_kv_cache_groups`. |
| GPU worker initialization | [`vllm/v1/worker/gpu_worker.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py) | Drives memory-request logic and calls `initialize_kv_cache`. |
| Model runner orchestration | [`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py) | Invokes allocation utilities and stores resulting tensors. |

## Summary

- **Entry point**: Configure KV cache behavior through `CacheConfig` fields (`block_size`, `gpu_memory_utilization`, `kv_cache_memory_bytes`) in the CLI or Python API.
- **Internal representation**: The system converts `CacheConfig` into `KVCacheConfig` objects that specify tensor geometry and layer grouping.
- **Allocation path**: Uniform layouts use `allocate_uniform_kv_caches` in [`kv_connector_model_runner_mixin.py`](https://github.com/vllm-project/vllm/blob/main/kv_connector_model_runner_mixin.py) to create a single contiguous buffer sliced per layer; non-uniform layouts use per-group allocation.
- **Binding**: `bind_kv_cache` in [`utils.py`](https://github.com/vllm-project/vllm/blob/main/utils.py) attaches allocated tensors to the model’s forward context.
- **Optimizations**: Enable `kv_sharing_fast_prefill` to reduce memory during prefill when layers share KV tensors, and use specialized specs (`SlidingWindowSpec`, `MambaSpec`) for non-standard attention mechanisms.

## Frequently Asked Questions

### How do I manually set the total KV cache size instead of using gpu_memory_utilization?

Pass the `kv_cache_memory_bytes` argument to the `LLM` constructor or set it in your configuration. According to the source code in [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py), this value overrides the automatic calculation derived from `gpu_memory_utilization`, allowing you to specify an exact byte count (e.g., `30 * 1024**3` for 30 GiB).

### What is the difference between uniform and non-uniform KV cache allocation in vLLM?

Uniform allocation occurs when all model layers share identical KV cache specifications, detected by `use_uniform_kv_cache()` returning `True`. In this case, `allocate_uniform_kv_caches` creates one large contiguous buffer sliced per layer. Non-uniform allocation handles heterogeneous layer configurations (e.g., mixed attention types) by allocating separate tensors per `KVCacheGroupSpec` in [`gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/gpu_model_runner.py).

### How does KV-sharing reduce memory usage, and how do I enable it?

KV-sharing allows multiple transformer layers to reference the same physical block table, reducing the total number of allocated blocks. Set `kv_sharing_fast_prefill=True` in `CacheConfig` to enable the optimization; the prefill stage then skips slot allocation for shared layers, as implemented in `add_kv_sharing_layers_to_kv_cache_groups` within [`vllm/v1/worker/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/utils.py).

### When should I use `SlidingWindowSpec` versus `FullAttentionSpec`?

Use `SlidingWindowSpec` (automatically selected when `sliding_window` is configured) for long-context models that employ sliding-window attention to limit memory growth; it allocates a fixed window size plus one additional block. Use `FullAttentionSpec` for standard transformers requiring global attention over the entire sequence, which is the default when no window constraints are specified.