# How Prefix Caching Works in vLLM: Configuration and Implementation Guide

> Learn how prefix caching in vLLM optimizes LLM inference by sharing KV-cache blocks. Discover configuration options for reduced compute overhead in prefill.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**Prefix caching in vLLM automatically shares KV-cache blocks across requests with identical token prefixes, reducing compute overhead during prefill phases and configurable via CLI flags or the CacheConfig API.**

The vLLM inference engine implements an optimization called Automatic Prefix Caching (APC) to eliminate redundant computation when processing prompts that share common beginnings. This mechanism, defined in the `vllm-project/vLLM` repository, allows the engine to reuse pre-computed key-value (KV) cache blocks across different requests, significantly accelerating throughput for workloads like multi-turn conversations or batch prompts with shared system instructions.

## What Is Prefix Caching in vLLM?

Prefix caching (APC) is a memory management optimization that identifies and shares KV-cache blocks containing identical token sequences across concurrent or subsequent requests. When a new request arrives with a token prefix matching a previously processed sequence, vLLM attaches the request to existing cached blocks rather than recomputing the attention states for those tokens.

The system operates through a hash-based block pool that maps token sequences to physical KV blocks. As implemented in [`vllm/v1/core/block_pool.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/block_pool.py), the `BlockPool` class maintains a `cached_block_hash_to_block` dictionary that enables O(1) lookups for reusable cache entries.

## How Prefix Caching Works Under the Hood

### Block Hashing and the BlockPool

At the core of the implementation lies the `BlockPool` class in [`vllm/v1/core/block_pool.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/block_pool.py) (lines 44-68). This component stores the mapping between block hashes and physical block objects. When the KV cache manager allocates blocks for a prefill operation, it computes a cryptographic hash of the block contents and stores it in `cached_block_hash_to_block`.

The `KVCacheManager` class in [`vllm/v1/core/kv_cache_manager.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/kv_cache_manager.py) (line 24) initializes this pool with caching enabled based on the `CacheConfig` settings. It acts as the intermediary between the scheduler and the low-level block storage.

### Scheduler Integration

The scheduler determines cache eligibility during each scheduling step. In [`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py) (line 224), the constructor receives the `enable_caching` flag from the configuration and initializes the `KVCacheManager` accordingly.

During scheduling, the system calls `KVCacheManager.get_num_common_prefix_blocks()` to compute the longest common prefix between new requests and cached sequences. If the computed hash matches an entry in the block pool, the scheduler allocates the request to the existing blocks, skipping the prefill computation for those tokens.

### Cache Lookup Flow

The lookup process follows this sequence:

1. A request enters the scheduling queue with its token sequence.
2. The scheduler queries the cache manager for the number of blocks sharing the request's prefix.
3. The cache manager hashes the prefix tokens and checks against `cached_block_hash_to_block`.
4. On a match, the request references the cached blocks; on a miss, new blocks are allocated and hashed.

## Configuring Prefix Caching in vLLM

Configuration occurs through `CacheConfig` defined in [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py) (line 76), which exposes the `enable_prefix_caching` boolean and `prefix_caching_hash_algo` string options.

### CLI Configuration

The engine argument parser in [`vllm/engine/arg_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/engine/arg_utils.py) (lines 440-445) exposes the following flags:

```bash

# Enable prefix caching (default behavior)

vllm serve meta-llama/Llama-2-7b-hf --enable-prefix-caching

# Explicitly disable

vllm serve meta-llama/Llama-2-7b-hf --disable-prefix-caching

```

### Python API Configuration

For programmatic control, instantiate `VllmConfig` with a custom `CacheConfig`:

```python
from vllm import VLLM, VllmConfig, CacheConfig

config = VllmConfig(
    cache_config=CacheConfig(
        enable_prefix_caching=True,
        prefix_caching_hash_algo="xxhash_cbor"
    )
)

engine = VLLM(vllm_config=config)

```

### Hash Algorithm Selection

vLLM supports multiple hashing algorithms for block identification:

- **sha256** (default): Cryptographically secure, no additional dependencies.
- **sha256_cbor**: CBOR-encoded SHA256.
- **xxhash**: Fast non-cryptographic hash (requires `pip install xxhash`).
- **xxhash_cbor**: CBOR-encoded xxhash.

Select via CLI:

```bash
vllm serve meta-llama/Llama-2-7b-hf --prefix-caching-hash-algo xxhash

```

## Managing the Cache at Runtime

vLLM provides mechanisms to invalidate the prefix cache without restarting the server. This is essential after model weight updates (e.g., post-RLHF fine-tuning) or when benchmarking cold-start performance.

The `BlockPool.reset_prefix_cache()` method clears all stored hashes and frees associated blocks, verifying that only the null block remains allocated before clearing `cached_block_hash_to_block`.

Programmatic reset:

```python
engine.reset_prefix_cache(reset_running_requests=True)

```

HTTP endpoint (defined in [`vllm/entrypoints/serve/cache/api_router.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/serve/cache/api_router.py), lines 21-38):

```bash
curl -X POST "http://localhost:8000/v1/reset_prefix_cache?reset_external=true"

```

## Monitoring Prefix Cache Performance

The engine exposes Prometheus metrics to evaluate cache effectiveness. Key metrics include `vllm_prefix_cache_hits` and `vllm_prefix_cache_queries`, available at the `/metrics` endpoint.

Calculate hit rate with:

```bash
rate(vllm_prefix_cache_hits[5m]) / rate(vllm_prefix_cache_queries[5m])

```

High hit rates indicate effective prefix sharing, typically observed in chat applications with consistent system prompts or batch inference with shared prefixes.

## Summary

- Prefix caching in vLLM shares KV-cache blocks across requests with identical token prefixes, eliminating redundant prefill computation.
- The `CacheConfig` class in [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py) controls activation via `enable_prefix_caching` and algorithm selection via `prefix_caching_hash_algo`.
- The `BlockPool` class manages the hash-to-block mapping in [`vllm/v1/core/block_pool.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/block_pool.py), while the scheduler in [`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py) determines cache eligibility.
- Configure through CLI flags (`--enable-prefix-caching`, `--prefix-caching-hash-algo`) or the Python API (`CacheConfig`).
- Reset the cache at runtime using `engine.reset_prefix_cache()` or the HTTP POST endpoint to `/v1/reset_prefix_cache`.
- Monitor performance using the `vllm_prefix_cache_hits` and `vllm_prefix_cache_queries` Prometheus metrics.

## Frequently Asked Questions

### What is the default hash algorithm for vLLM prefix caching?

The default algorithm is **sha256**, which provides cryptographic security without requiring additional dependencies. For performance-critical workloads where cryptographic security is unnecessary, you can switch to **xxhash** or **xxhash_cbor** by installing the `xxhash` package and setting `--prefix-caching-hash-algo xxhash`.

### How do I clear the prefix cache without restarting the server?

Use the `reset_prefix_cache` method on the engine instance with `reset_running_requests=True`, or send a POST request to the `/v1/reset_prefix_cache` endpoint with `reset_external=true`. This clears the `cached_block_hash_to_block` mapping in the `BlockPool` and frees associated blocks while the server continues running.

### Does enabling prefix caching increase memory usage?

Prefix caching can reduce overall memory pressure by sharing identical KV blocks across multiple requests, but it requires additional metadata storage for the hash mappings. The memory overhead is typically negligible compared to the savings from avoiding redundant prefill computations.

### When should I disable prefix caching?

Disable prefix caching when processing requests with highly unique prefixes where cache hits are unlikely, or when deterministic benchmarking requires cold-start behavior for each request. Use the `--disable-prefix-caching` CLI flag or set `enable_prefix_caching=False` in `CacheConfig` for these scenarios.