How Prefix Caching Works in vLLM: Configuration and Implementation Guide

Prefix caching in vLLM automatically shares KV-cache blocks across requests with identical token prefixes, reducing compute overhead during prefill phases and configurable via CLI flags or the CacheConfig API.

The vLLM inference engine implements an optimization called Automatic Prefix Caching (APC) to eliminate redundant computation when processing prompts that share common beginnings. This mechanism, defined in the vllm-project/vLLM repository, allows the engine to reuse pre-computed key-value (KV) cache blocks across different requests, significantly accelerating throughput for workloads like multi-turn conversations or batch prompts with shared system instructions.

What Is Prefix Caching in vLLM?

Prefix caching (APC) is a memory management optimization that identifies and shares KV-cache blocks containing identical token sequences across concurrent or subsequent requests. When a new request arrives with a token prefix matching a previously processed sequence, vLLM attaches the request to existing cached blocks rather than recomputing the attention states for those tokens.

The system operates through a hash-based block pool that maps token sequences to physical KV blocks. As implemented in vllm/v1/core/block_pool.py, the BlockPool class maintains a cached_block_hash_to_block dictionary that enables O(1) lookups for reusable cache entries.

How Prefix Caching Works Under the Hood

Block Hashing and the BlockPool

At the core of the implementation lies the BlockPool class in vllm/v1/core/block_pool.py (lines 44-68). This component stores the mapping between block hashes and physical block objects. When the KV cache manager allocates blocks for a prefill operation, it computes a cryptographic hash of the block contents and stores it in cached_block_hash_to_block.

The KVCacheManager class in vllm/v1/core/kv_cache_manager.py (line 24) initializes this pool with caching enabled based on the CacheConfig settings. It acts as the intermediary between the scheduler and the low-level block storage.

Scheduler Integration

The scheduler determines cache eligibility during each scheduling step. In vllm/v1/core/sched/scheduler.py (line 224), the constructor receives the enable_caching flag from the configuration and initializes the KVCacheManager accordingly.

During scheduling, the system calls KVCacheManager.get_num_common_prefix_blocks() to compute the longest common prefix between new requests and cached sequences. If the computed hash matches an entry in the block pool, the scheduler allocates the request to the existing blocks, skipping the prefill computation for those tokens.

Cache Lookup Flow

The lookup process follows this sequence:

  1. A request enters the scheduling queue with its token sequence.
  2. The scheduler queries the cache manager for the number of blocks sharing the request's prefix.
  3. The cache manager hashes the prefix tokens and checks against cached_block_hash_to_block.
  4. On a match, the request references the cached blocks; on a miss, new blocks are allocated and hashed.

Configuring Prefix Caching in vLLM

Configuration occurs through CacheConfig defined in vllm/config/cache.py (line 76), which exposes the enable_prefix_caching boolean and prefix_caching_hash_algo string options.

CLI Configuration

The engine argument parser in vllm/engine/arg_utils.py (lines 440-445) exposes the following flags:


# Enable prefix caching (default behavior)

vllm serve meta-llama/Llama-2-7b-hf --enable-prefix-caching

# Explicitly disable

vllm serve meta-llama/Llama-2-7b-hf --disable-prefix-caching

Python API Configuration

For programmatic control, instantiate VllmConfig with a custom CacheConfig:

from vllm import VLLM, VllmConfig, CacheConfig

config = VllmConfig(
    cache_config=CacheConfig(
        enable_prefix_caching=True,
        prefix_caching_hash_algo="xxhash_cbor"
    )
)

engine = VLLM(vllm_config=config)

Hash Algorithm Selection

vLLM supports multiple hashing algorithms for block identification:

  • sha256 (default): Cryptographically secure, no additional dependencies.
  • sha256_cbor: CBOR-encoded SHA256.
  • xxhash: Fast non-cryptographic hash (requires pip install xxhash).
  • xxhash_cbor: CBOR-encoded xxhash.

Select via CLI:

vllm serve meta-llama/Llama-2-7b-hf --prefix-caching-hash-algo xxhash

Managing the Cache at Runtime

vLLM provides mechanisms to invalidate the prefix cache without restarting the server. This is essential after model weight updates (e.g., post-RLHF fine-tuning) or when benchmarking cold-start performance.

The BlockPool.reset_prefix_cache() method clears all stored hashes and frees associated blocks, verifying that only the null block remains allocated before clearing cached_block_hash_to_block.

Programmatic reset:

engine.reset_prefix_cache(reset_running_requests=True)

HTTP endpoint (defined in vllm/entrypoints/serve/cache/api_router.py, lines 21-38):

curl -X POST "http://localhost:8000/v1/reset_prefix_cache?reset_external=true"

Monitoring Prefix Cache Performance

The engine exposes Prometheus metrics to evaluate cache effectiveness. Key metrics include vllm_prefix_cache_hits and vllm_prefix_cache_queries, available at the /metrics endpoint.

Calculate hit rate with:

rate(vllm_prefix_cache_hits[5m]) / rate(vllm_prefix_cache_queries[5m])

High hit rates indicate effective prefix sharing, typically observed in chat applications with consistent system prompts or batch inference with shared prefixes.

Summary

  • Prefix caching in vLLM shares KV-cache blocks across requests with identical token prefixes, eliminating redundant prefill computation.
  • The CacheConfig class in vllm/config/cache.py controls activation via enable_prefix_caching and algorithm selection via prefix_caching_hash_algo.
  • The BlockPool class manages the hash-to-block mapping in vllm/v1/core/block_pool.py, while the scheduler in vllm/v1/core/sched/scheduler.py determines cache eligibility.
  • Configure through CLI flags (--enable-prefix-caching, --prefix-caching-hash-algo) or the Python API (CacheConfig).
  • Reset the cache at runtime using engine.reset_prefix_cache() or the HTTP POST endpoint to /v1/reset_prefix_cache.
  • Monitor performance using the vllm_prefix_cache_hits and vllm_prefix_cache_queries Prometheus metrics.

Frequently Asked Questions

What is the default hash algorithm for vLLM prefix caching?

The default algorithm is sha256, which provides cryptographic security without requiring additional dependencies. For performance-critical workloads where cryptographic security is unnecessary, you can switch to xxhash or xxhash_cbor by installing the xxhash package and setting --prefix-caching-hash-algo xxhash.

How do I clear the prefix cache without restarting the server?

Use the reset_prefix_cache method on the engine instance with reset_running_requests=True, or send a POST request to the /v1/reset_prefix_cache endpoint with reset_external=true. This clears the cached_block_hash_to_block mapping in the BlockPool and frees associated blocks while the server continues running.

Does enabling prefix caching increase memory usage?

Prefix caching can reduce overall memory pressure by sharing identical KV blocks across multiple requests, but it requires additional metadata storage for the hash mappings. The memory overhead is typically negligible compared to the savings from avoiding redundant prefill computations.

When should I disable prefix caching?

Disable prefix caching when processing requests with highly unique prefixes where cache hits are unlikely, or when deterministic benchmarking requires cold-start behavior for each request. Use the --disable-prefix-caching CLI flag or set enable_prefix_caching=False in CacheConfig for these scenarios.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →