How PagedAttention Works in vLLM to Maximize Memory Efficiency
PagedAttention eliminates GPU memory overhead in vLLM by reinterpreting KV cache tensors as zero-copy views rather than performing expensive data copies, allowing the ROCm and CUDA backends to process attention directly on compact flash-layout buffers.
Efficient key-value (KV) cache management is the bottleneck for high-throughput LLM serving. In the vllm-project/vllm repository, PagedAttention bridges the gap between the memory-efficient flash layout used for storage and the paged layout required by GPU attention kernels, achieving zero-copy tensor translation.
The Memory Layout Challenge: Flash vs. Paged
vLLM stores KV caches in two distinct formats depending on whether the data is being stored or computed upon. The friction between these layouts creates the need for PagedAttention’s view-based translation layer.
Flash Layout: The Canonical Storage Format
The default storage uses a compact five-dimensional tensor that minimizes memory fragmentation:
(2, num_blocks, block_size, num_kv_heads, head_size)
- Dimension 0 separates keys (index 0) from values (index 1).
- This layout is defined in the flash attention backend and is ideal for bulk allocation and deallocation.
Paged Layout: The Kernel-Optimized View
When the ROCM attention backend processes requests, it requires a paged layout that matches the internal memory format of the underlying kernel:
key : (num_blocks, num_kv_heads, head_size//x, block_size, x)
value : (num_blocks, num_kv_heads, head_size, block_size)
Here, x = 16 // element_size (e.g., 2 for fp16). Without PagedAttention, converting between these formats would require expensive permute or contiguous operations that allocate new GPU memory.
Core PagedAttention Operations
The PagedAttention utility class is defined in vllm/v1/attention/ops/paged_attn.py (lines 15-52). It exposes two static methods that handle layout conversion and cache updates without data movement.
split_kv_cache: Zero-Copy Tensor Reinterpretation
The split_kv_cache(kv_cache, num_kv_heads, head_size) method creates paged views of the flash-layout cache. According to the source code, this is a pure view operation—no data is copied, only the tensor strides are reinterpreted.
This method is invoked in vllm/v1/attention/backends/rocm_attn.py at lines 397, 449, and 503 immediately before attention computation to prepare the kernel-ready layout.
write_to_paged_cache: Direct In-Place Updates
The write_to_paged_cache(key, value, key_cache, value_cache, slot_mapping, kv_cache_dtype, k_scale, v_scale) method writes incoming K/V tensors directly into the paged cache slots. It calls the low-level custom operator ops.reshape_and_cache to reshape inputs and scatter them into the correct memory locations specified by slot_mapping.
In rocm_attn.py (lines 60-71), this method is used within RocmAttentionImpl.do_kv_cache_update for power-of-two block sizes, eliminating the need for intermediate staging buffers.
Memory Efficiency Mechanisms
PagedAttention delivers memory efficiency through four specific architectural choices implemented in the vLLM source code:
Zero-Copy Layout Switching. By treating the paged layout as a view on the same underlying storage as the flash layout, split_kv_cache avoids allocating duplicate GPU memory buffers. This prevents the memory bloat that would occur from explicit transpose or permute operations.
Cache-Friendly Direct Writes. The write_to_paged_cache method reshapes incoming K/V tensors to match the paged layout and writes them directly into the destination slots. This eliminates temporary buffers that would otherwise consume memory during the write phase.
Block-Size Agnostic Optimization. For non-standard block sizes—such as 544 used in Qwen-3 models—the system automatically falls back to Triton-based logic while still reusing the same underlying memory buffer. This ensures memory efficiency regardless of model-specific block configurations.
Reduced Kernel Launch Overhead. Because the cache remains in the layout expected by the attention kernel, the ROCm and CUDA backends can read keys and values without index gymnastics or on-the-fly reshaping, leading to higher throughput and lower latency.
Code Example: Direct PagedAttention Usage
The following example demonstrates how to use PagedAttention to convert flash-layout caches to paged views and write new tokens efficiently:
import torch
from vllm.v1.attention.ops.paged_attn import PagedAttention
# 1. Create a flash-layout KV cache (canonical storage format)
num_blocks = 64
block_size = 32
num_kv_heads = 8
head_dim = 128
kv_cache = torch.randn(
2, # [key, value]
num_blocks,
block_size,
num_kv_heads,
head_dim,
dtype=torch.float16,
device="cuda",
)
# 2. Split into paged views required by the kernel
key_cache, value_cache = PagedAttention.split_kv_cache(
kv_cache, num_kv_heads=num_kv_heads, head_size=head_dim
)
# 3. Write new K/V pairs into specific cache slots (e.g., for new tokens)
batch_size = 4
new_key = torch.randn(batch_size, num_kv_heads, head_dim, dtype=torch.float16, device="cuda")
new_value = torch.randn(batch_size, num_kv_heads, head_dim, dtype=torch.float16, device="cuda")
slot_mapping = torch.arange(batch_size, device="cuda", dtype=torch.int32)
PagedAttention.write_to_paged_cache(
key=new_key,
value=new_value,
key_cache=key_cache,
value_cache=value_cache,
slot_mapping=slot_mapping,
kv_cache_dtype="auto",
k_scale=torch.tensor(1.0, device="cuda"),
v_scale=torch.tensor(1.0, device="cuda"),
)
In this implementation, key_cache and value_cache point to the same GPU memory as the original kv_cache tensor. The write_to_paged_cache call updates these buffers in-place without allocating temporary storage.
Summary
- PagedAttention in
vllm/v1/attention/ops/paged_attn.pyprovides zero-copy translation between flash and paged KV cache layouts. - The
split_kv_cachemethod creates kernel-compatible views without data duplication, invoked inrocm_attn.pybefore attention computation. - The
write_to_paged_cachemethod performs in-place scatter updates viaops.reshape_and_cache, eliminating intermediate buffers. - This architecture supports both standard power-of-two block sizes and irregular sizes (e.g., 544) while maintaining memory efficiency.
- The design reduces GPU memory churn and kernel launch overhead for both ROCm and CUDA backends.
Frequently Asked Questions
What is PagedAttention in vLLM?
PagedAttention is a utility class in vllm/v1/attention/ops/paged_attn.py that manages the memory layout translation for KV caches. It enables the inference engine to store caches in a compact flash format while presenting them to GPU kernels in a paged layout, all without copying data between buffers.
How does PagedAttention differ from standard KV cache management?
Standard implementations often transpose or permute tensor dimensions during the attention forward pass, which allocates temporary GPU memory. PagedAttention uses PyTorch view operations to reinterpret the same memory buffer with different strides, achieving zero-copy layout switching that preserves memory bandwidth.
Does PagedAttention support non-standard block sizes?
Yes. While the optimized path in rocm_attn.py handles power-of-two block sizes, PagedAttention automatically falls back to Triton-based implementations for non-standard sizes like 544 (used in Qwen-3 models). In both cases, the underlying memory buffer is reused without allocation overhead.
Which vLLM components use PagedAttention?
The primary consumer is the ROCm attention backend (vllm/v1/attention/backends/rocm_attn.py), which calls split_kv_cache before decoding (lines 397, 449, 503) and write_to_paged_cache during KV updates (lines 60-71). The utility is also referenced in test suites such as tests/v1/spec_decode/test_tree_attention.py to validate layout compatibility.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →