# How PagedAttention Works in vLLM to Maximize Memory Efficiency

> Discover how PagedAttention in vLLM optimizes GPU memory by eliminating copies and processing KV cache directly on compact buffers. Maximize your LLM's efficiency.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: internals
- Published: 2026-03-03

---

**PagedAttention eliminates GPU memory overhead in vLLM by reinterpreting KV cache tensors as zero-copy views rather than performing expensive data copies, allowing the ROCm and CUDA backends to process attention directly on compact flash-layout buffers.**

Efficient key-value (KV) cache management is the bottleneck for high-throughput LLM serving. In the vllm-project/vllm repository, **PagedAttention** bridges the gap between the memory-efficient flash layout used for storage and the paged layout required by GPU attention kernels, achieving zero-copy tensor translation.

## The Memory Layout Challenge: Flash vs. Paged

vLLM stores KV caches in two distinct formats depending on whether the data is being stored or computed upon. The friction between these layouts creates the need for PagedAttention’s view-based translation layer.

### Flash Layout: The Canonical Storage Format

The default storage uses a compact five-dimensional tensor that minimizes memory fragmentation:

```

(2, num_blocks, block_size, num_kv_heads, head_size)

```

- Dimension 0 separates keys (index 0) from values (index 1).
- This layout is defined in the flash attention backend and is ideal for bulk allocation and deallocation.

### Paged Layout: The Kernel-Optimized View

When the **ROCM attention backend** processes requests, it requires a paged layout that matches the internal memory format of the underlying kernel:

```

key : (num_blocks, num_kv_heads, head_size//x, block_size, x)
value : (num_blocks, num_kv_heads, head_size, block_size)

```

Here, `x = 16 // element_size` (e.g., 2 for fp16). Without PagedAttention, converting between these formats would require expensive `permute` or `contiguous` operations that allocate new GPU memory.

## Core PagedAttention Operations

The `PagedAttention` utility class is defined in [`vllm/v1/attention/ops/paged_attn.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/ops/paged_attn.py) (lines 15-52). It exposes two static methods that handle layout conversion and cache updates without data movement.

### `split_kv_cache`: Zero-Copy Tensor Reinterpretation

The `split_kv_cache(kv_cache, num_kv_heads, head_size)` method creates paged views of the flash-layout cache. According to the source code, this is a **pure view operation**—no data is copied, only the tensor strides are reinterpreted.

This method is invoked in [`vllm/v1/attention/backends/rocm_attn.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/backends/rocm_attn.py) at lines 397, 449, and 503 immediately before attention computation to prepare the kernel-ready layout.

### `write_to_paged_cache`: Direct In-Place Updates

The `write_to_paged_cache(key, value, key_cache, value_cache, slot_mapping, kv_cache_dtype, k_scale, v_scale)` method writes incoming K/V tensors directly into the paged cache slots. It calls the low-level custom operator `ops.reshape_and_cache` to reshape inputs and scatter them into the correct memory locations specified by `slot_mapping`.

In [`rocm_attn.py`](https://github.com/vllm-project/vllm/blob/main/rocm_attn.py) (lines 60-71), this method is used within `RocmAttentionImpl.do_kv_cache_update` for power-of-two block sizes, eliminating the need for intermediate staging buffers.

## Memory Efficiency Mechanisms

PagedAttention delivers memory efficiency through four specific architectural choices implemented in the vLLM source code:

**Zero-Copy Layout Switching.** By treating the paged layout as a view on the same underlying storage as the flash layout, `split_kv_cache` avoids allocating duplicate GPU memory buffers. This prevents the memory bloat that would occur from explicit transpose or permute operations.

**Cache-Friendly Direct Writes.** The `write_to_paged_cache` method reshapes incoming K/V tensors to match the paged layout and writes them directly into the destination slots. This eliminates temporary buffers that would otherwise consume memory during the write phase.

**Block-Size Agnostic Optimization.** For non-standard block sizes—such as 544 used in Qwen-3 models—the system automatically falls back to Triton-based logic while still reusing the same underlying memory buffer. This ensures memory efficiency regardless of model-specific block configurations.

**Reduced Kernel Launch Overhead.** Because the cache remains in the layout expected by the attention kernel, the ROCm and CUDA backends can read keys and values without index gymnastics or on-the-fly reshaping, leading to higher throughput and lower latency.

## Code Example: Direct PagedAttention Usage

The following example demonstrates how to use PagedAttention to convert flash-layout caches to paged views and write new tokens efficiently:

```python
import torch
from vllm.v1.attention.ops.paged_attn import PagedAttention

# 1. Create a flash-layout KV cache (canonical storage format)

num_blocks = 64
block_size = 32
num_kv_heads = 8
head_dim = 128
kv_cache = torch.randn(
    2,                # [key, value]

    num_blocks,
    block_size,
    num_kv_heads,
    head_dim,
    dtype=torch.float16,
    device="cuda",
)

# 2. Split into paged views required by the kernel

key_cache, value_cache = PagedAttention.split_kv_cache(
    kv_cache, num_kv_heads=num_kv_heads, head_size=head_dim
)

# 3. Write new K/V pairs into specific cache slots (e.g., for new tokens)

batch_size = 4
new_key = torch.randn(batch_size, num_kv_heads, head_dim, dtype=torch.float16, device="cuda")
new_value = torch.randn(batch_size, num_kv_heads, head_dim, dtype=torch.float16, device="cuda")
slot_mapping = torch.arange(batch_size, device="cuda", dtype=torch.int32)

PagedAttention.write_to_paged_cache(
    key=new_key,
    value=new_value,
    key_cache=key_cache,
    value_cache=value_cache,
    slot_mapping=slot_mapping,
    kv_cache_dtype="auto",
    k_scale=torch.tensor(1.0, device="cuda"),
    v_scale=torch.tensor(1.0, device="cuda"),
)

```

In this implementation, `key_cache` and `value_cache` point to the same GPU memory as the original `kv_cache` tensor. The `write_to_paged_cache` call updates these buffers in-place without allocating temporary storage.

## Summary

- **PagedAttention** in [`vllm/v1/attention/ops/paged_attn.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/ops/paged_attn.py) provides zero-copy translation between flash and paged KV cache layouts.
- The `split_kv_cache` method creates kernel-compatible views without data duplication, invoked in [`rocm_attn.py`](https://github.com/vllm-project/vllm/blob/main/rocm_attn.py) before attention computation.
- The `write_to_paged_cache` method performs in-place scatter updates via `ops.reshape_and_cache`, eliminating intermediate buffers.
- This architecture supports both standard power-of-two block sizes and irregular sizes (e.g., 544) while maintaining memory efficiency.
- The design reduces GPU memory churn and kernel launch overhead for both ROCm and CUDA backends.

## Frequently Asked Questions

### What is PagedAttention in vLLM?

PagedAttention is a utility class in [`vllm/v1/attention/ops/paged_attn.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/ops/paged_attn.py) that manages the memory layout translation for KV caches. It enables the inference engine to store caches in a compact flash format while presenting them to GPU kernels in a paged layout, all without copying data between buffers.

### How does PagedAttention differ from standard KV cache management?

Standard implementations often transpose or permute tensor dimensions during the attention forward pass, which allocates temporary GPU memory. PagedAttention uses PyTorch view operations to reinterpret the same memory buffer with different strides, achieving zero-copy layout switching that preserves memory bandwidth.

### Does PagedAttention support non-standard block sizes?

Yes. While the optimized path in [`rocm_attn.py`](https://github.com/vllm-project/vllm/blob/main/rocm_attn.py) handles power-of-two block sizes, PagedAttention automatically falls back to Triton-based implementations for non-standard sizes like 544 (used in Qwen-3 models). In both cases, the underlying memory buffer is reused without allocation overhead.

### Which vLLM components use PagedAttention?

The primary consumer is the ROCm attention backend ([`vllm/v1/attention/backends/rocm_attn.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/backends/rocm_attn.py)), which calls `split_kv_cache` before decoding (lines 397, 449, 503) and `write_to_paged_cache` during KV updates (lines 60-71). The utility is also referenced in test suites such as [`tests/v1/spec_decode/test_tree_attention.py`](https://github.com/vllm-project/vllm/blob/main/tests/v1/spec_decode/test_tree_attention.py) to validate layout compatibility.