# How Chunked Prefill Works in vLLM: Implementation Guide and Configuration Best Practices

> Learn how chunked prefill works in vLLM to process long prompts efficiently by splitting them into smaller chunks. Discover implementation details and configuration best practices for optimal performance.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: deep-dive
- Published: 2026-03-03

---

**Chunked prefill enables vLLM to split long prompt processing across multiple scheduler iterations by breaking the prefill phase into smaller chunks that fit within the `max_num_batched_tokens` budget, allowing prompts longer than the per-iteration GPU memory limit to execute successfully.**

In the vLLM inference engine, **chunked prefill** is a scheduling optimization that addresses memory and compute constraints when handling extremely long input sequences. By default enabled in modern vLLM versions, this mechanism ensures that prompts exceeding the configured token budget do not stall the system or cause out-of-memory errors. Understanding how this feature operates at the source code level helps you optimize throughput for production deployments.

## What is Chunked Prefill?

When vLLM receives a request with a prompt longer than the scheduler's per-iteration token budget (controlled by `max_num_batched_tokens`), the system must either reject the request or process it incrementally. Chunked prefill takes the latter approach by dividing the prompt's **KV-cache population phase** into discrete segments.

Each scheduler iteration processes only a portion of the prompt, allocates the corresponding KV-cache slots, and preserves the partial state. Subsequent iterations resume computation from the last cached position until the entire prompt resides in the KV cache. The final result is identical to monolithic prefill, but GPU memory usage remains bounded and predictable.

### The Core Mechanism

The implementation centers on the `num_computed_tokens` field tracked per request. In [`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py), the scheduler calculates `num_new_tokens` as the remaining uncached portion of the prompt. When `enable_chunked_prefill` is active and `num_new_tokens` exceeds the current iteration budget, the scheduler caps the value to the budget limit rather than blocking the request entirely.

## Configuration and Control

### The enable_chunked_prefill Flag

The primary configuration resides in [`vllm/config/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/scheduler.py) within the `SchedulerConfig` class. The boolean flag `enable_chunked_prefill` defaults to `True` in current vLLM versions, but you can explicitly control it via the Python API or CLI:

```python
from vllm import LLMEngine, EngineArgs

engine = LLMEngine(
    EngineArgs(
        model="meta-llama/Llama-2-7b-hf",
        max_num_batched_tokens=2048,
        enable_chunked_prefill=True,  # Explicitly enable (default)

    )
)

```

Via the OpenAI-compatible server CLI:

```bash
python -m vllm.entrypoints.openai.api_server \
    --model mistralai/Mistral-7B-v0.1 \
    --max-num-batched-tokens 2048 \
    --enable-chunked-prefill

```

### Automatic Disabling for Unsupported Models

The engine automatically disables chunked prefill for incompatible architectures. In [`vllm/v1/engine/core.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/core.py) (lines 135-138), the `EngineCore` checks whether the model provides a KV cache. Encoder-decoder models (like T5 or BART) and certain state-space models (SSMs) lack the unified KV-cache structure required for incremental prompt processing, forcing `enable_chunked_prefill` to `False` regardless of user configuration.

### Scheduler Decision Logic

The critical logic resides in [`vllm/v1/core/sched/scheduler.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py) (lines 64-72). Before allocating tokens, the scheduler compares `num_new_tokens` against the remaining budget:

- **If chunked prefill is enabled**: The scheduler caps `num_new_tokens` to the available budget and proceeds with partial allocation.
- **If chunked prefill is disabled**: The scheduler breaks the loop, leaving the request in the waiting queue until a future iteration has sufficient capacity.

This distinction ensures that disabling the flag guarantees atomic prompt processing—either the entire prompt fits in one iteration or the request blocks indefinitely.

## How It Works Under the Hood

### Partial KV-Cache Allocation

During a chunked iteration, `kv_cache_manager.allocate_slots` reserves exactly `num_new_tokens` blocks rather than the full prompt length. The request's `num_computed_tokens` counter is not incremented until after the attention kernel completes, ensuring that failed iterations do not corrupt the cached state.

This incremental approach appears in the allocation logic within [`vllm/v1/kv_cache_interface.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_cache_interface.py), where the manager handles partial block mapping for ongoing prefill requests.

### The Attention Kernel

The actual computation occurs in [`vllm/v1/attention/ops/chunked_prefill_paged_decode.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/ops/chunked_prefill_paged_decode.py) (lines 31-41). This Triton kernel receives the current query slice parameters (`query_start_loc` and slice length) and executes standard attention computation on that subset. When the slice length equals 1 (single token), the kernel skips unnecessary masking operations, optimizing the decode path.

The kernel writes attention outputs directly to the pre-allocated KV-cache slots, making the state persistent across scheduler iterations.

### State Management Across Iterations

After each chunk completes, the scheduler updates `request.num_computed_tokens` by adding the processed `num_new_tokens`. The request remains in the **RUNNING** state if additional prompt tokens exist, or transitions to generation phase once `num_computed_tokens` equals the full prompt length. This state machine ensures no token is processed twice while maintaining strict memory bounds.

## When to Enable or Disable Chunked Prefill

- **Long prompts exceeding `max_num_batched_tokens`**: Keep `enable_chunked_prefill=True` (default). This is the primary use case—without chunking, requests longer than 2048 tokens (default budget) would block or fail.
  
- **Encoder-decoder architectures**: The system automatically disables this feature. Do not attempt to force-enable for T5, BART, or similar models lacking decoder-side KV caches.

- **Multimodal inputs**: When processing images or audio where atomic scheduling is required, consider pairing chunked prefill with `disable_chunked_mm_input=True` to prevent media segments from splitting across iterations.

- **Low-latency short-prompt scenarios**: You may disable chunked prefill to eliminate scheduling overhead, though the performance gain is typically marginal compared to decode costs.

- **Models without KV caches**: Certain sliding-window or SSM architectures disable chunked prefill automatically via the check in `EngineCore`.

## Implementation Examples

### Python API Configuration

```python
from vllm import LLMEngine, EngineArgs

# Handle a 4096-token prompt with a 2048-token budget

engine = LLMEngine(
    EngineArgs(
        model="meta-llama/Llama-2-7b-hf",
        max_num_batched_tokens=2048,
        enable_chunked_prefill=True,
    )
)

request_id = engine.add_request(prompt="..." * 4000, max_new_tokens=256)

while not engine.is_finished(request_id):
    engine.step()

```

### CLI Usage Patterns

```bash

# Standard deployment with chunked prefill

python -m vllm.entrypoints.openai.api_server \
    --model facebook/opt-125m \
    --max-num-batched-tokens 1024 \
    --enable-chunked-prefill

# Explicitly disable for guaranteed atomic processing

python -m vllm.entrypoints.openai.api_server \
    --model facebook/opt-125m \
    --max-num-batched-tokens 1024 \
    --disable-chunked-prefill

```

### Debugging Scheduler Decisions

To verify chunked prefill behavior, enable INFO-level logging:

```python
import logging
logging.basicConfig(level=logging.INFO)

```

The scheduler emits specific messages when automatically disabling chunked prefill for incompatible models, such as *"Disabling chunked prefill for model without KVCache"* in [`vllm/v1/engine/core.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/core.py).

## Summary

- **Chunked prefill** splits long prompt processing across scheduler iterations by capping `num_new_tokens` to the `max_num_batched_tokens` budget and tracking progress via `num_computed_tokens`.
- The feature defaults to enabled (`True`) in `SchedulerConfig` but automatically disables itself for encoder-decoder models and architectures lacking KV caches.
- The `chunked_prefill_paged_decode` kernel in `vllm/v1/attention/ops/` handles the actual attention computation for each slice, writing results incrementally to the KV cache.
- Disable chunked prefill only when you require atomic prompt processing and can guarantee all inputs fit within your `max_num_batched_tokens` limit, or when working with specific multimodal inputs requiring atomic scheduling.

## Frequently Asked Questions

### Does chunked prefill affect output quality?

No. Chunked prefill is mathematically equivalent to monolithic prefill. The `chunked_prefill_paged_decode` kernel computes identical attention scores and KV-cache values, merely distributing the work across multiple GPU kernel launches. The final model state after the last chunk completes matches what would result from single-iteration processing.

### What happens if I disable chunked prefill and send a long prompt?

If chunked prefill is disabled and a prompt exceeds `max_num_batched_tokens`, the scheduler blocks the request in the waiting queue indefinitely. The request will not begin processing until the budget increases or other requests complete, effectively stalling that specific input. No error is raised—the request simply remains pending.

### Is there a performance overhead to chunked prefill?

The overhead consists of additional scheduler iterations and kernel launch latency. For very long prompts, this is negligible compared to the compute saved by avoiding out-of-memory errors. For short prompts that fit comfortably within the budget, the overhead is minimal (microseconds per request), though you may disable the feature if you require deterministic, single-pass scheduling behavior.

### Which models cannot use chunked prefill?

Encoder-decoder models like T5 and BART cannot use chunked prefill because they lack a unified KV cache for the decoder. Additionally, certain state-space models (SSMs) or architectures with alternative attention mechanisms that do not expose a standard KV cache interface will trigger the automatic disabling logic in [`vllm/v1/engine/core.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/core.py). The system handles this detection automatically based on model capabilities.