How Chunked Prefill Works in vLLM: Implementation Guide and Configuration Best Practices

Chunked prefill enables vLLM to split long prompt processing across multiple scheduler iterations by breaking the prefill phase into smaller chunks that fit within the max_num_batched_tokens budget, allowing prompts longer than the per-iteration GPU memory limit to execute successfully.

In the vLLM inference engine, chunked prefill is a scheduling optimization that addresses memory and compute constraints when handling extremely long input sequences. By default enabled in modern vLLM versions, this mechanism ensures that prompts exceeding the configured token budget do not stall the system or cause out-of-memory errors. Understanding how this feature operates at the source code level helps you optimize throughput for production deployments.

What is Chunked Prefill?

When vLLM receives a request with a prompt longer than the scheduler's per-iteration token budget (controlled by max_num_batched_tokens), the system must either reject the request or process it incrementally. Chunked prefill takes the latter approach by dividing the prompt's KV-cache population phase into discrete segments.

Each scheduler iteration processes only a portion of the prompt, allocates the corresponding KV-cache slots, and preserves the partial state. Subsequent iterations resume computation from the last cached position until the entire prompt resides in the KV cache. The final result is identical to monolithic prefill, but GPU memory usage remains bounded and predictable.

The Core Mechanism

The implementation centers on the num_computed_tokens field tracked per request. In vllm/v1/core/sched/scheduler.py, the scheduler calculates num_new_tokens as the remaining uncached portion of the prompt. When enable_chunked_prefill is active and num_new_tokens exceeds the current iteration budget, the scheduler caps the value to the budget limit rather than blocking the request entirely.

Configuration and Control

The enable_chunked_prefill Flag

The primary configuration resides in vllm/config/scheduler.py within the SchedulerConfig class. The boolean flag enable_chunked_prefill defaults to True in current vLLM versions, but you can explicitly control it via the Python API or CLI:

from vllm import LLMEngine, EngineArgs

engine = LLMEngine(
    EngineArgs(
        model="meta-llama/Llama-2-7b-hf",
        max_num_batched_tokens=2048,
        enable_chunked_prefill=True,  # Explicitly enable (default)

    )
)

Via the OpenAI-compatible server CLI:

python -m vllm.entrypoints.openai.api_server \
    --model mistralai/Mistral-7B-v0.1 \
    --max-num-batched-tokens 2048 \
    --enable-chunked-prefill

Automatic Disabling for Unsupported Models

The engine automatically disables chunked prefill for incompatible architectures. In vllm/v1/engine/core.py (lines 135-138), the EngineCore checks whether the model provides a KV cache. Encoder-decoder models (like T5 or BART) and certain state-space models (SSMs) lack the unified KV-cache structure required for incremental prompt processing, forcing enable_chunked_prefill to False regardless of user configuration.

Scheduler Decision Logic

The critical logic resides in vllm/v1/core/sched/scheduler.py (lines 64-72). Before allocating tokens, the scheduler compares num_new_tokens against the remaining budget:

  • If chunked prefill is enabled: The scheduler caps num_new_tokens to the available budget and proceeds with partial allocation.
  • If chunked prefill is disabled: The scheduler breaks the loop, leaving the request in the waiting queue until a future iteration has sufficient capacity.

This distinction ensures that disabling the flag guarantees atomic prompt processing—either the entire prompt fits in one iteration or the request blocks indefinitely.

How It Works Under the Hood

Partial KV-Cache Allocation

During a chunked iteration, kv_cache_manager.allocate_slots reserves exactly num_new_tokens blocks rather than the full prompt length. The request's num_computed_tokens counter is not incremented until after the attention kernel completes, ensuring that failed iterations do not corrupt the cached state.

This incremental approach appears in the allocation logic within vllm/v1/kv_cache_interface.py, where the manager handles partial block mapping for ongoing prefill requests.

The Attention Kernel

The actual computation occurs in vllm/v1/attention/ops/chunked_prefill_paged_decode.py (lines 31-41). This Triton kernel receives the current query slice parameters (query_start_loc and slice length) and executes standard attention computation on that subset. When the slice length equals 1 (single token), the kernel skips unnecessary masking operations, optimizing the decode path.

The kernel writes attention outputs directly to the pre-allocated KV-cache slots, making the state persistent across scheduler iterations.

State Management Across Iterations

After each chunk completes, the scheduler updates request.num_computed_tokens by adding the processed num_new_tokens. The request remains in the RUNNING state if additional prompt tokens exist, or transitions to generation phase once num_computed_tokens equals the full prompt length. This state machine ensures no token is processed twice while maintaining strict memory bounds.

When to Enable or Disable Chunked Prefill

  • Long prompts exceeding max_num_batched_tokens: Keep enable_chunked_prefill=True (default). This is the primary use case—without chunking, requests longer than 2048 tokens (default budget) would block or fail.

  • Encoder-decoder architectures: The system automatically disables this feature. Do not attempt to force-enable for T5, BART, or similar models lacking decoder-side KV caches.

  • Multimodal inputs: When processing images or audio where atomic scheduling is required, consider pairing chunked prefill with disable_chunked_mm_input=True to prevent media segments from splitting across iterations.

  • Low-latency short-prompt scenarios: You may disable chunked prefill to eliminate scheduling overhead, though the performance gain is typically marginal compared to decode costs.

  • Models without KV caches: Certain sliding-window or SSM architectures disable chunked prefill automatically via the check in EngineCore.

Implementation Examples

Python API Configuration

from vllm import LLMEngine, EngineArgs

# Handle a 4096-token prompt with a 2048-token budget

engine = LLMEngine(
    EngineArgs(
        model="meta-llama/Llama-2-7b-hf",
        max_num_batched_tokens=2048,
        enable_chunked_prefill=True,
    )
)

request_id = engine.add_request(prompt="..." * 4000, max_new_tokens=256)

while not engine.is_finished(request_id):
    engine.step()

CLI Usage Patterns


# Standard deployment with chunked prefill

python -m vllm.entrypoints.openai.api_server \
    --model facebook/opt-125m \
    --max-num-batched-tokens 1024 \
    --enable-chunked-prefill

# Explicitly disable for guaranteed atomic processing

python -m vllm.entrypoints.openai.api_server \
    --model facebook/opt-125m \
    --max-num-batched-tokens 1024 \
    --disable-chunked-prefill

Debugging Scheduler Decisions

To verify chunked prefill behavior, enable INFO-level logging:

import logging
logging.basicConfig(level=logging.INFO)

The scheduler emits specific messages when automatically disabling chunked prefill for incompatible models, such as "Disabling chunked prefill for model without KVCache" in vllm/v1/engine/core.py.

Summary

  • Chunked prefill splits long prompt processing across scheduler iterations by capping num_new_tokens to the max_num_batched_tokens budget and tracking progress via num_computed_tokens.
  • The feature defaults to enabled (True) in SchedulerConfig but automatically disables itself for encoder-decoder models and architectures lacking KV caches.
  • The chunked_prefill_paged_decode kernel in vllm/v1/attention/ops/ handles the actual attention computation for each slice, writing results incrementally to the KV cache.
  • Disable chunked prefill only when you require atomic prompt processing and can guarantee all inputs fit within your max_num_batched_tokens limit, or when working with specific multimodal inputs requiring atomic scheduling.

Frequently Asked Questions

Does chunked prefill affect output quality?

No. Chunked prefill is mathematically equivalent to monolithic prefill. The chunked_prefill_paged_decode kernel computes identical attention scores and KV-cache values, merely distributing the work across multiple GPU kernel launches. The final model state after the last chunk completes matches what would result from single-iteration processing.

What happens if I disable chunked prefill and send a long prompt?

If chunked prefill is disabled and a prompt exceeds max_num_batched_tokens, the scheduler blocks the request in the waiting queue indefinitely. The request will not begin processing until the budget increases or other requests complete, effectively stalling that specific input. No error is raised—the request simply remains pending.

Is there a performance overhead to chunked prefill?

The overhead consists of additional scheduler iterations and kernel launch latency. For very long prompts, this is negligible compared to the compute saved by avoiding out-of-memory errors. For short prompts that fit comfortably within the budget, the overhead is minimal (microseconds per request), though you may disable the feature if you require deterministic, single-pass scheduling behavior.

Which models cannot use chunked prefill?

Encoder-decoder models like T5 and BART cannot use chunked prefill because they lack a unified KV cache for the decoder. Additionally, certain state-space models (SSMs) or architectures with alternative attention mechanisms that do not expose a standard KV cache interface will trigger the automatic disabling logic in vllm/v1/engine/core.py. The system handles this detection automatically based on model capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →