How to Configure Attention Backends (FlashAttention and FlashInfer) in vLLM
Use the --attention-backend CLI flag, the attention_config parameter in the Python API, or JSON configuration files to select between FlashAttention (FLASH_ATTN) and FlashInfer (FLASHINFER), with automatic platform-specific fallback when unspecified.
vLLM exposes a pluggable attention architecture that lets you swap the kernel implementation powering transformer inference. According to the vllm-project/vllm source code, the system uses an AttentionConfig dataclass and backend registry to route attention operations to optimized implementations based on your hardware and model constraints.
Configuration Methods for Attention Backends
You can specify the attention backend through three primary interfaces: command-line arguments, direct Python API calls, or structured configuration files.
Command-Line Interface Configuration
The simplest method uses the --attention-backend flag defined in vllm/engine/arg_utils.py (around line 777). Pass the enum value directly to force a specific implementation:
vllm serve meta-llama/Meta-Llama-3-8B \
--attention-backend FLASH_ATTN
For structured configuration with backend-specific options, use the --attention-config prefix (short form -ac):
vllm serve meta-llama/Meta-Llama-3-8B \
--attention-config.backend FLASH_ATTN \
--attention-config.flash_attn_version 4
To enable FlashInfer with TRT-LLM ragged prefill support:
vllm serve mistralai/Mistral-7B-v0.1 \
--attention-config.backend FLASHINFER \
--attention-config.use_trtllm_attention true
Python API Configuration
When constructing an LLM instance programmatically, pass an AttentionConfig object from vllm/config/attention.py to explicitly control the backend:
from vllm import LLM
from vllm.config import AttentionConfig
from vllm.v1.attention.backends.registry import AttentionBackendEnum
# Configure FlashAttention version 3
attn_cfg = AttentionConfig(
backend=AttentionBackendEnum.FLASH_ATTN,
flash_attn_version=3,
use_prefill_decode_attention=True,
)
llm = LLM(
model="meta-llama/Meta-Llama-3-8B",
attention_config=attn_cfg,
)
For convenience, you can also pass the backend as a string using the attention_backend argument, which resolves to the enum internally:
from vllm import LLM
llm = LLM(
model="openai-community/gpt2",
attention_backend="FLASHINFER",
)
JSON Configuration Files
vLLM supports loading parameters from .json or .yaml files. Define the attention_config object with the backend name and specific flags:
{
"model": "facebook/opt-13b",
"attention_config": {
"backend": "FLASH_ATTN",
"flash_attn_version": 4,
"use_prefill_decode_attention": true
}
}
Launch the server with:
vllm serve --json-config ./my_config.json
Backend Selection Architecture
When you instantiate a model, vLLM builds a global VllmConfig (defined in vllm/config/vllm.py) containing your attention_config. The selection logic in vllm/v1/attention/selector.py processes this configuration through the get_attn_backend function, which constructs an AttentionSelectorConfig describing your model's head size, dtype, and KV-cache requirements.
If you set backend=None, the selector iterates over a platform-specific priority list (documented in tools/pre_commit/generate_attention_backend_docs.py) and selects the first compatible implementation based on compute capability and data type support. The registry in vllm/v1/attention/backends/registry.py maps enum values to concrete classes:
FLASH_ATTN→vllm.v1.attention.backends.flash_attn.FlashAttentionBackendFLASHINFER→vllm.v1.attention.backends.flashinfer.FlashInferBackend
The selector automatically adjusts the KV-cache layout via set_kv_cache_layout if the chosen backend requires a specific memory format, as implemented in the backend class's get_required_kv_cache_layout() method.
FlashAttention vs. FlashInfer Implementation Details
Both backends implement the AttentionBackend interface but target different optimization strategies.
FlashAttention (FLASH_ATTN) in vllm/v1/attention/backends/flash_attn.py supports versions 2, 3, and 4 through the flash_attn_version field, working on NVIDIA GPUs with compute capability ≥ 8.0.
FlashInfer (FLASHINFER) in vllm/v1/attention/backends/flashinfer.py provides additional features including TRT-LLM-style ragged prefill (enabled via use_trtllm_attention), quantization support, and fine-grained toggles like disable_flashinfer_prefill or disable_flashinfer_q_quantization for specialized deployment scenarios.
Advanced Configuration Parameters
Several flags interact with the attention backend selection to fine-tune performance:
use_prefill_decode_attention– Splits prefill and decode kernels for the selected backend to optimize throughput.flash_attn_version– Forces a specific FlashAttention implementation (2, 3, or 4) when usingFLASH_ATTN.use_trtllm_attention– Enables TRT-LLM attention paths within FlashInfer for specific quantization and memory layouts.disable_flashinfer_prefillanddisable_flashinfer_q_quantization– Disable specific FlashInfer optimizations when compatibility issues arise.
These fields reside in the AttentionConfig dataclass and propagate to the kernel implementation at runtime.
Summary
- Use
--attention-backendfor quick CLI selection betweenFLASH_ATTNandFLASHINFER. - Leverage
AttentionConfigin Python for programmatic control over backend versions and advanced flags. - Rely on automatic selection by leaving the backend unspecified, allowing
vllm/v1/attention/selector.pyto choose the first compatible implementation from the platform priority list. - Reference implementation files
vllm/v1/attention/backends/flash_attn.pyandvllm/v1/attention/backends/flashinfer.pyfor kernel-specific capabilities and constraints. - Adjust KV-cache layouts automatically or manually based on backend requirements via the configuration interface.
Frequently Asked Questions
What is the default attention backend if I do not specify one?
When backend is None in AttentionConfig, vLLM invokes the automatic selector in vllm/v1/attention/selector.py. This function iterates over a hardware-specific priority list defined in the backend registry and selects the first implementation that supports your GPU's compute capability, data type, and KV-cache layout requirements.
Can I use FlashInfer on older GPUs that do not support FlashAttention?
FlashInfer generally requires similar or more recent hardware capabilities than FlashAttention, depending on the specific kernel features enabled. The selector validates each candidate against your platform's constraints; if neither backend is compatible, vLLM raises a detailed error indicating which hardware or configuration requirements are unmet. Check tools/pre_commit/generate_attention_backend_docs.py for the compatibility matrix.
How do I force a specific FlashAttention version?
Pass the flash_attn_version parameter through the structured CLI syntax (--attention-config.flash_attn_version 3) or include it in your AttentionConfig instantiation in Python. Valid values are 2, 3, and 4, corresponding to different kernel optimizations in the FlashAttention family.
What is the difference between --attention-backend and --attention-config?
The --attention-backend flag accepts a simple string enum value (e.g., FLASH_ATTN) for quick backend selection. The --attention-config interface (or -ac shorthand) provides a namespaced way to set nested configuration fields like flash_attn_version or use_trtllm_attention, offering granular control over backend-specific parameters that the simple flag cannot express.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →