# How to Configure Attention Backends (FlashAttention and FlashInfer) in vLLM

> Learn to configure attention backends like FlashAttention and FlashInfer in vLLM using CLI flags API or config files for faster inference. Optimize your LLM performance effortlessly.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**Use the `--attention-backend` CLI flag, the `attention_config` parameter in the Python API, or JSON configuration files to select between FlashAttention (`FLASH_ATTN`) and FlashInfer (`FLASHINFER`), with automatic platform-specific fallback when unspecified.**

vLLM exposes a pluggable attention architecture that lets you swap the kernel implementation powering transformer inference. According to the vllm-project/vllm source code, the system uses an `AttentionConfig` dataclass and backend registry to route attention operations to optimized implementations based on your hardware and model constraints.

## Configuration Methods for Attention Backends

You can specify the attention backend through three primary interfaces: command-line arguments, direct Python API calls, or structured configuration files.

### Command-Line Interface Configuration

The simplest method uses the `--attention-backend` flag defined in [`vllm/engine/arg_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/engine/arg_utils.py) (around line 777). Pass the enum value directly to force a specific implementation:

```bash
vllm serve meta-llama/Meta-Llama-3-8B \
    --attention-backend FLASH_ATTN

```

For structured configuration with backend-specific options, use the `--attention-config` prefix (short form `-ac`):

```bash
vllm serve meta-llama/Meta-Llama-3-8B \
    --attention-config.backend FLASH_ATTN \
    --attention-config.flash_attn_version 4

```

To enable FlashInfer with TRT-LLM ragged prefill support:

```bash
vllm serve mistralai/Mistral-7B-v0.1 \
    --attention-config.backend FLASHINFER \
    --attention-config.use_trtllm_attention true

```

### Python API Configuration

When constructing an `LLM` instance programmatically, pass an `AttentionConfig` object from [`vllm/config/attention.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/attention.py) to explicitly control the backend:

```python
from vllm import LLM
from vllm.config import AttentionConfig
from vllm.v1.attention.backends.registry import AttentionBackendEnum

# Configure FlashAttention version 3

attn_cfg = AttentionConfig(
    backend=AttentionBackendEnum.FLASH_ATTN,
    flash_attn_version=3,
    use_prefill_decode_attention=True,
)

llm = LLM(
    model="meta-llama/Meta-Llama-3-8B",
    attention_config=attn_cfg,
)

```

For convenience, you can also pass the backend as a string using the `attention_backend` argument, which resolves to the enum internally:

```python
from vllm import LLM

llm = LLM(
    model="openai-community/gpt2",
    attention_backend="FLASHINFER",
)

```

### JSON Configuration Files

vLLM supports loading parameters from `.json` or `.yaml` files. Define the `attention_config` object with the backend name and specific flags:

```json
{
  "model": "facebook/opt-13b",
  "attention_config": {
    "backend": "FLASH_ATTN",
    "flash_attn_version": 4,
    "use_prefill_decode_attention": true
  }
}

```

Launch the server with:

```bash
vllm serve --json-config ./my_config.json

```

## Backend Selection Architecture

When you instantiate a model, vLLM builds a global `VllmConfig` (defined in [`vllm/config/vllm.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/vllm.py)) containing your `attention_config`. The selection logic in [`vllm/v1/attention/selector.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/selector.py) processes this configuration through the `get_attn_backend` function, which constructs an `AttentionSelectorConfig` describing your model's head size, dtype, and KV-cache requirements.

If you set `backend=None`, the selector iterates over a platform-specific priority list (documented in [`tools/pre_commit/generate_attention_backend_docs.py`](https://github.com/vllm-project/vllm/blob/main/tools/pre_commit/generate_attention_backend_docs.py)) and selects the first compatible implementation based on compute capability and data type support. The registry in [`vllm/v1/attention/backends/registry.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/backends/registry.py) maps enum values to concrete classes:

- `FLASH_ATTN` → `vllm.v1.attention.backends.flash_attn.FlashAttentionBackend`
- `FLASHINFER` → `vllm.v1.attention.backends.flashinfer.FlashInferBackend`

The selector automatically adjusts the KV-cache layout via `set_kv_cache_layout` if the chosen backend requires a specific memory format, as implemented in the backend class's `get_required_kv_cache_layout()` method.

## FlashAttention vs. FlashInfer Implementation Details

Both backends implement the `AttentionBackend` interface but target different optimization strategies.

**FlashAttention** (`FLASH_ATTN`) in [`vllm/v1/attention/backends/flash_attn.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/backends/flash_attn.py) supports versions 2, 3, and 4 through the `flash_attn_version` field, working on NVIDIA GPUs with compute capability ≥ 8.0.

**FlashInfer** (`FLASHINFER`) in [`vllm/v1/attention/backends/flashinfer.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/backends/flashinfer.py) provides additional features including TRT-LLM-style ragged prefill (enabled via `use_trtllm_attention`), quantization support, and fine-grained toggles like `disable_flashinfer_prefill` or `disable_flashinfer_q_quantization` for specialized deployment scenarios.

## Advanced Configuration Parameters

Several flags interact with the attention backend selection to fine-tune performance:

- **`use_prefill_decode_attention`** – Splits prefill and decode kernels for the selected backend to optimize throughput.
- **`flash_attn_version`** – Forces a specific FlashAttention implementation (2, 3, or 4) when using `FLASH_ATTN`.
- **`use_trtllm_attention`** – Enables TRT-LLM attention paths within FlashInfer for specific quantization and memory layouts.
- **`disable_flashinfer_prefill`** and **`disable_flashinfer_q_quantization`** – Disable specific FlashInfer optimizations when compatibility issues arise.

These fields reside in the `AttentionConfig` dataclass and propagate to the kernel implementation at runtime.

## Summary

- **Use `--attention-backend`** for quick CLI selection between `FLASH_ATTN` and `FLASHINFER`.
- **Leverage `AttentionConfig`** in Python for programmatic control over backend versions and advanced flags.
- **Rely on automatic selection** by leaving the backend unspecified, allowing [`vllm/v1/attention/selector.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/selector.py) to choose the first compatible implementation from the platform priority list.
- **Reference implementation files** [`vllm/v1/attention/backends/flash_attn.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/backends/flash_attn.py) and [`vllm/v1/attention/backends/flashinfer.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/backends/flashinfer.py) for kernel-specific capabilities and constraints.
- **Adjust KV-cache layouts** automatically or manually based on backend requirements via the configuration interface.

## Frequently Asked Questions

### What is the default attention backend if I do not specify one?

When `backend` is `None` in `AttentionConfig`, vLLM invokes the automatic selector in [`vllm/v1/attention/selector.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/attention/selector.py). This function iterates over a hardware-specific priority list defined in the backend registry and selects the first implementation that supports your GPU's compute capability, data type, and KV-cache layout requirements.

### Can I use FlashInfer on older GPUs that do not support FlashAttention?

FlashInfer generally requires similar or more recent hardware capabilities than FlashAttention, depending on the specific kernel features enabled. The selector validates each candidate against your platform's constraints; if neither backend is compatible, vLLM raises a detailed error indicating which hardware or configuration requirements are unmet. Check [`tools/pre_commit/generate_attention_backend_docs.py`](https://github.com/vllm-project/vllm/blob/main/tools/pre_commit/generate_attention_backend_docs.py) for the compatibility matrix.

### How do I force a specific FlashAttention version?

Pass the `flash_attn_version` parameter through the structured CLI syntax (`--attention-config.flash_attn_version 3`) or include it in your `AttentionConfig` instantiation in Python. Valid values are 2, 3, and 4, corresponding to different kernel optimizations in the FlashAttention family.

### What is the difference between `--attention-backend` and `--attention-config`?

The `--attention-backend` flag accepts a simple string enum value (e.g., `FLASH_ATTN`) for quick backend selection. The `--attention-config` interface (or `-ac` shorthand) provides a namespaced way to set nested configuration fields like `flash_attn_version` or `use_trtllm_attention`, offering granular control over backend-specific parameters that the simple flag cannot express.