# How to Configure Compilation Modes (Eager, CUDAGraph, Inductor) in vLLM

> Master vLLM compilation modes eager CUDAGraph and Inductor. Learn to configure settings via CLI or Python API for optimized performance.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: deep-dive
- Published: 2026-03-03

---

**vLLM controls its compilation pipeline through three orthogonal settings in `CompilationConfig`: `mode` (selecting eager versus Inductor), `cudagraph_mode` (controlling CUDA Graph capture strategy), and `backend` (specifying the compiler implementation), which can be set via CLI flags or the Python API.**

vLLM's inference engine uses a sophisticated compilation system to maximize GPU utilization for transformer models. Starting with the v1 architecture, the framework centralizes these optimizations in the `CompilationConfig` class, which is part of the top-level `VllmConfig` defined in [`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py). Understanding how to configure compilation modes in vLLM allows you to balance debugging flexibility against maximum throughput for production deployments.

## Understanding vLLM's Compilation Architecture

### The Three Configuration Knobs

In [`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py), the `CompilationConfig` dataclass defines three primary fields that determine runtime behavior:

- **`mode`**: A `CompilationMode` enum that determines whether `torch.compile` is invoked and which pipeline executes.
- **`cudagraph_mode`**: A `CUDAGraphMode` enum that controls how `torch.cuda.CUDAGraph` capture is applied.
- **`backend`**: A string specifying the compiler backend, such as `"inductor"` or `"eager"`.

These settings are orthogonal. For example, you can enable CUDAGraph capture while running in eager PyTorch mode, or use the Inductor backend without CUDA Graphs.

## Compilation Modes Explained

The `CompilationMode` enum, defined at lines 35-48 of [`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py), offers four distinct strategies:

### NONE (Pure Eager Mode)

Setting `mode=CompilationMode.NONE` disables all `torch.compile` machinery. The model executes in standard PyTorch eager mode, which is essential for debugging or when running on platforms without Inductor support.

### VLLM_COMPILE (Custom Inductor Backend)

`mode=CompilationMode.VLLM_COMPILE` activates vLLM's optimized compilation pipeline. According to the source code, this mode uses a piecewise compilation strategy with shape specialization and custom passes through the Inductor backend. This is the default for v1 and provides the best performance for most CUDA workloads.

### STOCK_TORCH_COMPILE and DYNAMO_TRACE_ONCE

- **`STOCK_TORCH_COMPILE`**: Uses PyTorch's default `torch.compile` pipeline without vLLM customizations.
- **`DYNAMO_TRACE_ONCE`**: Runs Dynamo tracing a single time without re-traces, useful for dynamic shape scenarios where full compilation overhead is undesirable.

## CUDAGraph Modes Explained

The `CUDAGraphMode` enum (lines 51-62 of [`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py)) determines how CUDA Graphs are captured and replayed. This is implemented in [`vllm/v1/worker/gpu_worker.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py) and [`vllm/v1/worker/gpu_ubatch_wrapper.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_ubatch_wrapper.py).

### NONE

`CUDAGraphMode.NONE` disables graph capture entirely. Every forward pass executes kernels individually, which is necessary when debugging kernel launches or using custom operators incompatible with graph capture.

### PIECEWISE

`CUDAGraphMode.PIECEWISE` captures only operations that are safe for CUDAGraph, splitting out attention operations. The `UBatchWrapper` class in [`vllm/v1/worker/gpu_ubatch_wrapper.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_ubatch_wrapper.py) manages these per-size graph objects. This mode is required when using attention kernels that cannot be captured in a full graph.

### FULL and FULL_DECODE_ONLY

- **`FULL`**: Captures the entire forward pass in a single graph. This provides maximum speed for small, fixed-size batches but may cause out-of-memory errors with larger batch sizes.
- **`FULL_DECODE_ONLY`**: Captures the full graph only for decode-only batches (no prefill), balancing throughput and memory.

### FULL_AND_PIECEWISE (Default)

`CUDAGraphMode.FULL_AND_PIECEWISE` is the default for v1. As implemented in [`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py), this mode uses full capture for decode operations and piecewise capture for prefill or mixed batches, automatically selecting the optimal strategy per request type.

## Configuration Methods

You can configure these settings via command-line arguments when using the OpenAI-compatible server, or programmatically through the Python API.

### Command-Line Interface

The entry point in [`vllm/entrypoints/openai/api_server.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/api_server.py) accepts flags that map directly to `CompilationConfig` fields:

```bash

# Default: VLLM_COMPILE with FULL_AND_PIECEWISE CUDAGraph

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-hf

# Pure eager mode (debugging)

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-hf \
    --compilation-mode none \
    --cudagraph-mode none

# Disable CUDAGraph but keep Inductor (useful for debugging graph-partition)

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-hf \
    --cudagraph-mode none

# Force full CUDAGraph capture (no piecewise)

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-hf \
    --cudagraph-mode full

```

### Python API

For embedded use cases, instantiate `CompilationConfig` and pass it to `VllmConfig`:

```python
from vllm import LLM
from vllm.config import VllmConfig, CompilationConfig, CompilationMode, CUDAGraphMode

# Eager mode configuration

eager_cfg = VllmConfig(
    model="meta-llama/Llama-2-7b-hf",
    compilation_config=CompilationConfig(
        mode=CompilationMode.NONE,
        cudagraph_mode=CUDAGraphMode.NONE,
    ),
)
llm_eager = LLM(vllm_config=eager_cfg)

# High-performance Inductor configuration

inductor_cfg = VllmConfig(
    model="meta-llama/Llama-2-7b-hf",
    compilation_config=CompilationConfig(
        mode=CompilationMode.VLLM_COMPILE,
        cudagraph_mode=CUDAGraphMode.FULL_AND_PIECEWISE,
        backend="inductor",
    ),
)
llm_inductor = LLM(vllm_config=inductor_cfg)

```

## Advanced: Customizing Splitting Operations

When using `CUDAGraphMode.PIECEWISE`, vLLM automatically determines which operations to split out based on `CompilationConfig._attention_ops`. You can override this behavior by manually setting `splitting_ops`:

```python
from vllm.config import CompilationConfig, CUDAGraphMode, CompilationMode

cfg = CompilationConfig(
    mode=CompilationMode.VLLM_COMPILE,
    cudagraph_mode=CUDAGraphMode.PIECEWISE,
    splitting_ops=["flash_attention", "custom_op"],  # Split these ops out of graphs

)

```

The logic that populates default splitting ops is implemented in `set_splitting_ops_for_v1` at lines 972-1010 of [`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py).

## Summary

Configuring compilation modes in vLLM involves three orthogonal settings in `CompilationConfig`:

- **`mode`** determines the torch.compile strategy: `NONE` for eager execution, `VLLM_COMPILE` for the optimized Inductor backend with custom passes, or `STOCK_TORCH_COMPILE` for standard PyTorch behavior.
- **`cudagraph_mode`** controls CUDA Graph capture: `NONE` for debugging, `PIECEWISE` for safe partial capture, or `FULL_AND_PIECEWISE` (default) for maximum throughput via full decode graphs and piecewise prefill.
- **`backend`** specifies the compiler implementation, with `"inductor"` delivering best performance on CUDA devices and `"eager"` aiding debugging.

Apply these via CLI flags (`--compilation-mode`, `--cudagraph-mode`), environment variables, or programmatically through `VllmConfig` and `CompilationConfig` to optimize inference for your specific hardware and debugging requirements.

## Frequently Asked Questions

### What is the default compilation mode in vLLM v1?

The v1 engine defaults to `CompilationMode.VLLM_COMPILE` with `CUDAGraphMode.FULL_AND_PIECEWISE` and the Inductor backend. This configuration, defined in [`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py), provides optimal throughput for most CUDA workloads by combining custom Inductor passes with intelligent CUDA Graph capture.

### When should I use Eager mode instead of Inductor?

Use `CompilationMode.NONE` (Eager mode) when debugging kernel launches, working with custom operators that are incompatible with `torch.compile`, or deploying on hardware without Inductor support. Eager mode executes PyTorch operations immediately without graph compilation, providing clearer stack traces at the cost of reduced inference speed.

### Can I use CUDAGraph without torch.compile?

Yes. The `cudagraph_mode` setting is orthogonal to the `mode` setting. You can configure `mode=CompilationMode.NONE` and `cudagraph_mode=CUDAGraphMode.FULL` to capture and replay CUDA Graphs while running in pure eager PyTorch mode. This is useful when you want the runtime efficiency of graph replay without the compilation overhead of Inductor.

### How do I debug compilation issues in vLLM?

Start by running with `--compilation-mode none --cudagraph-mode none` to establish a baseline in pure eager mode. If the issue persists, it is not compilation-related. If the issue only appears with compilation enabled, use `--backend eager` to isolate whether the problem is specific to the Inductor backend. Check the initialization logs for messages from `CompilationConfig.__post_init__` in [`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py) to verify your settings are being applied correctly.