How to Configure Compilation Modes (Eager, CUDAGraph, Inductor) in vLLM

vLLM controls its compilation pipeline through three orthogonal settings in CompilationConfig: mode (selecting eager versus Inductor), cudagraph_mode (controlling CUDA Graph capture strategy), and backend (specifying the compiler implementation), which can be set via CLI flags or the Python API.

vLLM's inference engine uses a sophisticated compilation system to maximize GPU utilization for transformer models. Starting with the v1 architecture, the framework centralizes these optimizations in the CompilationConfig class, which is part of the top-level VllmConfig defined in vllm/config/compilation.py. Understanding how to configure compilation modes in vLLM allows you to balance debugging flexibility against maximum throughput for production deployments.

Understanding vLLM's Compilation Architecture

The Three Configuration Knobs

In vllm/config/compilation.py, the CompilationConfig dataclass defines three primary fields that determine runtime behavior:

  • mode: A CompilationMode enum that determines whether torch.compile is invoked and which pipeline executes.
  • cudagraph_mode: A CUDAGraphMode enum that controls how torch.cuda.CUDAGraph capture is applied.
  • backend: A string specifying the compiler backend, such as "inductor" or "eager".

These settings are orthogonal. For example, you can enable CUDAGraph capture while running in eager PyTorch mode, or use the Inductor backend without CUDA Graphs.

Compilation Modes Explained

The CompilationMode enum, defined at lines 35-48 of vllm/config/compilation.py, offers four distinct strategies:

NONE (Pure Eager Mode)

Setting mode=CompilationMode.NONE disables all torch.compile machinery. The model executes in standard PyTorch eager mode, which is essential for debugging or when running on platforms without Inductor support.

VLLM_COMPILE (Custom Inductor Backend)

mode=CompilationMode.VLLM_COMPILE activates vLLM's optimized compilation pipeline. According to the source code, this mode uses a piecewise compilation strategy with shape specialization and custom passes through the Inductor backend. This is the default for v1 and provides the best performance for most CUDA workloads.

STOCK_TORCH_COMPILE and DYNAMO_TRACE_ONCE

  • STOCK_TORCH_COMPILE: Uses PyTorch's default torch.compile pipeline without vLLM customizations.
  • DYNAMO_TRACE_ONCE: Runs Dynamo tracing a single time without re-traces, useful for dynamic shape scenarios where full compilation overhead is undesirable.

CUDAGraph Modes Explained

The CUDAGraphMode enum (lines 51-62 of vllm/config/compilation.py) determines how CUDA Graphs are captured and replayed. This is implemented in vllm/v1/worker/gpu_worker.py and vllm/v1/worker/gpu_ubatch_wrapper.py.

NONE

CUDAGraphMode.NONE disables graph capture entirely. Every forward pass executes kernels individually, which is necessary when debugging kernel launches or using custom operators incompatible with graph capture.

PIECEWISE

CUDAGraphMode.PIECEWISE captures only operations that are safe for CUDAGraph, splitting out attention operations. The UBatchWrapper class in vllm/v1/worker/gpu_ubatch_wrapper.py manages these per-size graph objects. This mode is required when using attention kernels that cannot be captured in a full graph.

FULL and FULL_DECODE_ONLY

  • FULL: Captures the entire forward pass in a single graph. This provides maximum speed for small, fixed-size batches but may cause out-of-memory errors with larger batch sizes.
  • FULL_DECODE_ONLY: Captures the full graph only for decode-only batches (no prefill), balancing throughput and memory.

FULL_AND_PIECEWISE (Default)

CUDAGraphMode.FULL_AND_PIECEWISE is the default for v1. As implemented in vllm/v1/worker/gpu_model_runner.py, this mode uses full capture for decode operations and piecewise capture for prefill or mixed batches, automatically selecting the optimal strategy per request type.

Configuration Methods

You can configure these settings via command-line arguments when using the OpenAI-compatible server, or programmatically through the Python API.

Command-Line Interface

The entry point in vllm/entrypoints/openai/api_server.py accepts flags that map directly to CompilationConfig fields:


# Default: VLLM_COMPILE with FULL_AND_PIECEWISE CUDAGraph

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-hf

# Pure eager mode (debugging)

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-hf \
    --compilation-mode none \
    --cudagraph-mode none

# Disable CUDAGraph but keep Inductor (useful for debugging graph-partition)

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-hf \
    --cudagraph-mode none

# Force full CUDAGraph capture (no piecewise)

python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-2-7b-hf \
    --cudagraph-mode full

Python API

For embedded use cases, instantiate CompilationConfig and pass it to VllmConfig:

from vllm import LLM
from vllm.config import VllmConfig, CompilationConfig, CompilationMode, CUDAGraphMode

# Eager mode configuration

eager_cfg = VllmConfig(
    model="meta-llama/Llama-2-7b-hf",
    compilation_config=CompilationConfig(
        mode=CompilationMode.NONE,
        cudagraph_mode=CUDAGraphMode.NONE,
    ),
)
llm_eager = LLM(vllm_config=eager_cfg)

# High-performance Inductor configuration

inductor_cfg = VllmConfig(
    model="meta-llama/Llama-2-7b-hf",
    compilation_config=CompilationConfig(
        mode=CompilationMode.VLLM_COMPILE,
        cudagraph_mode=CUDAGraphMode.FULL_AND_PIECEWISE,
        backend="inductor",
    ),
)
llm_inductor = LLM(vllm_config=inductor_cfg)

Advanced: Customizing Splitting Operations

When using CUDAGraphMode.PIECEWISE, vLLM automatically determines which operations to split out based on CompilationConfig._attention_ops. You can override this behavior by manually setting splitting_ops:

from vllm.config import CompilationConfig, CUDAGraphMode, CompilationMode

cfg = CompilationConfig(
    mode=CompilationMode.VLLM_COMPILE,
    cudagraph_mode=CUDAGraphMode.PIECEWISE,
    splitting_ops=["flash_attention", "custom_op"],  # Split these ops out of graphs

)

The logic that populates default splitting ops is implemented in set_splitting_ops_for_v1 at lines 972-1010 of vllm/config/compilation.py.

Summary

Configuring compilation modes in vLLM involves three orthogonal settings in CompilationConfig:

  • mode determines the torch.compile strategy: NONE for eager execution, VLLM_COMPILE for the optimized Inductor backend with custom passes, or STOCK_TORCH_COMPILE for standard PyTorch behavior.
  • cudagraph_mode controls CUDA Graph capture: NONE for debugging, PIECEWISE for safe partial capture, or FULL_AND_PIECEWISE (default) for maximum throughput via full decode graphs and piecewise prefill.
  • backend specifies the compiler implementation, with "inductor" delivering best performance on CUDA devices and "eager" aiding debugging.

Apply these via CLI flags (--compilation-mode, --cudagraph-mode), environment variables, or programmatically through VllmConfig and CompilationConfig to optimize inference for your specific hardware and debugging requirements.

Frequently Asked Questions

What is the default compilation mode in vLLM v1?

The v1 engine defaults to CompilationMode.VLLM_COMPILE with CUDAGraphMode.FULL_AND_PIECEWISE and the Inductor backend. This configuration, defined in vllm/config/compilation.py, provides optimal throughput for most CUDA workloads by combining custom Inductor passes with intelligent CUDA Graph capture.

When should I use Eager mode instead of Inductor?

Use CompilationMode.NONE (Eager mode) when debugging kernel launches, working with custom operators that are incompatible with torch.compile, or deploying on hardware without Inductor support. Eager mode executes PyTorch operations immediately without graph compilation, providing clearer stack traces at the cost of reduced inference speed.

Can I use CUDAGraph without torch.compile?

Yes. The cudagraph_mode setting is orthogonal to the mode setting. You can configure mode=CompilationMode.NONE and cudagraph_mode=CUDAGraphMode.FULL to capture and replay CUDA Graphs while running in pure eager PyTorch mode. This is useful when you want the runtime efficiency of graph replay without the compilation overhead of Inductor.

How do I debug compilation issues in vLLM?

Start by running with --compilation-mode none --cudagraph-mode none to establish a baseline in pure eager mode. If the issue persists, it is not compilation-related. If the issue only appears with compilation enabled, use --backend eager to isolate whether the problem is specific to the Inductor backend. Check the initialization logs for messages from CompilationConfig.__post_init__ in vllm/config/compilation.py to verify your settings are being applied correctly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →