How to Configure Compilation Modes (Eager, CUDAGraph, Inductor) in vLLM
vLLM controls its compilation pipeline through three orthogonal settings in CompilationConfig: mode (selecting eager versus Inductor), cudagraph_mode (controlling CUDA Graph capture strategy), and backend (specifying the compiler implementation), which can be set via CLI flags or the Python API.
vLLM's inference engine uses a sophisticated compilation system to maximize GPU utilization for transformer models. Starting with the v1 architecture, the framework centralizes these optimizations in the CompilationConfig class, which is part of the top-level VllmConfig defined in vllm/config/compilation.py. Understanding how to configure compilation modes in vLLM allows you to balance debugging flexibility against maximum throughput for production deployments.
Understanding vLLM's Compilation Architecture
The Three Configuration Knobs
In vllm/config/compilation.py, the CompilationConfig dataclass defines three primary fields that determine runtime behavior:
mode: ACompilationModeenum that determines whethertorch.compileis invoked and which pipeline executes.cudagraph_mode: ACUDAGraphModeenum that controls howtorch.cuda.CUDAGraphcapture is applied.backend: A string specifying the compiler backend, such as"inductor"or"eager".
These settings are orthogonal. For example, you can enable CUDAGraph capture while running in eager PyTorch mode, or use the Inductor backend without CUDA Graphs.
Compilation Modes Explained
The CompilationMode enum, defined at lines 35-48 of vllm/config/compilation.py, offers four distinct strategies:
NONE (Pure Eager Mode)
Setting mode=CompilationMode.NONE disables all torch.compile machinery. The model executes in standard PyTorch eager mode, which is essential for debugging or when running on platforms without Inductor support.
VLLM_COMPILE (Custom Inductor Backend)
mode=CompilationMode.VLLM_COMPILE activates vLLM's optimized compilation pipeline. According to the source code, this mode uses a piecewise compilation strategy with shape specialization and custom passes through the Inductor backend. This is the default for v1 and provides the best performance for most CUDA workloads.
STOCK_TORCH_COMPILE and DYNAMO_TRACE_ONCE
STOCK_TORCH_COMPILE: Uses PyTorch's defaulttorch.compilepipeline without vLLM customizations.DYNAMO_TRACE_ONCE: Runs Dynamo tracing a single time without re-traces, useful for dynamic shape scenarios where full compilation overhead is undesirable.
CUDAGraph Modes Explained
The CUDAGraphMode enum (lines 51-62 of vllm/config/compilation.py) determines how CUDA Graphs are captured and replayed. This is implemented in vllm/v1/worker/gpu_worker.py and vllm/v1/worker/gpu_ubatch_wrapper.py.
NONE
CUDAGraphMode.NONE disables graph capture entirely. Every forward pass executes kernels individually, which is necessary when debugging kernel launches or using custom operators incompatible with graph capture.
PIECEWISE
CUDAGraphMode.PIECEWISE captures only operations that are safe for CUDAGraph, splitting out attention operations. The UBatchWrapper class in vllm/v1/worker/gpu_ubatch_wrapper.py manages these per-size graph objects. This mode is required when using attention kernels that cannot be captured in a full graph.
FULL and FULL_DECODE_ONLY
FULL: Captures the entire forward pass in a single graph. This provides maximum speed for small, fixed-size batches but may cause out-of-memory errors with larger batch sizes.FULL_DECODE_ONLY: Captures the full graph only for decode-only batches (no prefill), balancing throughput and memory.
FULL_AND_PIECEWISE (Default)
CUDAGraphMode.FULL_AND_PIECEWISE is the default for v1. As implemented in vllm/v1/worker/gpu_model_runner.py, this mode uses full capture for decode operations and piecewise capture for prefill or mixed batches, automatically selecting the optimal strategy per request type.
Configuration Methods
You can configure these settings via command-line arguments when using the OpenAI-compatible server, or programmatically through the Python API.
Command-Line Interface
The entry point in vllm/entrypoints/openai/api_server.py accepts flags that map directly to CompilationConfig fields:
# Default: VLLM_COMPILE with FULL_AND_PIECEWISE CUDAGraph
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-7b-hf
# Pure eager mode (debugging)
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-7b-hf \
--compilation-mode none \
--cudagraph-mode none
# Disable CUDAGraph but keep Inductor (useful for debugging graph-partition)
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-7b-hf \
--cudagraph-mode none
# Force full CUDAGraph capture (no piecewise)
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-7b-hf \
--cudagraph-mode full
Python API
For embedded use cases, instantiate CompilationConfig and pass it to VllmConfig:
from vllm import LLM
from vllm.config import VllmConfig, CompilationConfig, CompilationMode, CUDAGraphMode
# Eager mode configuration
eager_cfg = VllmConfig(
model="meta-llama/Llama-2-7b-hf",
compilation_config=CompilationConfig(
mode=CompilationMode.NONE,
cudagraph_mode=CUDAGraphMode.NONE,
),
)
llm_eager = LLM(vllm_config=eager_cfg)
# High-performance Inductor configuration
inductor_cfg = VllmConfig(
model="meta-llama/Llama-2-7b-hf",
compilation_config=CompilationConfig(
mode=CompilationMode.VLLM_COMPILE,
cudagraph_mode=CUDAGraphMode.FULL_AND_PIECEWISE,
backend="inductor",
),
)
llm_inductor = LLM(vllm_config=inductor_cfg)
Advanced: Customizing Splitting Operations
When using CUDAGraphMode.PIECEWISE, vLLM automatically determines which operations to split out based on CompilationConfig._attention_ops. You can override this behavior by manually setting splitting_ops:
from vllm.config import CompilationConfig, CUDAGraphMode, CompilationMode
cfg = CompilationConfig(
mode=CompilationMode.VLLM_COMPILE,
cudagraph_mode=CUDAGraphMode.PIECEWISE,
splitting_ops=["flash_attention", "custom_op"], # Split these ops out of graphs
)
The logic that populates default splitting ops is implemented in set_splitting_ops_for_v1 at lines 972-1010 of vllm/config/compilation.py.
Summary
Configuring compilation modes in vLLM involves three orthogonal settings in CompilationConfig:
modedetermines the torch.compile strategy:NONEfor eager execution,VLLM_COMPILEfor the optimized Inductor backend with custom passes, orSTOCK_TORCH_COMPILEfor standard PyTorch behavior.cudagraph_modecontrols CUDA Graph capture:NONEfor debugging,PIECEWISEfor safe partial capture, orFULL_AND_PIECEWISE(default) for maximum throughput via full decode graphs and piecewise prefill.backendspecifies the compiler implementation, with"inductor"delivering best performance on CUDA devices and"eager"aiding debugging.
Apply these via CLI flags (--compilation-mode, --cudagraph-mode), environment variables, or programmatically through VllmConfig and CompilationConfig to optimize inference for your specific hardware and debugging requirements.
Frequently Asked Questions
What is the default compilation mode in vLLM v1?
The v1 engine defaults to CompilationMode.VLLM_COMPILE with CUDAGraphMode.FULL_AND_PIECEWISE and the Inductor backend. This configuration, defined in vllm/config/compilation.py, provides optimal throughput for most CUDA workloads by combining custom Inductor passes with intelligent CUDA Graph capture.
When should I use Eager mode instead of Inductor?
Use CompilationMode.NONE (Eager mode) when debugging kernel launches, working with custom operators that are incompatible with torch.compile, or deploying on hardware without Inductor support. Eager mode executes PyTorch operations immediately without graph compilation, providing clearer stack traces at the cost of reduced inference speed.
Can I use CUDAGraph without torch.compile?
Yes. The cudagraph_mode setting is orthogonal to the mode setting. You can configure mode=CompilationMode.NONE and cudagraph_mode=CUDAGraphMode.FULL to capture and replay CUDA Graphs while running in pure eager PyTorch mode. This is useful when you want the runtime efficiency of graph replay without the compilation overhead of Inductor.
How do I debug compilation issues in vLLM?
Start by running with --compilation-mode none --cudagraph-mode none to establish a baseline in pure eager mode. If the issue persists, it is not compilation-related. If the issue only appears with compilation enabled, use --backend eager to isolate whether the problem is specific to the Inductor backend. Check the initialization logs for messages from CompilationConfig.__post_init__ in vllm/config/compilation.py to verify your settings are being applied correctly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →