How CUDA Graph Optimization Improves Inference Performance in vLLM
CUDA graph optimization in vLLM eliminates per-step kernel launch overhead by recording the model's forward pass as a static execution trace, delivering 2×–5× speed-ups for decode operations and 1.5×–3× improvements for prefilling.
CUDA graph optimization transforms the dynamic, kernel-heavy execution of large language models into pre-recorded computational graphs. In the vLLM inference engine, this technique captures the entire forward pass during warmup and replays it for subsequent identical batch shapes, removing CPU-side bottlenecks from critical inference paths. Understanding how vLLM implements this optimization reveals why it achieves substantial throughput gains for production serving workloads.
The Three Stages of CUDA Graph Optimization
Stage 1: Graph Capture
During the initial warmup phase, vLLM records the complete forward computation—including all CUDA kernels, memory copies, and synchronization points—into a torch.cuda.CUDAGraph object. The capture_model method in vllm/v1/worker/gpu_model_runner.py orchestrates this process by freezing Python's garbage collector and executing the forward pass within a graph capture context. This stage occurs once per unique batch descriptor (token count and shape), ensuring that the static trace accurately represents the computation for that specific configuration.
Stage 2: Graph Replay
For subsequent inference requests matching a previously captured batch descriptor, vLLM bypasses individual kernel launches entirely. The CUDAGraphWrapper.__call__ method in vllm/compilation/cuda_graph.py retrieves the cached CUDAGraphEntry, validates input memory addresses, and invokes entry.cudagraph.replay(). This replay executes the pre-recorded kernel sequence with minimal CPU overhead, effectively reducing the launch latency of dozens of transformer kernels—matmuls, attention operations, and layer normalizations—to a single operation.
Stage 3: Memory Pool and Garbage Collection Management
To ensure deterministic memory usage during capture, vLLM disables Python's garbage collector and allocates all intermediate buffers from a dedicated CUDA graph memory pool using set_graph_pool_id. As implemented in vllm/compilation/cuda_graph.py, this prevents per-step cudaMalloc/cudaFree calls during replay, eliminating memory fragmentation and CPU-GPU synchronization points that typically stall inference.
Performance Benefits of CUDA Graph Optimization
The optimization yields measurable improvements across several dimensions:
- Kernel Launch Reduction: Recording the full transformer forward pass once removes the CPU overhead of launching individual kernels for every decode step. Instead of dispatching dozens of separate operations, the system executes a single replay command.
- Stable Memory Allocation: All tensors are allocated during the capture phase and persist throughout inference. This eliminates dynamic memory management overhead and prevents address space fragmentation common in long-running inference services.
- Deterministic Latency: Because the execution path is identical for every replay, latency percentiles tighten significantly. This consistency proves critical for production serving where p99 latency requirements are strict.
- Flexible Capture Modes: The
CUDAGraphModeconfiguration enum invllm/config/compilation.pysupports FULL (entire forward pass), PIECEWISE (per-layer subgraphs), and FULL_AND_PIECEWISE modes, enabling optimization even with dynamic features like LoRA adapters.
In production deployments, users observe 2×–5× speed-ups for decode-only (token-by-token) inference and 1.5×–3× improvements for prefilling large batches, depending on GPU generation and model size.
Implementation in vLLM Source Code
The optimization spans several critical components within the vLLM codebase:
vllm/v1/worker/gpu_model_runner.py: Containscapture_model, which manages thetorch.cuda.graphcapture lifecycle and coordinates warmup runs.vllm/compilation/cuda_graph.py: ImplementsCUDAGraphWrapperfor recording, caching, and replaying graphs, plus memory pool management viaset_graph_pool_id.vllm/config/compilation.py: DefinesCUDAGraphMode(FULL, PIECEWISE, NONE) and parsescudagraph_capture_sizesconfiguration.vllm/v1/worker/gpu_worker.py: Houses_warmup_and_capture, which performskernel_warmupbefore capture to avoid JIT compilation delays.
Configuring CUDA Graphs for Your Workload
Enabling Full Graph Capture
For maximum performance with static model architectures, configure vLLM to capture the complete forward pass:
from vllm import LLM, SamplingParams
# Enable full CUDA-graph capture for specified batch shapes
llm = LLM(
model="meta-llama/Llama-2-70b-chat-hf",
tensor_parallel_size=4,
cudagraph_mode="FULL",
cudagraph_capture_sizes=[64, 128, 256] # Pre-capture these batch sizes
)
sampling_params = SamplingParams(temperature=0.8, max_tokens=50)
# First call triggers capture for the request shape
output = llm.generate("Explain CUDA graphs", sampling_params)
Piecewise Capture for Dynamic Workloads
When using LoRA adapters or mixed-precision requirements, employ piecewise mode to capture subgraphs separately:
llm = LLM(
model="my/lora-enabled-model",
lora_weights="my/lora",
cudagraph_mode="FULL_AND_PIECEWISE", # Separate graphs for base and LoRA paths
)
# Base model inference triggers the first graph capture
llm.generate("Write a poem", SamplingParams(max_tokens=30))
# LoRA-enabled inference reuses the adapter-specific graph
output = llm.generate(
"Summarise the following text",
SamplingParams(max_tokens=20),
lora_adapter_names=["my/lora"]
)
The CUDAGraphWrapper maintains separate CUDAGraphEntry objects keyed by both batch descriptor and LoRA configuration, automatically selecting the appropriate graph during __call__ execution.
Summary
- CUDA graph optimization in vLLM records the model forward pass as static execution traces during warmup, eliminating per-step kernel launch overhead.
- The three-stage process involves capture (recording), replay (execution), and memory pool management to prevent allocation overhead.
- Performance gains range from 2×–5× for decode operations and 1.5×–3× for prefilling, driven by reduced CPU overhead and stable memory usage.
- Configuration flexibility via
CUDAGraphModesupports FULL, PIECEWISE, and FULL_AND_PIECEWISE strategies to accommodate LoRA adapters and dynamic batching. - Key implementation files include
vllm/compilation/cuda_graph.pyfor the wrapper logic andvllm/v1/worker/gpu_model_runner.pyfor capture orchestration.
Frequently Asked Questions
What is the warm-up cost when enabling CUDA graphs in vLLM?
The warm-up cost occurs only once per unique batch shape defined by token count and configuration. During initialization, vLLM executes a dummy forward pass through kernel_warmup in vllm/v1/worker/gpu_worker.py to trigger JIT compilation before capture, ensuring that subsequent replays execute without compilation delays. This cost amortizes over thousands of inference steps.
Does CUDA graph optimization work with LoRA adapters?
Yes. When using cudagraph_mode="FULL_AND_PIECEWISE", vLLM captures separate graphs for the base model and each LoRA configuration. The CUDAGraphWrapper stores entries keyed by both the batch descriptor and LoRA adapter names in vllm/compilation/cuda_graph.py, allowing the system to replay the appropriate graph when adapters are dynamically loaded or unloaded during serving.
How does vLLM handle memory allocation with CUDA graphs enabled?
During capture, vLLM disables Python's garbage collector and assigns a dedicated graph memory pool ID using set_graph_pool_id in vllm/compilation/cuda_graph.py. This ensures all intermediate tensors are allocated once and persist in GPU memory, eliminating cudaMalloc/cudaFree calls during inference replay and preventing memory fragmentation.
Can I use CUDA graphs with dynamic batch sizes?
vLLM supports dynamic batching by capturing multiple graphs for specific sizes defined in cudagraph_capture_sizes. If a request arrives with an unseen shape, the system either captures a new graph (if the size is in the configuration list) or falls back to eager mode execution. For optimal performance, pre-configure capture sizes that match your expected traffic distribution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →