# How CUDA Graph Optimization Improves Inference Performance in vLLM

> Discover how CUDA graph optimization in vLLM dramatically boosts inference speed. Experience 2x-5x faster decode and 1.5x-3x faster prefill by eliminating kernel launch overhead.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: performance
- Published: 2026-03-03

---

**CUDA graph optimization in vLLM eliminates per-step kernel launch overhead by recording the model's forward pass as a static execution trace, delivering 2×–5× speed-ups for decode operations and 1.5×–3× improvements for prefilling.**

CUDA graph optimization transforms the dynamic, kernel-heavy execution of large language models into pre-recorded computational graphs. In the vLLM inference engine, this technique captures the entire forward pass during warmup and replays it for subsequent identical batch shapes, removing CPU-side bottlenecks from critical inference paths. Understanding how vLLM implements this optimization reveals why it achieves substantial throughput gains for production serving workloads.

## The Three Stages of CUDA Graph Optimization

### Stage 1: Graph Capture

During the initial warmup phase, vLLM records the complete forward computation—including all CUDA kernels, memory copies, and synchronization points—into a `torch.cuda.CUDAGraph` object. The `capture_model` method in [`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py) orchestrates this process by freezing Python's garbage collector and executing the forward pass within a graph capture context. This stage occurs once per unique **batch descriptor** (token count and shape), ensuring that the static trace accurately represents the computation for that specific configuration.

### Stage 2: Graph Replay

For subsequent inference requests matching a previously captured batch descriptor, vLLM bypasses individual kernel launches entirely. The `CUDAGraphWrapper.__call__` method in [`vllm/compilation/cuda_graph.py`](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/cuda_graph.py) retrieves the cached `CUDAGraphEntry`, validates input memory addresses, and invokes `entry.cudagraph.replay()`. This replay executes the pre-recorded kernel sequence with minimal CPU overhead, effectively reducing the launch latency of dozens of transformer kernels—matmuls, attention operations, and layer normalizations—to a single operation.

### Stage 3: Memory Pool and Garbage Collection Management

To ensure deterministic memory usage during capture, vLLM disables Python's garbage collector and allocates all intermediate buffers from a dedicated CUDA graph memory pool using `set_graph_pool_id`. As implemented in [`vllm/compilation/cuda_graph.py`](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/cuda_graph.py), this prevents per-step `cudaMalloc`/`cudaFree` calls during replay, eliminating memory fragmentation and CPU-GPU synchronization points that typically stall inference.

## Performance Benefits of CUDA Graph Optimization

The optimization yields measurable improvements across several dimensions:

- **Kernel Launch Reduction**: Recording the full transformer forward pass once removes the CPU overhead of launching individual kernels for every decode step. Instead of dispatching dozens of separate operations, the system executes a single replay command.
- **Stable Memory Allocation**: All tensors are allocated during the capture phase and persist throughout inference. This eliminates dynamic memory management overhead and prevents address space fragmentation common in long-running inference services.
- **Deterministic Latency**: Because the execution path is identical for every replay, latency percentiles tighten significantly. This consistency proves critical for production serving where p99 latency requirements are strict.
- **Flexible Capture Modes**: The `CUDAGraphMode` configuration enum in [`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py) supports **FULL** (entire forward pass), **PIECEWISE** (per-layer subgraphs), and **FULL_AND_PIECEWISE** modes, enabling optimization even with dynamic features like LoRA adapters.

In production deployments, users observe **2×–5×** speed-ups for decode-only (token-by-token) inference and **1.5×–3×** improvements for prefilling large batches, depending on GPU generation and model size.

## Implementation in vLLM Source Code

The optimization spans several critical components within the vLLM codebase:

- **[`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py)**: Contains `capture_model`, which manages the `torch.cuda.graph` capture lifecycle and coordinates warmup runs.
- **[`vllm/compilation/cuda_graph.py`](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/cuda_graph.py)**: Implements `CUDAGraphWrapper` for recording, caching, and replaying graphs, plus memory pool management via `set_graph_pool_id`.
- **[`vllm/config/compilation.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/compilation.py)**: Defines `CUDAGraphMode` (FULL, PIECEWISE, NONE) and parses `cudagraph_capture_sizes` configuration.
- **[`vllm/v1/worker/gpu_worker.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py)**: Houses `_warmup_and_capture`, which performs `kernel_warmup` before capture to avoid JIT compilation delays.

## Configuring CUDA Graphs for Your Workload

### Enabling Full Graph Capture

For maximum performance with static model architectures, configure vLLM to capture the complete forward pass:

```python
from vllm import LLM, SamplingParams

# Enable full CUDA-graph capture for specified batch shapes

llm = LLM(
    model="meta-llama/Llama-2-70b-chat-hf",
    tensor_parallel_size=4,
    cudagraph_mode="FULL",
    cudagraph_capture_sizes=[64, 128, 256]  # Pre-capture these batch sizes

)

sampling_params = SamplingParams(temperature=0.8, max_tokens=50)

# First call triggers capture for the request shape

output = llm.generate("Explain CUDA graphs", sampling_params)

```

### Piecewise Capture for Dynamic Workloads

When using LoRA adapters or mixed-precision requirements, employ piecewise mode to capture subgraphs separately:

```python
llm = LLM(
    model="my/lora-enabled-model",
    lora_weights="my/lora",
    cudagraph_mode="FULL_AND_PIECEWISE",  # Separate graphs for base and LoRA paths

)

# Base model inference triggers the first graph capture

llm.generate("Write a poem", SamplingParams(max_tokens=30))

# LoRA-enabled inference reuses the adapter-specific graph

output = llm.generate(
    "Summarise the following text",
    SamplingParams(max_tokens=20),
    lora_adapter_names=["my/lora"]
)

```

The `CUDAGraphWrapper` maintains separate `CUDAGraphEntry` objects keyed by both batch descriptor and LoRA configuration, automatically selecting the appropriate graph during `__call__` execution.

## Summary

- **CUDA graph optimization** in vLLM records the model forward pass as static execution traces during warmup, eliminating per-step kernel launch overhead.
- The **three-stage process** involves capture (recording), replay (execution), and memory pool management to prevent allocation overhead.
- **Performance gains** range from 2×–5× for decode operations and 1.5×–3× for prefilling, driven by reduced CPU overhead and stable memory usage.
- **Configuration flexibility** via `CUDAGraphMode` supports FULL, PIECEWISE, and FULL_AND_PIECEWISE strategies to accommodate LoRA adapters and dynamic batching.
- Key implementation files include [`vllm/compilation/cuda_graph.py`](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/cuda_graph.py) for the wrapper logic and [`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py) for capture orchestration.

## Frequently Asked Questions

### What is the warm-up cost when enabling CUDA graphs in vLLM?

The warm-up cost occurs only once per unique batch shape defined by token count and configuration. During initialization, vLLM executes a dummy forward pass through `kernel_warmup` in [`vllm/v1/worker/gpu_worker.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py) to trigger JIT compilation before capture, ensuring that subsequent replays execute without compilation delays. This cost amortizes over thousands of inference steps.

### Does CUDA graph optimization work with LoRA adapters?

Yes. When using `cudagraph_mode="FULL_AND_PIECEWISE"`, vLLM captures separate graphs for the base model and each LoRA configuration. The `CUDAGraphWrapper` stores entries keyed by both the batch descriptor and LoRA adapter names in [`vllm/compilation/cuda_graph.py`](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/cuda_graph.py), allowing the system to replay the appropriate graph when adapters are dynamically loaded or unloaded during serving.

### How does vLLM handle memory allocation with CUDA graphs enabled?

During capture, vLLM disables Python's garbage collector and assigns a dedicated graph memory pool ID using `set_graph_pool_id` in [`vllm/compilation/cuda_graph.py`](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/cuda_graph.py). This ensures all intermediate tensors are allocated once and persist in GPU memory, eliminating `cudaMalloc`/`cudaFree` calls during inference replay and preventing memory fragmentation.

### Can I use CUDA graphs with dynamic batch sizes?

vLLM supports dynamic batching by capturing multiple graphs for specific sizes defined in `cudagraph_capture_sizes`. If a request arrives with an unseen shape, the system either captures a new graph (if the size is in the configuration list) or falls back to eager mode execution. For optimal performance, pre-configure capture sizes that match your expected traffic distribution.