How to Debug Memory Issues and OOM Errors in vLLM: A Complete Guide
Enable VLLM_DEBUG_WORKSPACE and use the memory_profiling context manager from vllm/utils/mem_utils.py to isolate whether OOM errors stem from workspace buffers, KV-cache, activation tensors, or non-torch allocations like NCCL.
vLLM aggressively optimizes GPU memory to maximize inference throughput, which can lead to out-of-memory (OOM) errors when requests exceed allocated buffers. Understanding how to debug memory issues and OOM errors in vLLM requires identifying whether the pressure comes from workspace resizes, KV-cache allocations, or temporary activation tensors. This guide walks through the specific source files and environment variables in the vllm-project/vllm repository that control memory behavior and debugging instrumentation.
Understanding vLLM Memory Components
Memory pressure in vLLM typically originates from four distinct areas. Knowing which component dominates your workload determines the correct fix.
| Component | Source File | OOM Trigger |
|---|---|---|
| Workspace manager | vllm/v1/worker/workspace.py |
Resizing the reusable buffer for intermediate tensors causes a temporary spike because resize_() allocates new memory before freeing the old. |
| KV-cache | vllm/v1/core/kv_cache_utils.py |
Grows linearly with max_model_len × num_kv_heads × head_dim; unbounded growth exhausts GPU memory. |
| Activation tensors | vllm/utils/mem_utils.py |
Temporary tensors created during the forward pass; measured by torch.cuda.memory_stats()["allocated_bytes.all.peak"]. |
| Non-torch allocations | vllm/utils/mem_utils.py |
NCCL buffers, FlashAttention workspace, and custom kernels tracked by MemorySnapshot.non_torch_memory. |
Enable Detailed Workspace Logging
The workspace manager in vllm/v1/worker/workspace.py pre-allocates a reusable buffer that resizes dynamically. When a batch grows, the resize operation can trigger an OOM spike.
Set the environment variable to log every resize operation:
export VLLM_DEBUG_WORKSPACE=1
When the engine grows the workspace, the log output shows the old size, new size, and micro-batch count:
[WORKSPACE DEBUG] Resized workspace from 'gpu_worker._run_one_step': 256.00 MB -> 512.00 MB (4 ubatches, total memory 2048.00 MB)
If you observe a sudden jump immediately before the OOM error, the workspace buffer is the culprit. Pre-allocate a larger initial workspace or reduce batch sizes to prevent mid-run resizes.
Profile Memory with the Built-in Context Manager
Wrap your engine initialization and inference calls with the memory_profiling context manager defined in vllm/utils/mem_utils.py to capture a detailed breakdown of memory usage.
from vllm.entrypoints.llm import LLM
from vllm.utils.mem_utils import MemorySnapshot, memory_profiling
baseline = MemorySnapshot() # Capture baseline before engine init
with memory_profiling(baseline, weights_memory=0) as prof:
llm = LLM(model="meta-llama/Llama-2-7b-chat-hf")
output = llm.generate(prompts=["Explain quantum computing."])
print(prof)
The MemoryProfilingResult object reports three critical metrics:
- non_kv_cache_memory: Total memory used outside the KV-cache.
- torch_peak_increase: Peak activation memory during the forward pass.
- non_torch_increase: NCCL buffers and kernel workspaces (FlashAttention, etc.).
Use these values to determine whether you need to shrink the KV-cache, reduce activation batch sizes, or limit non-torch buffers.
Tune KV-Cache Size
The KV-cache allocator in vllm/config/cache.py accepts two mutually exclusive parameters to control cache size.
Adjust gpu_memory_utilization
The gpu_memory_utilization parameter (default 0.9) specifies the fraction of free GPU memory that vLLM may use for the KV-cache and activations. Lower this value to leave more headroom for workspace spikes.
python -m vllm.entrypoints.llm \
--model meta-llama/Llama-2-70b-chat-hf \
--gpu-memory-utilization 0.6
Set kv_cache_memory_bytes Explicitly
For precise control, pass kv_cache_memory_bytes (in bytes) to set a fixed cache size. This overrides gpu_memory_utilization.
llm = LLM(
model="meta-llama/Llama-2-70b-chat-hf",
kv_cache_memory_bytes=12 * 1024**3, # 12 GiB per GPU
)
Reduce Model Context Length
The KV-cache memory requirement grows linearly with max_model_len. If you encounter OOM errors only on long prompts, cap the maximum sequence length.
llm = LLM(
model="meta-llama/Llama-2-7b-chat-hf",
max_model_len=4096, # Default may be 8192 or higher
)
Alternatively, set the VLLM_MAX_MODEL_LEN environment variable defined in vllm/envs.py to enforce a global limit across all model instances.
Diagnose Non-Torch Memory
When memory_profiling reports a large non_torch_increase, the issue lies outside PyTorch’s allocator. Two common sources are NCCL buffers and FlashAttention workspace.
Limit FlashAttention Workspace
Set VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE in vllm/envs.py to reduce the buffer size for FlashAttention kernels.
export VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE=$((256 * 1024 * 1024)) # 256 MiB
Constrain NCCL Buffers
On multi-GPU deployments, NCCL buffers can exhaust per-GPU memory. Reduce NCCL_BUFFSIZE or lower gpu_memory_utilization to accommodate these non-torch allocations.
Common OOM Scenarios and Fixes
| Symptom | Root Cause | Solution |
|---|---|---|
| OOM only on the first request | Workspace resize creates a temporary allocation spike before freeing old memory (see vllm/v1/worker/workspace.py). |
Increase initial workspace size or lower gpu_memory_utilization. |
| OOM after many requests | KV-cache grows unbounded without memory limits. | Set kv_cache_memory_bytes or reduce gpu_memory_utilization. |
| OOM on multi-GPU but not single-GPU | NCCL buffers exceed per-GPU budget. | Reduce NCCL_BUFFSIZE or gpu_memory_utilization. |
| OOM despite low utilization | FlashAttention workspace dominates memory. | Decrease VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE. |
| OOM on CPU-offload runs | cpu_offload_gb set too low, forcing weights onto GPU. |
Increase cpu_offload_gb or adjust gpu_memory_utilization. |
Complete Debug Script
Use this script to reproduce and isolate OOM errors systematically.
#!/usr/bin/env python3
import os
from vllm.entrypoints.llm import LLM
from vllm.utils.mem_utils import MemorySnapshot, memory_profiling
# Enable workspace logging
os.environ["VLLM_DEBUG_WORKSPACE"] = "1"
os.environ["VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE"] = str(256 * 1024 * 1024)
baseline = MemorySnapshot()
with memory_profiling(baseline) as prof:
llm = LLM(
model="meta-llama/Llama-2-7b-chat-hf",
gpu_memory_utilization=0.6,
max_model_len=4096,
kv_cache_memory_bytes=None,
)
out = llm.generate(prompts=["Explain quantum computing in 50 words."])
print("=== Memory Profile ===")
print(prof)
Compare the printed profile against the workspace debug output to identify whether the OOM occurs during workspace resize, KV-cache allocation, or activation peak.
Summary
- Enable
VLLM_DEBUG_WORKSPACEto catch workspace resize spikes invllm/v1/worker/workspace.py. - Use
memory_profilingfromvllm/utils/mem_utils.pyto separate activation, KV-cache, and non-torch memory usage. - Adjust KV-cache via
gpu_memory_utilizationor a fixedkv_cache_memory_bytesvalue. - Limit sequence length with
max_model_lenif long prompts trigger OOM. - Shrink non-torch buffers by setting
VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZEand tuning NCCL environment variables.
Frequently Asked Questions
What causes sudden OOM errors in vLLM after the engine starts running?
Sudden OOM errors typically occur when the workspace manager in vllm/v1/worker/workspace.py resizes its buffer to accommodate a larger batch. The resize_() operation allocates new memory before releasing the old buffer, creating a temporary spike. Enable VLLM_DEBUG_WORKSPACE to confirm this, then pre-allocate a larger workspace or reduce gpu_memory_utilization to prevent the resize.
How do I distinguish between KV-cache OOM and activation OOM?
Wrap your inference code with the memory_profiling context manager from vllm/utils/mem_utils.py. If non_kv_cache_memory remains stable but the process still OOMs, the KV-cache is too large. If torch_peak_increase spikes during generation, activation tensors are the bottleneck. Reduce max_model_len for cache issues or batch size for activation issues.
Can I fix OOM errors without reducing batch size?
Yes. Instead of reducing batch size, set a fixed kv_cache_memory_bytes to cap cache growth, lower max_model_len to reduce per-token cache requirements, or decrease VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE to limit non-torch allocations. These adjustments preserve throughput while eliminating memory pressure.
Why does vLLM OOM on multi-GPU setups when it works on a single GPU?
Multi-GPU deployments allocate NCCL communication buffers that do not appear in PyTorch’s memory statistics. These non-torch allocations, tracked in MemorySnapshot.non_torch_memory, can exhaust GPU memory even when gpu_memory_utilization is set low. Reduce NCCL_BUFFSIZE or further lower gpu_memory_utilization to accommodate these buffers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →