# How to Debug Memory Issues and OOM Errors in vLLM: A Complete Guide

> Debug vLLM memory issues and OOM errors with our guide. Isolate problems in workspace buffers, KV-cache, or activations using memory profiling tools.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**Enable `VLLM_DEBUG_WORKSPACE` and use the `memory_profiling` context manager from [`vllm/utils/mem_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/utils/mem_utils.py) to isolate whether OOM errors stem from workspace buffers, KV-cache, activation tensors, or non-torch allocations like NCCL.**

vLLM aggressively optimizes GPU memory to maximize inference throughput, which can lead to out-of-memory (OOM) errors when requests exceed allocated buffers. Understanding how to debug memory issues and OOM errors in vLLM requires identifying whether the pressure comes from workspace resizes, KV-cache allocations, or temporary activation tensors. This guide walks through the specific source files and environment variables in the vllm-project/vllm repository that control memory behavior and debugging instrumentation.

## Understanding vLLM Memory Components

Memory pressure in vLLM typically originates from four distinct areas. Knowing which component dominates your workload determines the correct fix.

| Component | Source File | OOM Trigger |
|-----------|-------------|-------------|
| **Workspace manager** | [`vllm/v1/worker/workspace.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/workspace.py) | Resizing the reusable buffer for intermediate tensors causes a temporary spike because `resize_()` allocates new memory before freeing the old. |
| **KV-cache** | [`vllm/v1/core/kv_cache_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/kv_cache_utils.py) | Grows linearly with `max_model_len × num_kv_heads × head_dim`; unbounded growth exhausts GPU memory. |
| **Activation tensors** | [`vllm/utils/mem_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/utils/mem_utils.py) | Temporary tensors created during the forward pass; measured by `torch.cuda.memory_stats()["allocated_bytes.all.peak"]`. |
| **Non-torch allocations** | [`vllm/utils/mem_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/utils/mem_utils.py) | NCCL buffers, FlashAttention workspace, and custom kernels tracked by `MemorySnapshot.non_torch_memory`. |

## Enable Detailed Workspace Logging

The workspace manager in [`vllm/v1/worker/workspace.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/workspace.py) pre-allocates a reusable buffer that resizes dynamically. When a batch grows, the resize operation can trigger an OOM spike.

Set the environment variable to log every resize operation:

```bash
export VLLM_DEBUG_WORKSPACE=1

```

When the engine grows the workspace, the log output shows the old size, new size, and micro-batch count:

```text
[WORKSPACE DEBUG] Resized workspace from 'gpu_worker._run_one_step': 256.00 MB -> 512.00 MB (4 ubatches, total memory 2048.00 MB)

```

If you observe a sudden jump immediately before the OOM error, the workspace buffer is the culprit. Pre-allocate a larger initial workspace or reduce batch sizes to prevent mid-run resizes.

## Profile Memory with the Built-in Context Manager

Wrap your engine initialization and inference calls with the `memory_profiling` context manager defined in [`vllm/utils/mem_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/utils/mem_utils.py) to capture a detailed breakdown of memory usage.

```python
from vllm.entrypoints.llm import LLM
from vllm.utils.mem_utils import MemorySnapshot, memory_profiling

baseline = MemorySnapshot()  # Capture baseline before engine init

with memory_profiling(baseline, weights_memory=0) as prof:
    llm = LLM(model="meta-llama/Llama-2-7b-chat-hf")
    output = llm.generate(prompts=["Explain quantum computing."])

print(prof)

```

The `MemoryProfilingResult` object reports three critical metrics:

- **non_kv_cache_memory**: Total memory used outside the KV-cache.
- **torch_peak_increase**: Peak activation memory during the forward pass.
- **non_torch_increase**: NCCL buffers and kernel workspaces (FlashAttention, etc.).

Use these values to determine whether you need to shrink the KV-cache, reduce activation batch sizes, or limit non-torch buffers.

## Tune KV-Cache Size

The KV-cache allocator in [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py) accepts two mutually exclusive parameters to control cache size.

### Adjust gpu_memory_utilization

The `gpu_memory_utilization` parameter (default `0.9`) specifies the fraction of free GPU memory that vLLM may use for the KV-cache and activations. Lower this value to leave more headroom for workspace spikes.

```bash
python -m vllm.entrypoints.llm \
    --model meta-llama/Llama-2-70b-chat-hf \
    --gpu-memory-utilization 0.6

```

### Set kv_cache_memory_bytes Explicitly

For precise control, pass `kv_cache_memory_bytes` (in bytes) to set a fixed cache size. This overrides `gpu_memory_utilization`.

```python
llm = LLM(
    model="meta-llama/Llama-2-70b-chat-hf",
    kv_cache_memory_bytes=12 * 1024**3,  # 12 GiB per GPU

)

```

## Reduce Model Context Length

The KV-cache memory requirement grows linearly with `max_model_len`. If you encounter OOM errors only on long prompts, cap the maximum sequence length.

```python
llm = LLM(
    model="meta-llama/Llama-2-7b-chat-hf",
    max_model_len=4096,  # Default may be 8192 or higher

)

```

Alternatively, set the `VLLM_MAX_MODEL_LEN` environment variable defined in [`vllm/envs.py`](https://github.com/vllm-project/vllm/blob/main/vllm/envs.py) to enforce a global limit across all model instances.

## Diagnose Non-Torch Memory

When `memory_profiling` reports a large `non_torch_increase`, the issue lies outside PyTorch’s allocator. Two common sources are NCCL buffers and FlashAttention workspace.

### Limit FlashAttention Workspace

Set `VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE` in [`vllm/envs.py`](https://github.com/vllm-project/vllm/blob/main/vllm/envs.py) to reduce the buffer size for FlashAttention kernels.

```bash
export VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE=$((256 * 1024 * 1024))  # 256 MiB

```

### Constrain NCCL Buffers

On multi-GPU deployments, NCCL buffers can exhaust per-GPU memory. Reduce `NCCL_BUFFSIZE` or lower `gpu_memory_utilization` to accommodate these non-torch allocations.

## Common OOM Scenarios and Fixes

| Symptom | Root Cause | Solution |
|---------|------------|----------|
| OOM only on the first request | Workspace resize creates a temporary allocation spike before freeing old memory (see [`vllm/v1/worker/workspace.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/workspace.py)). | Increase initial workspace size or lower `gpu_memory_utilization`. |
| OOM after many requests | KV-cache grows unbounded without memory limits. | Set `kv_cache_memory_bytes` or reduce `gpu_memory_utilization`. |
| OOM on multi-GPU but not single-GPU | NCCL buffers exceed per-GPU budget. | Reduce `NCCL_BUFFSIZE` or `gpu_memory_utilization`. |
| OOM despite low utilization | FlashAttention workspace dominates memory. | Decrease `VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE`. |
| OOM on CPU-offload runs | `cpu_offload_gb` set too low, forcing weights onto GPU. | Increase `cpu_offload_gb` or adjust `gpu_memory_utilization`. |

## Complete Debug Script

Use this script to reproduce and isolate OOM errors systematically.

```python
#!/usr/bin/env python3
import os
from vllm.entrypoints.llm import LLM
from vllm.utils.mem_utils import MemorySnapshot, memory_profiling

# Enable workspace logging

os.environ["VLLM_DEBUG_WORKSPACE"] = "1"
os.environ["VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE"] = str(256 * 1024 * 1024)

baseline = MemorySnapshot()

with memory_profiling(baseline) as prof:
    llm = LLM(
        model="meta-llama/Llama-2-7b-chat-hf",
        gpu_memory_utilization=0.6,
        max_model_len=4096,
        kv_cache_memory_bytes=None,
    )
    out = llm.generate(prompts=["Explain quantum computing in 50 words."])

print("=== Memory Profile ===")
print(prof)

```

Compare the printed profile against the workspace debug output to identify whether the OOM occurs during workspace resize, KV-cache allocation, or activation peak.

## Summary

- **Enable `VLLM_DEBUG_WORKSPACE`** to catch workspace resize spikes in [`vllm/v1/worker/workspace.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/workspace.py).
- **Use `memory_profiling`** from [`vllm/utils/mem_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/utils/mem_utils.py) to separate activation, KV-cache, and non-torch memory usage.
- **Adjust KV-cache** via `gpu_memory_utilization` or a fixed `kv_cache_memory_bytes` value.
- **Limit sequence length** with `max_model_len` if long prompts trigger OOM.
- **Shrink non-torch buffers** by setting `VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE` and tuning NCCL environment variables.

## Frequently Asked Questions

### What causes sudden OOM errors in vLLM after the engine starts running?

Sudden OOM errors typically occur when the workspace manager in [`vllm/v1/worker/workspace.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/workspace.py) resizes its buffer to accommodate a larger batch. The `resize_()` operation allocates new memory before releasing the old buffer, creating a temporary spike. Enable `VLLM_DEBUG_WORKSPACE` to confirm this, then pre-allocate a larger workspace or reduce `gpu_memory_utilization` to prevent the resize.

### How do I distinguish between KV-cache OOM and activation OOM?

Wrap your inference code with the `memory_profiling` context manager from [`vllm/utils/mem_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/utils/mem_utils.py). If `non_kv_cache_memory` remains stable but the process still OOMs, the KV-cache is too large. If `torch_peak_increase` spikes during generation, activation tensors are the bottleneck. Reduce `max_model_len` for cache issues or batch size for activation issues.

### Can I fix OOM errors without reducing batch size?

Yes. Instead of reducing batch size, set a fixed `kv_cache_memory_bytes` to cap cache growth, lower `max_model_len` to reduce per-token cache requirements, or decrease `VLLM_FLASHINFER_WORKSPACE_BUFFER_SIZE` to limit non-torch allocations. These adjustments preserve throughput while eliminating memory pressure.

### Why does vLLM OOM on multi-GPU setups when it works on a single GPU?

Multi-GPU deployments allocate NCCL communication buffers that do not appear in PyTorch’s memory statistics. These non-torch allocations, tracked in `MemorySnapshot.non_torch_memory`, can exhaust GPU memory even when `gpu_memory_utilization` is set low. Reduce `NCCL_BUFFSIZE` or further lower `gpu_memory_utilization` to accommodate these buffers.