# How to Optimize GPU Memory Utilization with gpu_memory_utilization in vLLM

> Optimize GPU memory utilization in vLLM by setting the gpu_memory_utilization parameter. Balance throughput and OOM risk by adjusting this value. Learn how now.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**Set the `gpu_memory_utilization` parameter between 0 and 1 to control the fraction of GPU memory reserved for model weights, activations, and the KV cache, where higher values improve throughput but increase OOM risk.**

The `gpu_memory_utilization` flag in the vLLM inference engine is the primary mechanism for managing GPU memory allocation. By tuning this ratio in the vLLM-project/vllm repository, you can balance between fitting larger models, maximizing throughput, and preventing out-of-memory errors during high-load scenarios.

## Understanding gpu_memory_utilization in vLLM

### Configuration and Default Behavior

In [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py), the `gpu_memory_utilization` field is defined within `CacheConfig` with a default value of `0.9`, constraining valid inputs to the range `(0, 1]`:

```python

# vllm/config/cache.py

gpu_memory_utilization: float = Field(default=0.9, gt=0, le=1)

```

This parameter determines what percentage of the GPU's total memory the engine is permitted to use for the model weights, activation tensors, and the KV cache. When you initialize an `LLM` instance, this value propagates through the configuration pipeline to the worker nodes.

### Memory Budgeting Mechanics

Under the hood, vLLM performs a two-phase memory calculation:

1. **Startup Validation**: The engine calculates `requested_memory` by multiplying the device's total memory by the utilization ratio. In [`vllm/v1/worker/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/utils.py), the `request_memory` function validates this against available free memory at initialization:

   ```python
   # vllm/v1/worker/utils.py

   requested_memory = math.ceil(
       init_snapshot.total_memory * cache_config.gpu_memory_utilization
   )
   ```

   If the requested memory exceeds available free memory, vLLM raises a `ValueError` immediately rather than failing later during inference.

2. **KV-Cache Allocation**: After profiling the model's non-KV memory requirements, the worker allocates the remaining budget to the KV cache. In [`vllm/v1/worker/gpu_worker.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py), this calculation appears as:

   ```python
   # vllm/v1/worker/gpu_worker.py

   self.available_kv_cache_memory_bytes = (
       self.requested_memory - profile_result.non_kv_cache_memory
   )
   ```

   This ensures the KV cache dynamically fills the space permitted by `gpu_memory_utilization` without exceeding the GPU's physical limits.

## Practical Use Cases and Tuning Strategies

### Fitting Larger Models on Fixed GPUs

To load a model that barely fits your hardware, **decrease** `gpu_memory_utilization` (e.g., from `0.9` to `0.7` or `0.5`). This reduces the memory reserved for the KV cache, potentially limiting maximum context length or batch size, but frees sufficient memory for the model weights to load successfully.

### Maximizing Throughput

For production deployments where latency is less critical than tokens-per-second, **increase** the value toward the maximum of `0.95` or keep the default `0.9`. Higher utilization allocates more memory to the KV cache, enabling longer sequences and larger batch sizes. This reduces the frequency of cache evictions and improves overall throughput.

### Avoiding Out-of-Memory Errors

When encountering CUDA OOM errors during peak load, incrementally lower `gpu_memory_utilization` until the startup validation passes. Unlike reactive memory management, this approach guarantees allocation safety at initialization time, though it may reduce the maximum batch size your deployment can handle.

## Implementation Details from Source Code

The parameter surfaces in both the Python API and command-line interface. In [`vllm/entrypoints/llm.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/llm.py), the `LLM` constructor accepts `gpu_memory_utilization` directly:

```python

# vllm/entrypoints/llm.py

def __init__(..., gpu_memory_utilization: float = 0.9, ...):

```

For CLI deployments, [`vllm/engine/arg_utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/engine/arg_utils.py) exposes the flag as `--gpu-memory-utilization`, parsing it into `EngineArgs` before passing it to the underlying engine.

## Interaction with Other Memory Parameters

Understanding how `gpu_memory_utilization` interacts with other flags is critical for advanced tuning:

- **`kv_cache_memory_bytes`**: When set explicitly, this **overrides** the utilization-based calculation. The engine allocates exactly the specified bytes for the KV cache and ignores the `gpu_memory_utilization` value for cache sizing, though the parameter still constrains total memory.

- **`cpu_offload_gb`**: Offloading model weights to CPU effectively increases the GPU memory budget available for the KV cache. You can maintain a high `gpu_memory_utilization` (e.g., `0.9`) while using `cpu_offload_gb` to handle large models that would otherwise OOM.

- **`swap_space`**: This allocates CPU memory for block swapping when the KV cache is full. While it does not directly affect GPU utilization, increasing swap space allows you to lower `gpu_memory_utilization` without sacrificing the ability to handle long contexts.

## Code Examples

### Python API Configuration

Configure `gpu_memory_utilization` when instantiating the `LLM` class to reserve 70% of GPU memory:

```python
from vllm import LLM, SamplingParams

llm = LLM(
    model="meta-llama/Meta-Llama-3-8B",
    gpu_memory_utilization=0.7,
)

sampling_params = SamplingParams(temperature=0.7, top_p=0.9)
output = llm.generate(["Explain quantum computing"], sampling_params)
print(output[0].outputs[0].text)

```

### Command-Line Deployment

Launch the OpenAI-compatible API server with reduced memory utilization to accommodate shared GPU environments:

```bash
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3-8B \
    --gpu-memory-utilization 0.6 \
    --port 8000

```

### Combining with CPU Offloading

For extremely large models, combine high utilization with weight offloading:

```python
llm = LLM(
    model="bigscience/T0pp",
    gpu_memory_utilization=0.85,
    cpu_offload_gb=8,
    swap_space=2,
)

```

This configuration keeps 85% of GPU memory for active inference while offloading 8 GB of weights to system RAM.

### Runtime Memory Inspection

Verify the actual memory allocation after engine initialization:

```python
engine = llm.llm_engine
print(f"GPU memory utilization: {engine.cache_config.gpu_memory_utilization}")
print(f"KV cache bytes: {engine.cache_config.kv_cache_memory_bytes}")

```

## Summary

- **`gpu_memory_utilization`** controls the fraction of total GPU memory vLLM can use, defaulting to `0.9` in [`vllm/config/cache.py`](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py).
- Lower values (e.g., `0.5`–`0.7`) help fit larger models or prevent OOMs on shared GPUs, while higher values maximize KV-cache capacity for throughput.
- The engine validates memory requests at startup in [`vllm/v1/worker/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/utils.py), raising errors immediately if the allocation exceeds available free memory.
- Set `kv_cache_memory_bytes` directly to override utilization-based calculations for deterministic multi-tenant deployments.
- Combine with `cpu_offload_gb` to load models exceeding GPU memory while maintaining high utilization ratios for the KV cache.

## Frequently Asked Questions

### What happens if I set gpu_memory_utilization too high?

If you set the value too high for your current GPU state, vLLM raises a `ValueError` during initialization in [`vllm/v1/worker/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/utils.py) when the `request_memory` calculation exceeds available free memory. Even if initialization succeeds, setting it to `0.95` or higher on a shared GPU risks runtime OOM errors if other processes allocate memory after vLLM starts.

### How does gpu_memory_utilization differ from kv_cache_memory_bytes?

`gpu_memory_utilization` is a ratio (0.0–1.0) that constrains total GPU memory usage for weights, activations, and KV cache combined. In contrast, `kv_cache_memory_bytes` is an absolute byte value that explicitly sets the KV cache size, causing the engine to ignore `gpu_memory_utilization` for cache calculations. Use the latter when you need deterministic memory allocation across different hardware configurations.

### Can I change gpu_memory_utilization after the engine has started?

No, this parameter is fixed at initialization. The memory budget is calculated and locked during the worker startup phase in [`vllm/v1/worker/gpu_worker.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py). To change the utilization ratio, you must restart the vLLM engine with the new value.

### Why does lowering gpu_memory_utilization sometimes reduce maximum context length?

When you lower the ratio, the engine has less total memory to allocate after accounting for model weights and activations. Since the KV cache size is derived from the remaining budget in `available_kv_cache_memory_bytes`, a smaller cache stores fewer tokens, effectively reducing the maximum context length or batch size your deployment can handle simultaneously.