# How ComfyUI VRAM Management and Model Offloading Works: A Technical Deep Dive

> Discover how ComfyUI manages VRAM and offloads models using a three-tier system dynamic VRAM usage and GPU CPU weight transfer for efficient AI image generation.

- Repository: [Comfy Org/ComfyUI](https://github.com/Comfy-Org/ComfyUI)
- Tags: deep-dive
- Published: 2026-02-26

---

**ComfyUI uses a three-tier VRAM management subsystem—comprising VRAMState detection, ModelPatcher weight orchestration, and asynchronous offloading streams—to dynamically move model weights between GPU and CPU memory based on available VRAM and user flags.**

ComfyUI’s VRAM management and model offloading system allows users to run large diffusion models on hardware with limited GPU memory. According to the Comfy-Org/ComfyUI source code, the framework automatically detects hardware capabilities, applies low-VRAM patches when necessary, and orchestrates weight movement between devices to prevent out-of-memory errors while maintaining inference speed.

## The VRAMState Detection System

At startup, ComfyUI evaluates the hardware configuration and user preferences to establish a **VRAMState**. This enum, defined in [`comfy/model_management.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_management.py) (lines 38‑45), categorizes the system into one of six tiers:

```python
class VRAMState(Enum):
    DISABLED = 0    # No GPU available; pure CPU execution

    NO_VRAM = 1     # Force all offloading due to extremely limited VRAM

    LOW_VRAM = 2    # Aggressive offloading with patches

    NORMAL_VRAM = 3 # Standard GPU residency with selective offloading

    HIGH_VRAM = 4   # Keep models permanently on GPU

    SHARED = 5      # Unified memory (e.g., Apple Silicon)

```

The state is determined by parsing CLI arguments such as `--lowvram`, `--novram`, `--highvram`, or `--normalvram` (lines 18‑44). Once established, helper functions like `unet_offload_device()`, `vae_offload_device()`, and `text_encoder_offload_device()` return the appropriate target device for each model component:

```python
def unet_offload_device():
    if vram_state == VRAMState.HIGH_VRAM:
        return get_torch_device()          # Retain on GPU

    else:
        return torch.device("cpu")         # Offload to system RAM

```

These selectors ensure that the UNet, VAE, and text encoders are routed to the correct device based on the current VRAM budget.

## ModelPatcher: The Core Offloading Engine

The `ModelPatcher` class in [`comfy/model_patcher.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_patcher.py) wraps PyTorch models and implements the actual weight movement logic. It serves as the primary interface for loading, unloading, and pinning model parameters.

### Low-VRAM Patching Strategy

When `ModelPatcher.load()` is invoked (lines 267‑447), it iterates through every submodule and calculates memory requirements. If the accumulated memory exceeds the `lowvram_model_memory` threshold, the system enters **low-VRAM mode**:

- **Full-load path**: If sufficient free VRAM exists (`mem_counter + module_mem + potential_offload < lowvram_model_memory`), weights are copied directly to the compute device.
- **Low-VRAM path**: Weights remain on the CPU, and `LowVramPatch` objects are injected (lines 66‑78). These patches lazily materialize weights during the forward pass, applying LoRA or style-transfer modifications on-the-fly without permanently occupying GPU memory.

### Weight Pinning for Performance

To avoid repeated copy operations during iterative sampling, `ModelPatcher` supports **weight pinning** via `pin_weight_to_device()` (lines 80‑90). When a weight is pinned, it remains resident in GPU memory even if the module is technically "offloaded," allowing rapid access during subsequent inference steps. The method `unpin_all_weights()` automatically releases these reservations during prompt cleanup.

## Asynchronous Offloading and Memory Streams

ComfyUI minimizes transfer latency by overlapping data movement with computation. When the `args.async_offload` flag is enabled (default for NVIDIA and AMD GPUs), the system initializes a pool of CUDA or XPU streams (`NUM_STREAMS`, lines 71‑85 in [`comfy/model_management.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_management.py)).

In [`comfy/ops.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/ops.py), the `cast_bias_weight()` function (lines 197‑226) checks if a weight resides on the offload device. If so, it uses `get_offload_stream()` to perform asynchronous copies via `cast_to_gathered()`, allowing the next computation kernel to execute while the previous weight transfers. This pipeline hides memory transfer costs behind active GPU operations.

## Global Loading Workflow

Before executing a prompt, the execution engine (in [`execution.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/execution.py)) calls `load_models_gpu()` (lines 590‑652 in [`comfy/model_management.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_management.py)) to ensure all required models fit within the allocated memory budget:

1. **Memory Calculation**: Aggregates `memory_required` across all pending `ModelPatcher` instances and adds a safety margin (`extra_mem`).
2. **Memory Reclamation**: Invokes `free_memory()` to unload existing models that exceed the current budget.
3. **Model Loading**: Iterates through the model list, calling `ModelPatcher.load()` with the computed constraints.
4. **Fallback Handling**: If models cannot fit even after memory freeing, the system either forces high-VRAM mode or raises an `OOM_EXCEPTION`.

This orchestration guarantees that the active working set resides on the GPU while dormant weights remain on the CPU or pinned in shared memory.

## Practical Implementation Examples

### Enabling Low-VRAM Mode via Command Line

Force the system into aggressive offloading to run large models on 4‑8 GB GPUs:

```bash
python main.py --lowvram

```

This sets `VRAMState.LOW_VRAM`, triggering the low-VRAM patch insertion logic in `ModelPatcher.load()`.

### Manual Model Loading with Custom Memory Budgets

Programmatically control how much VRAM a specific model may consume:

```python
import comfy.model_management as mm
from comfy.model_patcher import ModelPatcher
from comfy.sd import load_checkpoint

# Load checkpoint without immediate device placement

raw_model = load_checkpoint("sd15.ckpt")

# Wrap with device directives

patcher = ModelPatcher(
    raw_model,
    load_device=mm.get_torch_device(),      # Target GPU

    offload_device=mm.unet_offload_device() # Fallback CPU

)

# Allocate 2 GiB; system decides between full-load or low-VRAM patches

mm.load_models_gpu(
    [patcher], 
    memory_required=2 * 1024**3, 
    force_full_load=False
)

```

### Pinning Critical Weights to GPU

Prevent frequent offloading for layers accessed repeatedly during sampling:

```python

# Inside a custom node execution

model = inputs["model"]  # ModelPatcher instance

weight_key = "mid_block.0.attn1.to_q.weight"

# Pin to GPU for the duration of the prompt

model.pin_weight_to_device(weight_key)

# Execute inference...

# Cleanup occurs automatically via unpin_all_weights()

```

### Configuring Asynchronous Offload Streams

Increase the number of concurrent transfer streams to reduce latency on high-end GPUs:

```python
import comfy.model_management as mm

# Default is 2 for CUDA/AMD; increase for very large models

mm.NUM_STREAMS = 4

```

Alternatively, set via CLI: `--async_offload 4`.

## Summary

- **VRAMState** categorizes hardware capabilities and drives device selection for UNet, VAE, and text encoder components via helpers in [`comfy/model_management.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_management.py).
- **ModelPatcher** orchestrates weight placement, implementing low-VRAM patches for memory-constrained environments and pinning for performance-critical weights.
- **Asynchronous streams** hide transfer latency by overlapping CPU-to-GPU copies with active computation in [`comfy/ops.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/ops.py).
- The **global loader** (`load_models_gpu()`) coordinates memory budgets across the entire prompt execution, automatically offloading or freeing models to prevent OOM errors.

## Frequently Asked Questions

### What is the difference between `--lowvram` and `--novram` flags?

**`--lowvram`** enables the low-VRAM patching system in `ModelPatcher`, which keeps weights on CPU and patches them into GPU memory only during forward passes. **`--novram`** forces `VRAMState.NO_VRAM`, which aggressively offloads everything to system RAM and may disable certain GPU optimizations entirely. Use `--lowvram` for 4‑6 GB cards and `--novram` only when the GPU has minimal memory (under 4 GB) or for debugging CPU-only execution.

### How does ComfyUI decide when to offload a model to CPU?

The decision occurs in `load_models_gpu()` within [`comfy/model_management.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/model_management.py). The function calculates the total memory required for all models in the current prompt, compares it against `get_free_memory()`, and calls `free_memory()` to unload non-essential models if the budget is exceeded. Individual weights are offloaded based on the `unet_offload_device()` and similar helpers, which return CPU for all states except `VRAMState.HIGH_VRAM`.

### What is weight pinning and when should I use it?

**Weight pinning** via `ModelPatcher.pin_weight_to_device()` keeps specific weight tensors resident in GPU memory even when their parent module is technically offloaded. This is useful for weights accessed repeatedly during iterative sampling (such as attention layers), as it eliminates the overhead of repeated CPU-to-GPU copies. Pin only the weights you access frequently, as excessive pinning reduces available VRAM for other operations.

### Can ComfyUI run entirely without a GPU?

Yes. When no CUDA, XPU, or MPS device is detected, ComfyUI defaults to `VRAMState.DISABLED`, routing all computation to the CPU via `torch.device("cpu")`. The `ModelPatcher` logic remains functional, though inference will be significantly slower. Use the `--cpu` flag to force this mode explicitly, ensuring that all `offload_device()` helpers return CPU and no GPU memory management is attempted.