How ComfyUI VRAM Management and Model Offloading Works: A Technical Deep Dive

ComfyUI uses a three-tier VRAM management subsystem—comprising VRAMState detection, ModelPatcher weight orchestration, and asynchronous offloading streams—to dynamically move model weights between GPU and CPU memory based on available VRAM and user flags.

ComfyUI’s VRAM management and model offloading system allows users to run large diffusion models on hardware with limited GPU memory. According to the Comfy-Org/ComfyUI source code, the framework automatically detects hardware capabilities, applies low-VRAM patches when necessary, and orchestrates weight movement between devices to prevent out-of-memory errors while maintaining inference speed.

The VRAMState Detection System

At startup, ComfyUI evaluates the hardware configuration and user preferences to establish a VRAMState. This enum, defined in comfy/model_management.py (lines 38‑45), categorizes the system into one of six tiers:

class VRAMState(Enum):
    DISABLED = 0    # No GPU available; pure CPU execution

    NO_VRAM = 1     # Force all offloading due to extremely limited VRAM

    LOW_VRAM = 2    # Aggressive offloading with patches

    NORMAL_VRAM = 3 # Standard GPU residency with selective offloading

    HIGH_VRAM = 4   # Keep models permanently on GPU

    SHARED = 5      # Unified memory (e.g., Apple Silicon)

The state is determined by parsing CLI arguments such as --lowvram, --novram, --highvram, or --normalvram (lines 18‑44). Once established, helper functions like unet_offload_device(), vae_offload_device(), and text_encoder_offload_device() return the appropriate target device for each model component:

def unet_offload_device():
    if vram_state == VRAMState.HIGH_VRAM:
        return get_torch_device()          # Retain on GPU

    else:
        return torch.device("cpu")         # Offload to system RAM

These selectors ensure that the UNet, VAE, and text encoders are routed to the correct device based on the current VRAM budget.

ModelPatcher: The Core Offloading Engine

The ModelPatcher class in comfy/model_patcher.py wraps PyTorch models and implements the actual weight movement logic. It serves as the primary interface for loading, unloading, and pinning model parameters.

Low-VRAM Patching Strategy

When ModelPatcher.load() is invoked (lines 267‑447), it iterates through every submodule and calculates memory requirements. If the accumulated memory exceeds the lowvram_model_memory threshold, the system enters low-VRAM mode:

  • Full-load path: If sufficient free VRAM exists (mem_counter + module_mem + potential_offload < lowvram_model_memory), weights are copied directly to the compute device.
  • Low-VRAM path: Weights remain on the CPU, and LowVramPatch objects are injected (lines 66‑78). These patches lazily materialize weights during the forward pass, applying LoRA or style-transfer modifications on-the-fly without permanently occupying GPU memory.

Weight Pinning for Performance

To avoid repeated copy operations during iterative sampling, ModelPatcher supports weight pinning via pin_weight_to_device() (lines 80‑90). When a weight is pinned, it remains resident in GPU memory even if the module is technically "offloaded," allowing rapid access during subsequent inference steps. The method unpin_all_weights() automatically releases these reservations during prompt cleanup.

Asynchronous Offloading and Memory Streams

ComfyUI minimizes transfer latency by overlapping data movement with computation. When the args.async_offload flag is enabled (default for NVIDIA and AMD GPUs), the system initializes a pool of CUDA or XPU streams (NUM_STREAMS, lines 71‑85 in comfy/model_management.py).

In comfy/ops.py, the cast_bias_weight() function (lines 197‑226) checks if a weight resides on the offload device. If so, it uses get_offload_stream() to perform asynchronous copies via cast_to_gathered(), allowing the next computation kernel to execute while the previous weight transfers. This pipeline hides memory transfer costs behind active GPU operations.

Global Loading Workflow

Before executing a prompt, the execution engine (in execution.py) calls load_models_gpu() (lines 590‑652 in comfy/model_management.py) to ensure all required models fit within the allocated memory budget:

  1. Memory Calculation: Aggregates memory_required across all pending ModelPatcher instances and adds a safety margin (extra_mem).
  2. Memory Reclamation: Invokes free_memory() to unload existing models that exceed the current budget.
  3. Model Loading: Iterates through the model list, calling ModelPatcher.load() with the computed constraints.
  4. Fallback Handling: If models cannot fit even after memory freeing, the system either forces high-VRAM mode or raises an OOM_EXCEPTION.

This orchestration guarantees that the active working set resides on the GPU while dormant weights remain on the CPU or pinned in shared memory.

Practical Implementation Examples

Enabling Low-VRAM Mode via Command Line

Force the system into aggressive offloading to run large models on 4‑8 GB GPUs:

python main.py --lowvram

This sets VRAMState.LOW_VRAM, triggering the low-VRAM patch insertion logic in ModelPatcher.load().

Manual Model Loading with Custom Memory Budgets

Programmatically control how much VRAM a specific model may consume:

import comfy.model_management as mm
from comfy.model_patcher import ModelPatcher
from comfy.sd import load_checkpoint

# Load checkpoint without immediate device placement

raw_model = load_checkpoint("sd15.ckpt")

# Wrap with device directives

patcher = ModelPatcher(
    raw_model,
    load_device=mm.get_torch_device(),      # Target GPU

    offload_device=mm.unet_offload_device() # Fallback CPU

)

# Allocate 2 GiB; system decides between full-load or low-VRAM patches

mm.load_models_gpu(
    [patcher], 
    memory_required=2 * 1024**3, 
    force_full_load=False
)

Pinning Critical Weights to GPU

Prevent frequent offloading for layers accessed repeatedly during sampling:


# Inside a custom node execution

model = inputs["model"]  # ModelPatcher instance

weight_key = "mid_block.0.attn1.to_q.weight"

# Pin to GPU for the duration of the prompt

model.pin_weight_to_device(weight_key)

# Execute inference...

# Cleanup occurs automatically via unpin_all_weights()

Configuring Asynchronous Offload Streams

Increase the number of concurrent transfer streams to reduce latency on high-end GPUs:

import comfy.model_management as mm

# Default is 2 for CUDA/AMD; increase for very large models

mm.NUM_STREAMS = 4

Alternatively, set via CLI: --async_offload 4.

Summary

  • VRAMState categorizes hardware capabilities and drives device selection for UNet, VAE, and text encoder components via helpers in comfy/model_management.py.
  • ModelPatcher orchestrates weight placement, implementing low-VRAM patches for memory-constrained environments and pinning for performance-critical weights.
  • Asynchronous streams hide transfer latency by overlapping CPU-to-GPU copies with active computation in comfy/ops.py.
  • The global loader (load_models_gpu()) coordinates memory budgets across the entire prompt execution, automatically offloading or freeing models to prevent OOM errors.

Frequently Asked Questions

What is the difference between --lowvram and --novram flags?

--lowvram enables the low-VRAM patching system in ModelPatcher, which keeps weights on CPU and patches them into GPU memory only during forward passes. --novram forces VRAMState.NO_VRAM, which aggressively offloads everything to system RAM and may disable certain GPU optimizations entirely. Use --lowvram for 4‑6 GB cards and --novram only when the GPU has minimal memory (under 4 GB) or for debugging CPU-only execution.

How does ComfyUI decide when to offload a model to CPU?

The decision occurs in load_models_gpu() within comfy/model_management.py. The function calculates the total memory required for all models in the current prompt, compares it against get_free_memory(), and calls free_memory() to unload non-essential models if the budget is exceeded. Individual weights are offloaded based on the unet_offload_device() and similar helpers, which return CPU for all states except VRAMState.HIGH_VRAM.

What is weight pinning and when should I use it?

Weight pinning via ModelPatcher.pin_weight_to_device() keeps specific weight tensors resident in GPU memory even when their parent module is technically offloaded. This is useful for weights accessed repeatedly during iterative sampling (such as attention layers), as it eliminates the overhead of repeated CPU-to-GPU copies. Pin only the weights you access frequently, as excessive pinning reduces available VRAM for other operations.

Can ComfyUI run entirely without a GPU?

Yes. When no CUDA, XPU, or MPS device is detected, ComfyUI defaults to VRAMState.DISABLED, routing all computation to the CPU via torch.device("cpu"). The ModelPatcher logic remains functional, though inference will be significantly slower. Use the --cpu flag to force this mode explicitly, ensuring that all offload_device() helpers return CPU and no GPU memory management is attempted.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →