How ComfyUI VRAM Management and Model Offloading Works: A Technical Deep Dive
ComfyUI uses a three-tier VRAM management subsystem—comprising VRAMState detection, ModelPatcher weight orchestration, and asynchronous offloading streams—to dynamically move model weights between GPU and CPU memory based on available VRAM and user flags.
ComfyUI’s VRAM management and model offloading system allows users to run large diffusion models on hardware with limited GPU memory. According to the Comfy-Org/ComfyUI source code, the framework automatically detects hardware capabilities, applies low-VRAM patches when necessary, and orchestrates weight movement between devices to prevent out-of-memory errors while maintaining inference speed.
The VRAMState Detection System
At startup, ComfyUI evaluates the hardware configuration and user preferences to establish a VRAMState. This enum, defined in comfy/model_management.py (lines 38‑45), categorizes the system into one of six tiers:
class VRAMState(Enum):
DISABLED = 0 # No GPU available; pure CPU execution
NO_VRAM = 1 # Force all offloading due to extremely limited VRAM
LOW_VRAM = 2 # Aggressive offloading with patches
NORMAL_VRAM = 3 # Standard GPU residency with selective offloading
HIGH_VRAM = 4 # Keep models permanently on GPU
SHARED = 5 # Unified memory (e.g., Apple Silicon)
The state is determined by parsing CLI arguments such as --lowvram, --novram, --highvram, or --normalvram (lines 18‑44). Once established, helper functions like unet_offload_device(), vae_offload_device(), and text_encoder_offload_device() return the appropriate target device for each model component:
def unet_offload_device():
if vram_state == VRAMState.HIGH_VRAM:
return get_torch_device() # Retain on GPU
else:
return torch.device("cpu") # Offload to system RAM
These selectors ensure that the UNet, VAE, and text encoders are routed to the correct device based on the current VRAM budget.
ModelPatcher: The Core Offloading Engine
The ModelPatcher class in comfy/model_patcher.py wraps PyTorch models and implements the actual weight movement logic. It serves as the primary interface for loading, unloading, and pinning model parameters.
Low-VRAM Patching Strategy
When ModelPatcher.load() is invoked (lines 267‑447), it iterates through every submodule and calculates memory requirements. If the accumulated memory exceeds the lowvram_model_memory threshold, the system enters low-VRAM mode:
- Full-load path: If sufficient free VRAM exists (
mem_counter + module_mem + potential_offload < lowvram_model_memory), weights are copied directly to the compute device. - Low-VRAM path: Weights remain on the CPU, and
LowVramPatchobjects are injected (lines 66‑78). These patches lazily materialize weights during the forward pass, applying LoRA or style-transfer modifications on-the-fly without permanently occupying GPU memory.
Weight Pinning for Performance
To avoid repeated copy operations during iterative sampling, ModelPatcher supports weight pinning via pin_weight_to_device() (lines 80‑90). When a weight is pinned, it remains resident in GPU memory even if the module is technically "offloaded," allowing rapid access during subsequent inference steps. The method unpin_all_weights() automatically releases these reservations during prompt cleanup.
Asynchronous Offloading and Memory Streams
ComfyUI minimizes transfer latency by overlapping data movement with computation. When the args.async_offload flag is enabled (default for NVIDIA and AMD GPUs), the system initializes a pool of CUDA or XPU streams (NUM_STREAMS, lines 71‑85 in comfy/model_management.py).
In comfy/ops.py, the cast_bias_weight() function (lines 197‑226) checks if a weight resides on the offload device. If so, it uses get_offload_stream() to perform asynchronous copies via cast_to_gathered(), allowing the next computation kernel to execute while the previous weight transfers. This pipeline hides memory transfer costs behind active GPU operations.
Global Loading Workflow
Before executing a prompt, the execution engine (in execution.py) calls load_models_gpu() (lines 590‑652 in comfy/model_management.py) to ensure all required models fit within the allocated memory budget:
- Memory Calculation: Aggregates
memory_requiredacross all pendingModelPatcherinstances and adds a safety margin (extra_mem). - Memory Reclamation: Invokes
free_memory()to unload existing models that exceed the current budget. - Model Loading: Iterates through the model list, calling
ModelPatcher.load()with the computed constraints. - Fallback Handling: If models cannot fit even after memory freeing, the system either forces high-VRAM mode or raises an
OOM_EXCEPTION.
This orchestration guarantees that the active working set resides on the GPU while dormant weights remain on the CPU or pinned in shared memory.
Practical Implementation Examples
Enabling Low-VRAM Mode via Command Line
Force the system into aggressive offloading to run large models on 4‑8 GB GPUs:
python main.py --lowvram
This sets VRAMState.LOW_VRAM, triggering the low-VRAM patch insertion logic in ModelPatcher.load().
Manual Model Loading with Custom Memory Budgets
Programmatically control how much VRAM a specific model may consume:
import comfy.model_management as mm
from comfy.model_patcher import ModelPatcher
from comfy.sd import load_checkpoint
# Load checkpoint without immediate device placement
raw_model = load_checkpoint("sd15.ckpt")
# Wrap with device directives
patcher = ModelPatcher(
raw_model,
load_device=mm.get_torch_device(), # Target GPU
offload_device=mm.unet_offload_device() # Fallback CPU
)
# Allocate 2 GiB; system decides between full-load or low-VRAM patches
mm.load_models_gpu(
[patcher],
memory_required=2 * 1024**3,
force_full_load=False
)
Pinning Critical Weights to GPU
Prevent frequent offloading for layers accessed repeatedly during sampling:
# Inside a custom node execution
model = inputs["model"] # ModelPatcher instance
weight_key = "mid_block.0.attn1.to_q.weight"
# Pin to GPU for the duration of the prompt
model.pin_weight_to_device(weight_key)
# Execute inference...
# Cleanup occurs automatically via unpin_all_weights()
Configuring Asynchronous Offload Streams
Increase the number of concurrent transfer streams to reduce latency on high-end GPUs:
import comfy.model_management as mm
# Default is 2 for CUDA/AMD; increase for very large models
mm.NUM_STREAMS = 4
Alternatively, set via CLI: --async_offload 4.
Summary
- VRAMState categorizes hardware capabilities and drives device selection for UNet, VAE, and text encoder components via helpers in
comfy/model_management.py. - ModelPatcher orchestrates weight placement, implementing low-VRAM patches for memory-constrained environments and pinning for performance-critical weights.
- Asynchronous streams hide transfer latency by overlapping CPU-to-GPU copies with active computation in
comfy/ops.py. - The global loader (
load_models_gpu()) coordinates memory budgets across the entire prompt execution, automatically offloading or freeing models to prevent OOM errors.
Frequently Asked Questions
What is the difference between --lowvram and --novram flags?
--lowvram enables the low-VRAM patching system in ModelPatcher, which keeps weights on CPU and patches them into GPU memory only during forward passes. --novram forces VRAMState.NO_VRAM, which aggressively offloads everything to system RAM and may disable certain GPU optimizations entirely. Use --lowvram for 4‑6 GB cards and --novram only when the GPU has minimal memory (under 4 GB) or for debugging CPU-only execution.
How does ComfyUI decide when to offload a model to CPU?
The decision occurs in load_models_gpu() within comfy/model_management.py. The function calculates the total memory required for all models in the current prompt, compares it against get_free_memory(), and calls free_memory() to unload non-essential models if the budget is exceeded. Individual weights are offloaded based on the unet_offload_device() and similar helpers, which return CPU for all states except VRAMState.HIGH_VRAM.
What is weight pinning and when should I use it?
Weight pinning via ModelPatcher.pin_weight_to_device() keeps specific weight tensors resident in GPU memory even when their parent module is technically offloaded. This is useful for weights accessed repeatedly during iterative sampling (such as attention layers), as it eliminates the overhead of repeated CPU-to-GPU copies. Pin only the weights you access frequently, as excessive pinning reduces available VRAM for other operations.
Can ComfyUI run entirely without a GPU?
Yes. When no CUDA, XPU, or MPS device is detected, ComfyUI defaults to VRAMState.DISABLED, routing all computation to the CPU via torch.device("cpu"). The ModelPatcher logic remains functional, though inference will be significantly slower. Use the --cpu flag to force this mode explicitly, ensuring that all offload_device() helpers return CPU and no GPU memory management is attempted.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →