How the ComfyUI Execution Caching System Works: LRU and RAM Pressure Explained

ComfyUI speeds up workflow execution by caching intermediate node outputs using three strategies: Classic (unlimited, prompt-scoped), LRU (fixed-size with generational eviction), and RAM-Pressure (dynamic eviction based on available system memory).

The execution caching system in ComfyUI eliminates redundant computation by storing node results between prompts. Depending on your hardware constraints and workflow complexity, you can choose between a simple unlimited cache, a size-restricted LRU (Least Recently Used) cache, or a RAM-Pressure cache that actively monitors system memory. This article breaks down the implementation details found in the Comfy-Org/ComfyUI repository, including the exact eviction algorithms and command-line flags that control them.

Where the Cache Is Initialized

The PromptExecutor class in execution.py instantiates the caching layer at the start of every prompt. Inside PromptExecutor.__init__, the cache_type and cache_args parameters determine which strategy to use:


# execution.py – PromptExecutor.__init__

class PromptExecutor:
    def __init__(self, server, cache_type=False, cache_args=None):
        self.cache_args = cache_args
        self.cache_type = cache_type
        self.server = server
        self.reset()

The reset() method constructs a fresh CacheSet object, which wraps two distinct stores: outputs (for tensors) and objects (for model handles). The CacheSet.__init__ method routes to the appropriate initialization logic based on the cache_type enum:


# execution.py – CacheSet.__init__

class CacheSet:
    def __init__(self, cache_type=None, cache_args={}):
        if cache_type == CacheType.NONE:
            self.init_null_cache()
        elif cache_type == CacheType.RAM_PRESSURE:
            self.init_ram_cache(cache_args.get("ram", 16.0))
        elif cache_type == CacheType.LRU:
            self.init_lru_cache(cache_args.get("lru", 0))
        else:
            self.init_classic_cache()
        self.all = [self.outputs, self.objects]

The Three Caching Strategies

ComfyUI implements three distinct caching behaviors, each optimized for different memory and performance trade-offs.

Classic Cache (Default Behavior)

When no --cache-* flags are provided, ComfyUI uses the Classic cache. This strategy stores results in a HierarchicalCache keyed by a node's input signature—a hash derived from the node's class type, input values, and "is_changed" flags. The cache persists for the duration of the prompt and is automatically cleared when execution completes, meaning it never evicts entries mid-prompt regardless of memory usage.


# execution.py – CacheSet.init_classic_cache

def init_classic_cache(self):
    self.outputs = HierarchicalCache(CacheKeySetInputSignature)
    self.objects = HierarchicalCache(CacheKeySetID)

LRU Cache (--cache-lru)

The LRU strategy wraps the classic cache with a hard limit on the number of stored entries. Enable it with --cache-lru N, where N represents the maximum number of node results to retain. The implementation in comfy_execution/caching.py uses a generation counter that increments at the start of each prompt to track recency:


# comfy_execution/caching.py – LRUCache.__init__

class LRUCache(BasicCache):
    def __init__(self, key_class, max_size=100):
        super().__init__(key_class)
        self.max_size = max_size
        self.min_generation = 0
        self.generation = 0
        self.used_generation = {}

When a node is accessed, _mark_used records the current generation. If the cache exceeds max_size, clean_unused removes entries from the oldest generation until the count falls below the limit:


# LRUCache.clean_unused

while len(self.cache) > self.max_size and self.min_generation < self.generation:
    self.min_generation += 1
    to_remove = [key for key in self.cache if self.used_generation[key] < self.min_generation]
    for key in to_remove:
        del self.cache[key]
        del self.used_generation[key]

RAM-Pressure Cache (--cache-ram)

The RAM-Pressure cache dynamically evicts entries based on real-time system memory availability. Activate it with --cache-ram M, where M specifies the minimum free RAM headroom in gigabytes. This mode extends LRUCache but disables size-based eviction (max_size=0) in favor of memory-driven cleanup.


# comfy_execution/caching.py – RAMPressureCache.__init__

class RAMPressureCache(LRUCache):
    def __init__(self, key_class):
        super().__init__(key_class, 0)  # size=0 disables LRU limit

        self.timestamps = {}

How RAM-Pressure Eviction Calculates OOM Scores

The RAM-Pressure strategy determines which entries to delete using an OOM-score (Out-Of-Memory score) that balances entry age against memory footprint. After each node executes, PromptExecutor calls poll() on the cache:


# execution.py – main execution loop

self.caches.outputs.poll(ram_headroom=self.cache_args["ram"])

The RAMPressureCache.poll method checks available RAM via psutil.virtual_memory().available. When free memory drops below the configured headroom, it constructs a prioritized clean-list:


# comfy_execution/caching.py – RAMPressureCache.poll

for key, (outputs, _) in self.cache.items():
    oom_score = RAM_CACHE_OLD_WORKFLOW_OOM_MULTIPLIER ** (self.generation - self.used_generation[key])
    ram_usage = RAM_CACHE_DEFAULT_RAM_USAGE
    # Scan outputs for CPU tensors and sum their memory (discounted by 0.5)

    oom_score *= ram_usage
    bisect.insort(clean_list, (oom_score, self.timestamps[key], key))

The OOM-score formula combines:

  • Age factor: Exponential weighting based on how many prompt generations have passed since the entry was last used
  • Memory footprint: Sum of all CPU tensor sizes (GPU tensors are excluded because CUDA manages that memory separately)

Entries with the lowest scores (old and large) are deleted first until free RAM exceeds the headroom multiplied by RAM_CACHE_HYSTERESIS (approximately 1.1), providing a safety margin against thrashing.

Enabling Caching Modes from the Command Line

ComfyUI exposes these caching strategies through command-line arguments parsed in comfy/cli_args.py. The flags propagate to PromptExecutor via the cache_type and cache_args parameters.

Flag Description Example
--cache-lru N Enable LRU cache with maximum N entries python main.py --cache-lru 200
--cache-ram M Enable RAM-Pressure cache with M GB headroom python main.py --cache-ram 8
--cache-none Disable all caching (useful for debugging) python main.py --cache-none

You can also instantiate a PromptExecutor programmatically with specific cache settings:

from comfy.execution import PromptExecutor, CacheType

executor = PromptExecutor(
    server=my_server,
    cache_type=CacheType.RAM_PRESSURE,
    cache_args={"ram": 4.0}  # Maintain 4 GB free RAM

)

Summary

  • ComfyUI's execution caching system offers three modes: Classic (unlimited, prompt-scoped), LRU (fixed entry count with generational eviction), and RAM-Pressure (dynamic eviction based on available RAM).
  • The CacheSet class in execution.py initializes the appropriate cache implementation during PromptExecutor.reset().
  • LRU eviction removes the least-recently-used entries by comparing generation counters when the cache exceeds its maximum size.
  • RAM-Pressure eviction calculates an OOM-score combining entry age and CPU tensor memory usage, deleting low-scoring entries when system memory falls below the configured headroom.
  • Control caching behavior via --cache-lru, --cache-ram, or --cache-none command-line flags.

Frequently Asked Questions

What is the default caching behavior in ComfyUI?

By default, ComfyUI uses the Classic cache, which stores unlimited node outputs for the duration of a single prompt. The cache is cleared automatically when the prompt finishes executing, ensuring no memory pressure between workflows but offering no persistence across separate prompt calls.

How does the LRU cache decide which entries to remove?

The LRU cache tracks a generation counter that increments at the start of each prompt. When the cache exceeds the size limit specified by --cache-lru, it removes entries from the oldest generation first. This ensures that nodes not accessed in recent prompts are evicted before more recently used results.

What happens when RAM-Pressure cache detects low memory?

When available RAM drops below the headroom specified by --cache-ram, the cache calculates an OOM-score for each entry based on its age and memory footprint. It then deletes entries with the lowest scores (typically old entries consuming large amounts of CPU memory) until sufficient RAM is reclaimed, using a hysteresis multiplier to prevent rapid cycling.

Can I disable caching entirely?

Yes. Launch ComfyUI with the --cache-none flag to disable all caching. This forces every node to recompute its output on every execution, which is useful for debugging node implementations or verifying that workflows produce deterministic results without cache interference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →