# How the ComfyUI Execution Caching System Works: LRU and RAM Pressure Explained

> Understand ComfyUI execution caching, LRU, and RAM pressure. Learn how ComfyUI speeds up workflows by intelligently caching intermediate node outputs to save time and resources.

- Repository: [Comfy Org/ComfyUI](https://github.com/Comfy-Org/ComfyUI)
- Tags: internals
- Published: 2026-02-26

---

**ComfyUI speeds up workflow execution by caching intermediate node outputs using three strategies: Classic (unlimited, prompt-scoped), LRU (fixed-size with generational eviction), and RAM-Pressure (dynamic eviction based on available system memory).**

The execution caching system in ComfyUI eliminates redundant computation by storing node results between prompts. Depending on your hardware constraints and workflow complexity, you can choose between a simple unlimited cache, a size-restricted **LRU** (Least Recently Used) cache, or a **RAM-Pressure** cache that actively monitors system memory. This article breaks down the implementation details found in the `Comfy-Org/ComfyUI` repository, including the exact eviction algorithms and command-line flags that control them.

## Where the Cache Is Initialized

The `PromptExecutor` class in [`execution.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/execution.py) instantiates the caching layer at the start of every prompt. Inside `PromptExecutor.__init__`, the `cache_type` and `cache_args` parameters determine which strategy to use:

```python

# execution.py – PromptExecutor.__init__

class PromptExecutor:
    def __init__(self, server, cache_type=False, cache_args=None):
        self.cache_args = cache_args
        self.cache_type = cache_type
        self.server = server
        self.reset()

```

The `reset()` method constructs a fresh `CacheSet` object, which wraps two distinct stores: **outputs** (for tensors) and **objects** (for model handles). The `CacheSet.__init__` method routes to the appropriate initialization logic based on the `cache_type` enum:

```python

# execution.py – CacheSet.__init__

class CacheSet:
    def __init__(self, cache_type=None, cache_args={}):
        if cache_type == CacheType.NONE:
            self.init_null_cache()
        elif cache_type == CacheType.RAM_PRESSURE:
            self.init_ram_cache(cache_args.get("ram", 16.0))
        elif cache_type == CacheType.LRU:
            self.init_lru_cache(cache_args.get("lru", 0))
        else:
            self.init_classic_cache()
        self.all = [self.outputs, self.objects]

```

## The Three Caching Strategies

ComfyUI implements three distinct caching behaviors, each optimized for different memory and performance trade-offs.

### Classic Cache (Default Behavior)

When no `--cache-*` flags are provided, ComfyUI uses the **Classic** cache. This strategy stores results in a `HierarchicalCache` keyed by a node's **input signature**—a hash derived from the node's class type, input values, and "is_changed" flags. The cache persists for the duration of the prompt and is automatically cleared when execution completes, meaning it never evicts entries mid-prompt regardless of memory usage.

```python

# execution.py – CacheSet.init_classic_cache

def init_classic_cache(self):
    self.outputs = HierarchicalCache(CacheKeySetInputSignature)
    self.objects = HierarchicalCache(CacheKeySetID)

```

### LRU Cache (--cache-lru)

The **LRU** strategy wraps the classic cache with a hard limit on the number of stored entries. Enable it with `--cache-lru N`, where **N** represents the maximum number of node results to retain. The implementation in [`comfy_execution/caching.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy_execution/caching.py) uses a generation counter that increments at the start of each prompt to track recency:

```python

# comfy_execution/caching.py – LRUCache.__init__

class LRUCache(BasicCache):
    def __init__(self, key_class, max_size=100):
        super().__init__(key_class)
        self.max_size = max_size
        self.min_generation = 0
        self.generation = 0
        self.used_generation = {}

```

When a node is accessed, `_mark_used` records the current generation. If the cache exceeds `max_size`, `clean_unused` removes entries from the oldest generation until the count falls below the limit:

```python

# LRUCache.clean_unused

while len(self.cache) > self.max_size and self.min_generation < self.generation:
    self.min_generation += 1
    to_remove = [key for key in self.cache if self.used_generation[key] < self.min_generation]
    for key in to_remove:
        del self.cache[key]
        del self.used_generation[key]

```

### RAM-Pressure Cache (--cache-ram)

The **RAM-Pressure** cache dynamically evicts entries based on real-time system memory availability. Activate it with `--cache-ram M`, where **M** specifies the minimum free RAM headroom in gigabytes. This mode extends `LRUCache` but disables size-based eviction (`max_size=0`) in favor of memory-driven cleanup.

```python

# comfy_execution/caching.py – RAMPressureCache.__init__

class RAMPressureCache(LRUCache):
    def __init__(self, key_class):
        super().__init__(key_class, 0)  # size=0 disables LRU limit

        self.timestamps = {}

```

## How RAM-Pressure Eviction Calculates OOM Scores

The RAM-Pressure strategy determines which entries to delete using an **OOM-score** (Out-Of-Memory score) that balances entry age against memory footprint. After each node executes, `PromptExecutor` calls `poll()` on the cache:

```python

# execution.py – main execution loop

self.caches.outputs.poll(ram_headroom=self.cache_args["ram"])

```

The `RAMPressureCache.poll` method checks available RAM via `psutil.virtual_memory().available`. When free memory drops below the configured headroom, it constructs a prioritized clean-list:

```python

# comfy_execution/caching.py – RAMPressureCache.poll

for key, (outputs, _) in self.cache.items():
    oom_score = RAM_CACHE_OLD_WORKFLOW_OOM_MULTIPLIER ** (self.generation - self.used_generation[key])
    ram_usage = RAM_CACHE_DEFAULT_RAM_USAGE
    # Scan outputs for CPU tensors and sum their memory (discounted by 0.5)

    oom_score *= ram_usage
    bisect.insort(clean_list, (oom_score, self.timestamps[key], key))

```

The OOM-score formula combines:
- **Age factor**: Exponential weighting based on how many prompt generations have passed since the entry was last used
- **Memory footprint**: Sum of all CPU tensor sizes (GPU tensors are excluded because CUDA manages that memory separately)

Entries with the lowest scores (old and large) are deleted first until free RAM exceeds the headroom multiplied by `RAM_CACHE_HYSTERESIS` (approximately 1.1), providing a safety margin against thrashing.

## Enabling Caching Modes from the Command Line

ComfyUI exposes these caching strategies through command-line arguments parsed in [`comfy/cli_args.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/comfy/cli_args.py). The flags propagate to `PromptExecutor` via the `cache_type` and `cache_args` parameters.

| Flag | Description | Example |
|------|-------------|---------|
| `--cache-lru N` | Enable LRU cache with maximum **N** entries | `python main.py --cache-lru 200` |
| `--cache-ram M` | Enable RAM-Pressure cache with **M** GB headroom | `python main.py --cache-ram 8` |
| `--cache-none` | Disable all caching (useful for debugging) | `python main.py --cache-none` |

You can also instantiate a `PromptExecutor` programmatically with specific cache settings:

```python
from comfy.execution import PromptExecutor, CacheType

executor = PromptExecutor(
    server=my_server,
    cache_type=CacheType.RAM_PRESSURE,
    cache_args={"ram": 4.0}  # Maintain 4 GB free RAM

)

```

## Summary

- ComfyUI's execution caching system offers three modes: **Classic** (unlimited, prompt-scoped), **LRU** (fixed entry count with generational eviction), and **RAM-Pressure** (dynamic eviction based on available RAM).
- The `CacheSet` class in [`execution.py`](https://github.com/Comfy-Org/ComfyUI/blob/main/execution.py) initializes the appropriate cache implementation during `PromptExecutor.reset()`.
- **LRU eviction** removes the least-recently-used entries by comparing generation counters when the cache exceeds its maximum size.
- **RAM-Pressure eviction** calculates an OOM-score combining entry age and CPU tensor memory usage, deleting low-scoring entries when system memory falls below the configured headroom.
- Control caching behavior via `--cache-lru`, `--cache-ram`, or `--cache-none` command-line flags.

## Frequently Asked Questions

### What is the default caching behavior in ComfyUI?

By default, ComfyUI uses the **Classic** cache, which stores unlimited node outputs for the duration of a single prompt. The cache is cleared automatically when the prompt finishes executing, ensuring no memory pressure between workflows but offering no persistence across separate prompt calls.

### How does the LRU cache decide which entries to remove?

The LRU cache tracks a **generation counter** that increments at the start of each prompt. When the cache exceeds the size limit specified by `--cache-lru`, it removes entries from the oldest generation first. This ensures that nodes not accessed in recent prompts are evicted before more recently used results.

### What happens when RAM-Pressure cache detects low memory?

When available RAM drops below the headroom specified by `--cache-ram`, the cache calculates an **OOM-score** for each entry based on its age and memory footprint. It then deletes entries with the lowest scores (typically old entries consuming large amounts of CPU memory) until sufficient RAM is reclaimed, using a hysteresis multiplier to prevent rapid cycling.

### Can I disable caching entirely?

Yes. Launch ComfyUI with the `--cache-none` flag to disable all caching. This forces every node to recompute its output on every execution, which is useful for debugging node implementations or verifying that workflows produce deterministic results without cache interference.