How Prefect Caching Works: Cache Policies and Cache Key Functions Explained
Prefect's caching mechanism uses composable cache policies that convert task runtime context into deterministic cache keys, enabling idempotent execution by short-circuiting tasks when matching prior results exist.
Prefect's caching system centers on policy-driven architecture that determines when task results should be reused across flow runs. By configuring cache policies and custom cache key functions in the PrefectHQ/prefect repository, developers gain precise control over result persistence, storage backends, and computational efficiency. The implementation spans src/prefect/cache_policies.py and src/prefect/tasks.py, providing both high-level decorators and low-level key computation hooks.
How Cache Policies Determine the Effective Caching Strategy
When a task is instantiated, Prefect resolves the effective caching strategy through a specific precedence order defined in Task.__init__.
Task Constructor Resolution Logic
In src/prefect/tasks.py (lines 511-525), the constructor inspects both cache_policy and cache_key_fn parameters. If a user supplies a custom cache_key_fn, it takes precedence and is automatically wrapped in a CacheKeyFnPolicy via CachePolicy.from_cache_key_fn. This policy object standardizes the interface between custom functions and the broader policy ecosystem.
If no explicit policy is provided and result persistence is enabled, Prefect falls back to the DEFAULT policy (lines 55-57). However, if persist_result=False is set, the policy is forced to NO_CACHE (a _None policy) regardless of other settings, and a warning is emitted to indicate the conflict (lines 49-55).
The Default Cache Policy Composition
The DEFAULT cache policy is not a monolithic object but a composed strategy defined in src/prefect/cache_policies.py (lines 12-13) as:
DEFAULT = INPUTS + TASK_SOURCE + RUN_ID
The + operator creates a CompoundCachePolicy, which validates that component policies have compatible storage, isolation, and lock settings before merging them (lines 20-55). This validation ensures that combined policies can safely share cache storage without configuration conflicts.
Built-in Cache Policy Types and Their Hashing Behavior
Individual policies contribute specific deterministic components to the final cache key through their compute_key implementations.
Inputs Policy
The Inputs policy hashes the runtime inputs passed to the task. It supports excluding specific keys via the exclude parameter and applies stable transforms for known mutable types like pandas DataFrames to ensure deterministic hashing. If inputs are empty or unhashable, it returns None. The implementation resides in Inputs.compute_key (lines 58-86).
TaskSource Policy
TaskSource ensures cache invalidation when task logic changes by hashing the function's source code. It captures the source at task creation time using inspect.getsource and handles the hashing in TaskSource.compute_key (lines 88-106).
RunId Policy
The RunId policy scopes cache keys to specific execution contexts by returning the current flow-run ID, or the task-run ID if no flow context exists. This prevents cache hits across unrelated flow runs. See RunId.compute_key (lines 108-122).
CacheKeyFnPolicy
When users provide a custom cache_key_fn, Prefect wraps it in a CacheKeyFnPolicy. This policy calls the user function with the TaskRunContext and raw inputs dictionary, expecting a string or None return value. The implementation is in CacheKeyFnPolicy.compute_key (lines 70-78).
NO_CACHE Policy
The _None policy (exposed as NO_CACHE) always yields None from its compute_key method (lines 62-68), effectively disabling caching by ensuring no cache key is ever generated or matched.
CompoundCachePolicy
When policies are combined using the + operator, CompoundCachePolicy coordinates key generation. It computes each sub-policy's key, discards any None values, sorts the remaining keys, and hashes the final tuple to produce a single deterministic string. This logic is implemented in CompoundCachePolicy.compute_key (lines 90-108).
The Cache Key Generation Flow
During task execution setup, specifically within Task.create_run, Prefect invokes the resolved policy's compute_key method, passing the TaskRunContext, resolved inputs, and flow parameters. The resulting string is stored on the task run object.
The engine then queries for a previously completed task state matching this key. If a match exists, the task execution is short-circuited and the prior result is restored without re-running the task function. This lookup occurs before the task enters the RUNNING state, ensuring true idempotency.
Customizing Cache Behavior
Implementing Custom Cache Key Functions
To override policy-based hashing entirely, provide a function to the cache_key_fn parameter:
from prefect import task, TaskRunContext
def my_key(ctx: TaskRunContext, inputs: dict) -> str | None:
# Cache only on the first argument, ignore everything else
return f"my-task-{inputs.get('first_arg')}"
@task(cache_key_fn=my_key)
def expensive_op(first_arg: int, noisy_arg: str):
# `noisy_arg` will never affect caching
return first_arg ** 2
The function receives the TaskRunContext and the raw input dictionary, allowing complete control over key generation logic.
Excluding Volatile Inputs
For cases where specific arguments change frequently but should not invalidate the cache, use the subtraction operator or explicit exclusion:
from prefect import task
from prefect.cache_policies import INPUTS
# Method 1: Subtract from existing policy
my_policy = INPUTS - "timestamp"
@task(cache_policy=my_policy)
def compute(value: int, timestamp: datetime.datetime):
# `timestamp` changes each run → excluded from cache key
return value * 2
Storage and Locking Configuration
Policies support advanced configuration for distributed scenarios. You can specify key_storage (accepting block references or filesystem paths) and configure lock_manager or isolation_level parameters to guarantee atomicity when multiple workers may write the same cache key (src/prefect/cache_policies.py lines 72-82).
Disabling Caching
To completely disable caching for a task, either set persist_result=False or explicitly pass cache_policy=NO_CACHE. When disabled, the engine ignores any cache key even if a custom cache_key_fn is supplied.
Stable Transforms for Deterministic Hashing
Some Python types, notably pandas DataFrames, produce non-deterministic byte streams when serialized with cloudpickle. Prefect addresses this by registering deterministic transforms in _register_stable_transforms (src/prefect/cache_policies.py lines 33-47). These transforms ensure that identical data structures generate identical hash values across different Python processes and runs, preventing unnecessary cache misses.
Summary
- Policy resolution follows a strict order: custom
cache_key_fntakes precedence, followed by explicitcache_policy, withDEFAULTas the fallback when persistence is enabled. - Default composition combines
INPUTS + TASK_SOURCE + RUN_IDto create robust cache keys that invalidate when inputs, code, or run context change. - Individual policies handle specific concerns: input hashing, source code validation, run scoping, or custom logic via user-provided functions.
- Key generation occurs in
Task.create_runand drives the engine's lookup of prior completed states to enable short-circuit execution. - Customization supports input exclusion, custom key functions, storage backends, and distributed locking for atomic cache writes.
Frequently Asked Questions
What is the difference between cache_policy and cache_key_fn in Prefect?
The cache_policy parameter accepts a policy object (like INPUTS, TASK_SOURCE, or NO_CACHE) that defines how the cache key is computed using Prefect's built-in rules. The cache_key_fn parameter accepts a custom Python function that receives the task context and inputs, allowing you to define the key string manually. When cache_key_fn is provided, it is wrapped in a CacheKeyFnPolicy and takes precedence over any cache_policy setting.
How does Prefect handle caching when persist_result is set to False?
When persist_result=False is configured on a task, Prefect forces the cache policy to NO_CACHE (the _None policy) regardless of any other configuration, as implemented in src/prefect/tasks.py (lines 49-55). The system emits a warning to indicate that the explicit persist_result=False setting is overriding caching behavior, ensuring no cache keys are generated or stored.
Can I exclude specific function arguments from the cache key calculation?
Yes, you can exclude volatile or irrelevant arguments using the subtraction operator on the INPUTS policy (e.g., INPUTS - "param_name") or by instantiating Inputs(exclude=[...]) with a list of keys to omit. This is useful for timestamps, random seeds, or other arguments that change between runs but should not trigger cache invalidation.
What happens when multiple cache policies are combined with the + operator?
Combining policies with + creates a CompoundCachePolicy that validates compatibility between storage, isolation, and lock settings before merging. During key generation, it computes keys for each component policy, filters out None values, sorts the remaining keys, and hashes them as a tuple to produce the final cache key. This ensures deterministic behavior while allowing flexible policy composition.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →