PersistenceMode.UDF_CACHING vs. Other Persistence Modes in Pathway
PersistenceMode.UDF_CACHING caches only the outputs of expensive user-defined functions (like LLM calls) to reduce API costs and latency, while PersistenceMode.PERSISTING saves the entire pipeline state for full resume capability.
Pathway's persistence layer allows LLM applications to optimize performance and cost by storing intermediate results between pipeline runs. In the pathwaycom/llm-app repository, the PersistenceMode enum controls whether and how the engine caches data, with UDF_CACHING serving as the default for most LLM templates.
What is PersistenceMode.UDF_CACHING?
PersistenceMode.UDF_CACHING is a targeted persistence strategy that stores only the outputs of user-defined functions (UDFs). When your pipeline calls external services like OpenAI, embedding models, or custom Python functions, Pathway writes the input-output pairs to a filesystem backend (defaulting to ./Cache). On subsequent runs, if the same input hash is encountered, Pathway returns the cached result without invoking the external service.
This mode is implemented in templates/question_answering_rag/app.py, where the App class configures the persistence backend and mode:
import pathway as pw
class App(pw.BaseModel):
persistence_backend: pw.persistence.Backend | None = None
persistence_mode: pw.PersistenceMode | None = pw.PersistenceMode.UDF_CACHING
def run(self) -> None:
# Default to filesystem backend in ./Cache
if self.persistence_backend is None:
persistence_backend = pw.persistence.Backend.filesystem("./Cache")
else:
persistence_backend = self.persistence_backend
persistence_config = pw.persistence.Config(
persistence_backend,
persistence_mode=self.persistence_mode,
)
pw.run(persistence_config=persistence_config)
Comparing Pathway Persistence Modes
Pathway provides three distinct persistence behaviors, each optimized for different operational requirements.
UDF_CACHING
PersistenceMode.UDF_CACHING minimizes API costs and latency by caching only expensive function calls. The rest of the pipeline—including data ingestion, filtering, and joins—recomputes on every run. This ensures fresh data processing while avoiding redundant LLM calls.
- Storage impact: Low (only UDF outputs stored)
- Memory footprint: Minimal (graph recomputes each run)
- Best for: LLM apps with expensive external API calls and frequently changing source data
PERSISTING
PersistenceMode.PERSISTING (referenced in templates/slides_ai_search/app.yaml comments) captures the entire pipeline state, including all intermediate tables and dataframes. When the application restarts, Pathway resumes computation from the last persisted checkpoint without reprocessing any prior data.
# In app.yaml, you can switch modes:
# persistence_mode: !pw.PersistenceMode.PERSISTING
persistence_mode: !pw.PersistenceMode.UDF_CACHING
- Storage impact: High (full data snapshots)
- Memory footprint: Larger (state must be loaded)
- Best for: Long-running batch jobs, incremental data processing, or crash recovery scenarios
NONE
When persistence_mode is set to None or omitted entirely, Pathway operates with no persistence. Every run starts from scratch, invoking all UDFs and recomputing the entire graph. This is useful for development, testing, or when deterministic replay is required.
When to Use Each Persistence Mode
Choose your persistence strategy based on cost constraints, data freshness requirements, and recovery needs:
-
Use
UDF_CACHINGwhen API costs dominate (OpenAI, embedding services) and you want fresh data ingestion on each run. This is the default intemplates/private_rag/app.pyand other LLM templates for good reason—it balances performance with freshness. -
Use
PERSISTINGwhen processing large datasets incrementally or when you need crash recovery. If your pipeline runs for hours and you cannot afford to restart from scratch, full persistence is essential. -
Use
NONEduring active development when you need to verify UDF behavior changes immediately without clearing cache directories manually.
Summary
-
PersistenceMode.UDF_CACHINGcaches only user-defined function outputs (LLM calls, embeddings) to reduce API costs while keeping data processing fresh. It stores data in./Cacheby default and is the recommended mode for most LLM applications in thepathwaycom/llm-apprepository. -
PersistenceMode.PERSISTINGsaves the entire pipeline state, enabling full resume capability after restarts or crashes, but requires significantly more storage. -
No persistence forces complete recomputation on every run, suitable for development but expensive for production LLM usage.
Frequently Asked Questions
How do I clear the UDF cache when using PersistenceMode.UDF_CACHING?
Delete the cache directory specified in your persistence_backend configuration. By default, this is the ./Cache folder in your project root. Pathway will recreate it on the next run. If you need programmatic clearing, instantiate a new backend pointing to a different directory or remove the existing cache files before initializing the App class.
Can I switch between UDF_CACHING and PERSISTING without losing data?
Switching from UDF_CACHING to PERSISTING is safe but starts fresh persistence snapshots, as the two modes store different data structures. Switching from PERSISTING to UDF_CACHING will ignore existing full-state snapshots and only cache UDF outputs going forward. Always test mode changes in a staging environment, as downstream operators may behave differently when receiving fully persisted versus recomputed inputs.
Does UDF_CACHING work with distributed or cloud storage backends?
Yes. While the pathwaycom/llm-app templates default to pw.persistence.Backend.filesystem("./Cache"), the Backend abstraction supports custom implementations. You can configure S3, GCS, or other object storage by providing a backend instance that implements the required persistence interface. This allows multiple workers or container restarts to share the same UDF cache across distributed deployments.
Why is UDF_CACHING the default in LLM templates instead of full PERSISTING?
UDF_CACHING strikes the optimal balance for LLM applications: it eliminates redundant API calls (the most expensive operation) while ensuring that data ingestion and preprocessing steps run fresh on each execution. Full persistence would capture potentially sensitive raw documents and require significantly more storage, which is unnecessary when the primary goal is to avoid repeated LLM inference costs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →