How to Configure persistence_backend for Distributed Pathway Deployments: A Complete Guide
Set persistence_backend in your template's app.yaml to a shared storage service (e.g., S3, GCS, or PostgreSQL) and ensure all distributed workers use identical configurations to maintain a coherent cache across the cluster.
The pathwaycom/llm-app repository provides production-ready templates for building LLM applications with Pathway. When you configure persistence_backend for distributed Pathway deployments, you ensure that every worker node shares the same durable storage layer, preventing data loss during scaling events and eliminating redundant computations across the cluster.
Why a Shared persistence_backend Matters in Distributed Deployments
Pathway separates caching (temporary results of UDF calls) from persistence (long-term storage that survives restarts). In a single-node deployment, the default local filesystem backend (./Cache) suffices. However, distributed deployments introduce critical failure modes when each worker maintains an isolated cache:
- Isolated local caches force redundant recomputations across nodes, increasing latency and compute costs.
- Node failure destroys all cached results on that instance, breaking pipeline continuity.
- Horizontal scaling prevents new replicas from benefiting from work already performed by existing nodes.
To eliminate these issues, you must point all replicas to a single, durable persistence store using the persistence_backend configuration option.
Architecture of persistence_backend in the pathwaycom/llm-app Repository
The app.yaml Configuration File
Each template ships with an app.yaml file that exposes optional persistence keys. The persistence_backend entry accepts a YAML tag that resolves to a concrete pw.persistence.Backend subclass at runtime.
# Example fragment from templates/question_answering_rag/app.yaml
persistence_mode: !pw.PersistenceMode.UDF_CACHING
persistence_backend: !pw.persistence.Backend.s3
bucket: "my-shared-cache"
region: "us-east-1"
The app.py Runtime Logic
The app.py file in each template reads the YAML configuration and constructs a pw.persistence.Config object. In templates/question_answering_rag/app.py (lines 53–86), the logic handles backend instantiation:
if persistence_mode is not None:
if self.persistence_backend is None:
# Default fallback – local filesystem under ./Cache
persistence_backend = pw.persistence.Backend.filesystem("./Cache")
else:
persistence_backend = self.persistence_backend
persistence_config = pw.persistence.Config(
persistence_backend,
persistence_mode=persistence_mode,
)
else:
persistence_config = None
When persistence_mode is set, the backend defined in app.yaml propagates to all distributed workers, ensuring every node reads from and writes to the same storage layer.
Choosing a persistence_backend for Distributed Pathway Deployments
Pathway provides several built-in backends suitable for distributed deployments. Select based on your infrastructure requirements:
| Backend | Use Case | YAML Configuration Example |
|---|---|---|
| Filesystem | Local testing, single-node demos | !pw.persistence.Backend.filesystem "./Cache" |
| Amazon S3 | Cloud-native, highly available object storage | !pw.persistence.Backend.s3 bucket: "my-bucket" region: "us-east-1" |
| Google Cloud Storage | GCP-based deployments | !pw.persistence.Backend.gcs bucket: "my-gcs-bucket" |
| Azure Blob Storage | Azure-based deployments | !pw.persistence.Backend.azure_blob account_name: "myaccount" container: "mycontainer" |
| PostgreSQL | Structured, queryable persistence | !pw.persistence.Backend.postgres dsn: "postgresql://user:pwd@host/db" |
| Redis | Low-latency key-value lookups | !pw.persistence.Backend.redis url: "redis://host:6379" |
The concrete class names (Backend.s3, Backend.gcs, etc.) are part of the Pathway SDK. The pathwaycom/llm-app repository forwards the YAML tag directly to the SDK without additional abstraction.
Step-by-Step Configuration Guide
Follow these steps to configure a shared persistence backend for your distributed deployment:
-
Provision shared storage (e.g., create an S3 bucket, GCS bucket, or PostgreSQL instance) accessible from all worker nodes.
-
Configure IAM permissions to allow read/write access from your containers. Use environment variables, IAM roles, or Kubernetes secrets to inject credentials securely.
-
Edit the template's
app.yamlto enable persistence and specify the backend. For example, intemplates/question_answering_rag/app.yaml:persistence_mode: !pw.PersistenceMode.PERSISTING persistence_backend: !pw.persistence.Backend.s3 bucket: "my-qa-rag-cache" region: "eu-central-1" -
Verify all replicas use identical configurations. In Docker Compose or Kubernetes, ensure the
app.yamlis mounted or built into every container image consistently. -
Deploy the stack using your orchestration tool (Docker Compose, Kubernetes, or Pathway deployment scripts). All nodes will instantiate the same
pw.persistence.Configand share the cache via the configured backend.
Complete Configuration Example
app.yaml Configuration
# path: templates/question_answering_rag/app.yaml
persistence_mode: !pw.PersistenceMode.PERSISTING
persistence_backend: !pw.persistence.Backend.s3
bucket: "rag-demo-cache"
region: "us-east-2"
app.py Integration
# path: templates/question_answering_rag/app.py
class App:
persistence_backend: pw.persistence.Backend | None = None
persistence_mode: pw.PersistenceMode | None = pw.PersistenceMode.UDF_CACHING
def build(self):
if self.persistence_mode is not None:
# Use the backend defined in YAML or fall back to local FS
persistence_backend = (
self.persistence_backend
or pw.persistence.Backend.filesystem("./Cache")
)
persistence_config = pw.persistence.Config(
persistence_backend,
persistence_mode=self.persistence_mode,
)
else:
persistence_config = None
# Pass persistence_config to the Pathway pipeline
pipeline = MyPipeline(persistence_config=persistence_config)
When the YAML declares persistence_backend: !pw.persistence.Backend.s3, the self.persistence_backend attribute resolves to an S3-backed instance, ensuring all distributed workers read from and write to the same bucket.
Summary
- Pathway separates caching from persistence; the
persistence_backendsetting controls long-term storage that survives restarts. - Distributed deployments require a shared backend (S3, GCS, PostgreSQL, etc.) to prevent isolated caches and data loss during scaling events.
- Configuration occurs in
app.yamlusing YAML tags like!pw.persistence.Backend.s3, which the template'sapp.pyresolves into apw.persistence.Configobject. - All workers must use identical configurations to ensure cache coherence across the cluster.
Frequently Asked Questions
How does persistence_backend differ from persistence_mode?
persistence_mode determines what gets persisted (e.g., UDF_CACHING for function results or PERSISTING for full state), while persistence_backend determines where it gets stored (e.g., local filesystem, S3, or PostgreSQL). You must set both for active persistence; setting only persistence_mode without a backend results in the default local filesystem backend, which is unsuitable for distributed deployments.
Can I use Redis as a persistence_backend for the llm-app templates?
Yes, Pathway supports Redis via !pw.persistence.Backend.redis. Configure it in your app.yaml by providing the connection URL, such as url: "redis://host:6379". This is ideal for low-latency key-value lookups, though for large-scale distributed caching, object storage like S3 or GCS is typically more cost-effective.
What happens if I don't configure persistence_backend in a distributed deployment?
If you omit persistence_backend while enabling persistence_mode, Pathway defaults to pw.persistence.Backend.filesystem("./Cache"), which writes to the local container filesystem. In a distributed setup, this creates isolated caches on each node, causing redundant computations, inconsistent state between workers, and complete data loss when any container restarts or scales down.
Where are the template configuration files located in the pathwaycom/llm-app repository?
Each template resides under the templates/ directory with its own subdirectory. For example, the Question Answering RAG template configuration is at templates/question_answering_rag/app.yaml with the runtime logic in templates/question_answering_rag/app.py. Other templates follow the same pattern, including templates/private_rag/, templates/multimodal_rag/, templates/adaptive_rag/, and templates/slides_ai_search/.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →