How to Implement Multi-Tenant Isolation in Pathway RAG Applications
To implement multi-tenant isolation in Pathway RAG applications, deploy separate pipeline instances per tenant with unique data source credentials, isolated persistence backends, and distinct network ports.
Pathway RAG templates in the pathwaycom/llm-app repository provide self-contained pipelines for indexing documents and serving queries via REST API. Achieving strict multi-tenant isolation requires ensuring each tenant operates within a completely segregated environment, from data ingestion to vector storage. The reference implementation achieves this through configuration-driven separation of data sources, cache directories, and runtime processes.
Core Isolation Strategies
The pathwaycom/llm-app templates support multi-tenancy through four architectural layers. Each layer prevents cross-tenant data leakage while allowing shared codebases.
Isolated Data Source Credentials
Every template's YAML configuration supports distinct connection parameters per tenant. In templates/question_answering_rag/app.yaml (lines 19-27), the SharePoint connector accepts a tenant placeholder that injects tenant-specific values via environment variables.
By configuring unique credentials for each pipeline instance, you ensure that Tenant A's SharePoint site or Google Drive folder remains inaccessible to Tenant B. The connector filters documents at the source, preventing data from entering the wrong pipeline.
Independent Persistence Backends
Vector indexes and UDF caches must remain physically separated. The App class in templates/question_answering_rag/app.py (lines 52-55) defaults to a local folder ./Cache for the persistence_backend.
Override this path per tenant using the persistence_backend parameter. This ensures that vector embeddings, intermediate state, and cached LLM responses never co-mingle on disk.
Runtime Process Isolation
Each App instance wraps pw.run() as shown in templates/question_answering_rag/app.py (lines 34-41). Running separate processes or containers for each tenant creates truly isolated Pathway runtimes with independent memory spaces and event loops.
Dedicated Network Endpoints
The YAML configuration in templates/question_answering_rag/app.yaml (lines 102-104) exposes host and port parameters. Binding each tenant to a distinct port—or routing through a reverse proxy with tenant-specific URL prefixes—ensures API-level separation.
Step-by-Step Implementation
Configuring Tenant-Specific Data Sources
Modify the app.yaml configuration to use environment variables for tenant-specific connection details. This example shows the SharePoint connector configured for dynamic tenant injection:
# templates/question_answering_rag/app.yaml
$source:
- !pw.xpacks.connectors.sharepoint.read
url: $SHAREPOINT_URL
tenant: $SHAREPOINT_TENANT # Unique per tenant
client_id: $SHAREPOINT_CLIENT_ID
cert_path: sharepointcert.pem
thumbprint: $SHAREPOINT_THUMBPRINT
root_path: $SHAREPOINT_ROOT
with_metadata: true
Replace each $SHAREPOINT_… variable with the credentials specific to the tenant that the instance serves. The same pattern applies to Google Drive connectors or local file system paths.
Isolating Vector Indexes with Per-Tenant Caching
In app.py, configure the persistence_backend to use a tenant-specific subdirectory. This prevents vector index collisions and ensures that retrieved documents belong only to the requesting tenant:
# templates/question_answering_rag/app.py
class App(BaseModel):
persistence_backend: pw.persistence.Backend | None = None
def run(self) -> None:
# ...
if persistence_mode is not None:
if self.persistence_backend is None:
# Tenant-specific sub-folder, e.g., ./Cache/tenant_a
persistence_backend = pw.persistence.Backend.filesystem(
f"./Cache/{os.getenv('TENANT_ID', 'default')}"
)
else:
persistence_backend = self.persistence_backend
# ...
Set the TENANT_ID environment variable during deployment to ensure physical separation of on-disk state.
Orchestrating Multi-Tenant Deployments
Use Docker Compose or Kubernetes to run identical container images with tenant-specific configurations. This example demonstrates two isolated tenants sharing the same image but operating on separate ports with distinct caches:
# docker-compose.yml
services:
rag_tenant_a:
image: pathway/llm-app:latest
command: python templates/question_answering_rag/app.py
environment:
- SHAREPOINT_TENANT=tenant_a_id
- TENANT_ID=tenant_a
- PATHWAY_PORT=8000
ports:
- "8000:8000"
rag_tenant_b:
image: pathway/llm-app:latest
command: python templates/question_answering_rag/app.py
environment:
- SHAREPOINT_TENANT=tenant_b_id
- TENANT_ID=tenant_b
- PATHWAY_PORT=8001
ports:
- "8001:8001"
Each service receives its own environment variables, cache directory (./Cache/tenant_a versus ./Cache/tenant_b), and network socket. This pattern scales to any orchestration platform by injecting tenant-specific configuration at runtime.
Summary
- Separate data sources by injecting tenant-specific credentials via environment variables in
app.yaml, preventing cross-tenant document access at the connector level according to thepathwaycom/llm-appimplementation. - Isolate persistence by overriding the
persistence_backendpath with a tenant-specific subdirectory, ensuring vector indexes and UDF caches remain physically separated as configured intemplates/question_answering_rag/app.py. - Run independent instances by deploying separate containers or processes per tenant, leveraging the
Appclass wrapper aroundpw.run()to create isolated Pathway runtimes. - Bind unique ports or use reverse-proxy routing to expose per-tenant API endpoints without port conflicts, utilizing the
hostandportsettings inapp.yaml. - Scale horizontally using the same Docker image across all tenants while relying on orchestration to provide isolation guarantees.
Frequently Asked Questions
Does Pathway RAG support multi-tenancy within a single process?
No. According to the pathwaycom/llm-app source code, multi-tenant isolation requires separate runtime instances. The App class in templates/question_answering_rag/app.py initializes a single pipeline with one persistence_backend and one set of data source credentials. Running multiple tenants in one process would share memory and disk state, violating isolation guarantees.
How do I handle vector index isolation between tenants?
Configure the persistence_backend parameter to use a unique directory per tenant. In templates/question_answering_rag/app.py, the default ./Cache path (lines 52-55) can be overridden by passing a tenant-specific path like ./Cache/{tenant_id} to pw.persistence.Backend.filesystem(). This ensures that vector embeddings and retrieval indexes remain segregated on disk.
Can I use the same Docker image for all tenants?
Yes. The pathwaycom/llm-app templates are designed to accept configuration via environment variables and YAML files. You can deploy the same image across multiple containers, injecting tenant-specific values (such as SHAREPOINT_TENANT, TENANT_ID, and PATHWAY_PORT) at runtime. This approach simplifies updates while maintaining strict isolation boundaries.
What data sources support per-tenant configuration?
All connector types support tenant-specific configuration through YAML templating. The SharePoint connector in templates/question_answering_rag/app.yaml (lines 19-27) demonstrates this with the tenant field, but the same pattern applies to Google Drive, S3 buckets, or local file systems. Each pipeline instance reads its own credentials, ensuring data separation regardless of the source type.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →