Best Practices for Deploying Pathway Apps on Kubernetes and Cloud Containers
Pathway applications deploy as standard OCI-compliant Docker containers that expose an HTTP API on port 8080, making them compatible with Kubernetes, Docker Compose, and managed cloud container services without modification.
Deploying Pathway apps on Kubernetes or cloud containers follows standard containerization patterns defined in the pathwaycom/llm-app repository. Each template ships with a production-ready Dockerfile that bundles the Rust-based Pathway engine, Python runtime, and your pipeline code into a portable image. The following practices ensure your deployment is reproducible, secure, and observable at scale.
Build a Reproducible Container Image
The foundation of any robust deployment starts with the Dockerfile located at templates/*/Dockerfile. This file defines a multi-stage build that begins from the official pathwaycom/pathway base image.
Pin Base Image Versions
Always pin the base image tag rather than relying on latest to avoid unexpected runtime changes:
# Pin to a specific version for reproducibility
FROM pathwaycom/pathway:v0.13.0
Optimize Layer Caching
Clean package manager caches and use pip with --no-cache-dir to minimize image size and attack surface:
RUN apt-get update \
&& apt-get install -y python3-opencv \
&& rm -rf /var/lib/apt/lists/* /var/cache/apt/archives/*
COPY requirements.txt .
RUN pip install --pre -U --no-cache-dir -r requirements.txt
Exclude Unnecessary Files
Add a .dockerignore file to prevent build context bloat and protect sensitive data:
data/
__pycache__/
.git/
.env
*.md
Deploy to Kubernetes
Pathway containers run as standard Kubernetes workloads. The repository's docker-compose.yml patterns translate directly to Kubernetes manifests.
Create a Highly Available Deployment
Define a Deployment with at least two replicas to ensure availability during updates or node failures:
apiVersion: apps/v1
kind: Deployment
metadata:
name: pathway-app
spec:
replicas: 2
selector:
matchLabels:
app: pathway-app
template:
metadata:
labels:
app: pathway-app
spec:
containers:
- name: pathway
image: ghcr.io/pathwaycom/llm-app:latest
ports:
- containerPort: 8080
resources:
requests:
cpu: "250m"
memory: "256Mi"
limits:
cpu: "500m"
memory: "512Mi"
Configure Health Probes
Pathway apps expose a /healthz endpoint by default. Configure liveness and readiness probes to enable automatic restarts and traffic management:
readinessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 15
periodSeconds: 20
Manage Secrets and Configuration
Never bake API keys into images. Pass sensitive values via Kubernetes Secret objects:
env:
- name: OPENAI_API_KEY
valueFrom:
secretKeyRef:
name: openai-secret
key: api-key
Mount Persistent Volumes
For pipelines that ingest static files (PDFs, CSVs), mount a PersistentVolumeClaim to /app/data, mirroring the Docker Compose volume pattern found in templates/*/docker-compose.yml:
volumeMounts:
- name: data-volume
mountPath: /app/data
volumes:
- name: data-volume
persistentVolumeClaim:
claimName: pathway-data-pvc
Enforce Security Contexts
Run containers as non-root users and drop unnecessary capabilities:
securityContext:
runAsNonRoot: true
runAsUser: 1000
capabilities:
drop:
- ALL
Run on Managed Cloud Containers
The same OCI image deploys to serverless container platforms without modification. The repository READMEs link to official guides for each provider.
Google Cloud Platform
Deploy to Cloud Run for automatic scaling to zero:
gcloud run deploy pathway-app \
--image gcr.io/PROJECT/pathway-app:latest \
--port 8080 \
--set-env-vars "OPENAI_API_KEY=secret" \
--allow-unauthenticated
For persistent workloads, use GKE with the Kubernetes manifests above.
Amazon Web Services
Push the image to ECR and run on Fargate:
aws ecr get-login-password | docker login --username AWS --password-stdin ACCOUNT.dkr.ecr.REGION.amazonaws.com
docker tag pathway-app:latest ACCOUNT.dkr.ecr.REGION.amazonaws.com/pathway-app:latest
docker push ACCOUNT.dkr.ecr.REGION.amazonaws.com/pathway-app:latest
Create an ECS task definition referencing the ECR image with desired count ≥ 2, or deploy to EKS using the same Kubernetes YAML.
Microsoft Azure
Push to Azure Container Registry (ACR) and deploy to Azure Container Apps:
az acr build --registry myregistry --image pathway-app:latest .
az containerapp create \
--name pathway-app \
--resource-group mygroup \
--image myregistry.azurecr.io/pathway-app:latest \
--target-port 8080 \
--ingress external
For AKS deployments, apply the standard Kubernetes manifests.
Observability and Monitoring
Pathway exposes standard interfaces for cloud-native observability.
Logging
Pathway writes structured logs to stdout, which Kubernetes captures automatically. Stream logs using:
kubectl logs -f deployment/pathway-app
Metrics
Enable the Prometheus-compatible /metrics endpoint by setting the environment variable:
env:
- name: PATHWAY_METRICS
value: "true"
Scrape this endpoint with Prometheus or cloud monitoring tools.
Distributed Tracing
Export traces to Jaeger or compatible collectors:
env:
- name: PATHWAY_TRACING_ENDPOINT
value: "http://jaeger-collector:4317"
CI/CD Pipeline
Automate the build and deployment process using GitHub Actions, leveraging the existing workflow at .github/workflows/python-lint.yml.
Continuous Integration
Extend the linting workflow to build and test the image:
name: CI
on: [push, pull_request]
jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Build Docker image
run: docker build -t pathway-app:${{ github.sha }} .
- name: Test container
run: docker run --rm pathway-app:${{ github.sha }} python -c "import pathway; print(pathway.__version__)"
Continuous Deployment
Push the image to a registry and update the Kubernetes deployment:
name: CD
on:
push:
branches: [main]
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Login to Registry
run: echo ${{ secrets.GITHUB_TOKEN }} | docker login ghcr.io -u ${{ github.actor }} --password-stdin
- name: Build and Push
run: |
docker build -t ghcr.io/pathwaycom/llm-app:${{ github.sha }} .
docker push ghcr.io/pathwaycom/llm-app:${{ github.sha }}
- name: Update K8s Deployment
run: |
kubectl set image deployment/pathway-app pathway=ghcr.io/pathwaycom/llm-app:${{ github.sha }}
kubectl rollout status deployment/pathway-app
Summary
- Pathway apps are standard containers that expose port 8080 and run via
python app.py, making them compatible with any OCI-compliant platform. - Pin your base image to a specific version like
pathwaycom/pathway:v0.13.0and clean package caches to ensure reproducible, minimal images. - Use Kubernetes best practices: run 2+ replicas, configure liveness/readiness probes on
/healthz, mount PVCs for static data, and inject secrets via environment variables. - Deploy to managed clouds (GCP Cloud Run, AWS Fargate, Azure Container Apps) using the same container image without modification.
- Enable observability by capturing stdout logs, exposing the
/metricsendpoint, and optionally configuring distributed tracing. - Automate with CI/CD by extending the existing GitHub Actions workflow to build images and update Kubernetes deployments on every commit.
Frequently Asked Questions
How do I handle secrets like API keys in a Pathway Kubernetes deployment?
Never hardcode secrets in your Dockerfile or source code. Instead, create a Kubernetes Secret object and reference it in your Deployment manifest using valueFrom.secretKeyRef. For example, mount your OpenAI API key as an environment variable sourced from a secret named openai-secret. This keeps credentials out of images and allows rotation without rebuilding containers.
Can I run Pathway apps on serverless platforms like Google Cloud Run?
Yes. Pathway containers follow the OCI spec and expose a single HTTP port (8080), making them ideal for serverless platforms. Deploy to Cloud Run using gcloud run deploy with your container image. Note that Cloud Run scales to zero, which is suitable for development or low-traffic scenarios, but for persistent streaming pipelines, use GKE or Cloud Run with minimum instances set to 1.
What is the recommended resource allocation for Pathway containers in production?
Resource requirements depend on your pipeline complexity and LLM model size. As a baseline, request 250m CPU and 256Mi memory per replica, with limits set to 500m CPU and 512Mi memory for lightweight pipelines. If your app loads large language models into memory (e.g., local LLM inference), increase memory limits accordingly—often to several GiB. Always set resource requests to ensure the scheduler places pods on nodes with adequate capacity.
How do I enable monitoring and health checks for my Pathway deployment?
Pathway exposes a /healthz endpoint by default for Kubernetes liveness and readiness probes. Configure these in your Deployment manifest to ensure unhealthy pods restart automatically. For metrics, set the environment variable PATHWAY_METRICS=true to enable the Prometheus-compatible /metrics endpoint. Additionally, Pathway writes structured logs to stdout, which Kubernetes captures automatically—stream them with kubectl logs or forward to your centralized logging stack.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →