How to Deploy and Manage Pathway Pipelines Using Kubernetes: Complete Guide
Deploy Pathway pipelines on Kubernetes by containerizing your Python code, configuring runtime arguments via the PATHWAY_SPAWN_ARGS environment variable using the spawn-from-env CLI, and orchestrating with a StatefulSet for distributed execution and persistent state.
Pathway is a Python stream-processing framework that compiles pipelines into in-memory data-flow graphs executed by the Pathway runtime. Because the runtime is packaged as a pure-Python library available on PyPI, you can deploy and manage Pathway pipelines using Kubernetes anywhere a container runtime exists. According to the pathwaycom/pathway source code, the framework natively supports distributed execution across Kubernetes pods with automatic discovery and OpenTelemetry observability.
Containerizing Pathway Pipelines
Pathway pipelines are standard Python programs, so you package them using the same container workflow as any Python application. The official approach uses a Docker image that includes the Pathway library and your pipeline code, then invokes the spawn-from-env entrypoint to read dynamic configuration at runtime.
As documented in docs/2.developers/4.user-guide/60.deployment/32.nebius-deploy.md, this containerization strategy allows the same image to run in Nebius, Azure, AWS, or on-premise Kubernetes clusters without modification.
FROM python:3.11-slim
# Install Pathway (the library is published on PyPI)
RUN pip install pathway=={{ latest_version }}
# Copy your pipeline source
COPY my_pipeline.py /app/
WORKDIR /app
# Default command runs the Pathway CLI `spawn-from-env`
ENTRYPOINT ["pathway", "spawn-from-env"]
Configuring Runtime Arguments with spawn-from-env
The pathway spawn-from-env command, introduced in CHANGELOG.md at lines 527-531, reads the PATHWAY_SPAWN_ARGS environment variable to determine the pipeline entry point, input connectors, and output sinks. This pattern decouples configuration from the container image, allowing you to inject secrets and pipeline parameters via Kubernetes Secrets and ConfigMaps without rebuilding the image.
export PATHWAY_SPAWN_ARGS="run my_pipeline.py --source kafka --sink prometheus"
# If you have a Pathway Scale license
export PATHWAY_LICENSE_KEY="YOUR-LICENSE-KEY"
Kubernetes Deployment Architecture
Using StatefulSets for Stable Identity
For production deployments, Pathway recommends a StatefulSet rather than a Deployment. As stated in docs/2.developers/4.user-guide/60.deployment/10.cloud-deployment.md at lines 31-33, a multi-server deployment "assumes a stateful set deployment with all pods present for a successful operation." StatefulSets provide stable network identities and persistent volumes required for external state retention and consistent pod-to-pod communication.
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: pathway-pipeline
spec:
serviceName: pathway
replicas: 3 # number of workers
selector:
matchLabels:
app: pathway
template:
metadata:
labels:
app: pathway
spec:
containers:
- name: pathway
image: ghcr.io/pathwaycom/pathway:latest
env:
- name: PATHWAY_SPAWN_ARGS
value: "run my_pipeline.py --source kafka --sink prometheus"
- name: PATHWAY_LICENSE_KEY
valueFrom:
secretKeyRef:
name: pathway-license
key: key
ports:
- containerPort: 8080 # optional HTTP endpoint
resources:
limits:
memory: "4Gi"
cpu: "2000m"
# Optional: mount a PersistentVolume for external persistence
volumeClaimTemplates:
- metadata:
name: pathway-data
spec:
accessModes: [ "ReadWriteOnce" ]
resources:
requests:
storage: 10Gi
Enabling Distributed Execution
When running multiple replicas, Pathway automatically discovers pods via the Kubernetes API and shards the computation across the cluster. As noted in README.md at lines 360-362, Pathway "natively supports … distributed using Kubernetes." Each worker processes a partition of the data-flow graph, and the runtime handles synchronization automatically. Ensure all pods are running simultaneously, as the distributed engine expects the full StatefulSet to be present.
Observability and Monitoring
Pathway emits OpenTelemetry traces, metrics, and logs by default. To collect these signals in Kubernetes, deploy an OpenTelemetry Collector using the official Helm chart, as shown in docs/2.developers/4.user-guide/60.deployment/50.pathway-monitoring.md at lines 36-37. The collector receives OTLP data from Pathway containers and forwards it to Grafana, Prometheus, Loki, or any compatible backend.
Deploying the OpenTelemetry Collector
Configure the collector as a sidecar or independent deployment within your Helm release. The following values.yaml combines the Pathway StatefulSet with an OpenTelemetry Collector pipeline for debugging or production monitoring.
pathway:
image: ghcr.io/pathwaycom/pathway:latest
replicaCount: 3
env:
PATHWAY_SPAWN_ARGS: "run my_pipeline.py --source kafka --sink prometheus"
PATHWAY_LICENSE_KEY: "<YOUR-LICENSE-KEY>"
otelCollector:
enabled: true
config:
receivers:
otlp:
protocols:
grpc:
exporters:
debug:
verbosity: detailed
service:
pipelines:
traces:
receivers: [otlp]
exporters: [debug]
metrics:
receivers: [otlp]
exporters: [debug]
logs:
receivers: [otlp]
exporters: [debug]
Local Testing Before Deployment
Validate your container configuration locally before applying it to Kubernetes. Build the image and run it with the same environment variables you will use in production to verify the spawn-from-env behavior and pipeline logic.
# Build the image locally
docker build -t my/pathway-pipeline .
# Run it with the same env variables you will use in K8s
docker run -e PATHWAY_SPAWN_ARGS="run my_pipeline.py --source kafka --sink prometheus" \
-e PATHWAY_LICENSE_KEY="YOUR-LICENSE-KEY" \
my/pathway-pipeline
Summary
- Containerize your pipeline using a Docker image with Pathway installed from PyPI and the
spawn-from-enventrypoint to enable runtime configuration. - Configure behavior via the
PATHWAY_SPAWN_ARGSenvironment variable, allowing you to inject pipeline parameters and secrets through Kubernetes Secrets without image rebuilds. - Deploy using a Kubernetes StatefulSet to provide stable pod identities and persistent volumes required for distributed stream processing.
- Scale horizontally by increasing the
replicascount; Pathway automatically discovers pods and shards the data-flow graph across the cluster. - Monitor by deploying an OpenTelemetry Collector via Helm to capture traces, metrics, and logs, forwarding them to OTLP-compatible backends like Grafana or Prometheus.
Frequently Asked Questions
Why must I use a StatefulSet instead of a Deployment for Pathway?
Pathway's distributed execution engine assumes a StatefulSet with all pods present for successful operation, as documented in docs/2.developers/4.user-guide/60.deployment/10.cloud-deployment.md. StatefulSets provide stable network identities and persistent volume claims necessary for stateful stream processing, whereas Deployments are better suited for stateless workloads where pod identity is ephemeral.
How does the spawn-from-env command handle configuration?
The pathway spawn-from-env CLI command, detailed in CHANGELOG.md and the Nebius deployment guide, reads the PATHWAY_SPAWN_ARGS environment variable to determine the Python file to run and its CLI arguments. This allows you to store the pipeline entry point, connector configurations, and license keys in Kubernetes Secrets or ConfigMaps, keeping the container image generic and reusable across environments.
Does Pathway support automatic horizontal scaling in Kubernetes?
Yes. Pathway natively supports distributed execution using Kubernetes, as stated in README.md. When you increase the replica count in your StatefulSet, the runtime automatically discovers new pods via the Kubernetes API and redistributes the data-flow graph partitions across the cluster. Ensure you maintain the PATHWAY_LICENSE_KEY for Scale licenses when scaling beyond single-node limits.
How do I integrate Pathway with existing monitoring stacks?
Pathway emits OpenTelemetry signals (traces, metrics, logs) that any OTLP-compatible collector can ingest. Deploy the OpenTelemetry Collector using the official Helm chart referenced in docs/2.developers/4.user-guide/60.deployment/50.pathway-monitoring.md, and configure exporters to forward data to your existing Prometheus, Grafana, or Loki instances for unified observability of your streaming pipelines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →