How to Deploy and Manage Pathway Pipelines Using Kubernetes: Complete Guide

Deploy Pathway pipelines on Kubernetes by containerizing your Python code, configuring runtime arguments via the PATHWAY_SPAWN_ARGS environment variable using the spawn-from-env CLI, and orchestrating with a StatefulSet for distributed execution and persistent state.

Pathway is a Python stream-processing framework that compiles pipelines into in-memory data-flow graphs executed by the Pathway runtime. Because the runtime is packaged as a pure-Python library available on PyPI, you can deploy and manage Pathway pipelines using Kubernetes anywhere a container runtime exists. According to the pathwaycom/pathway source code, the framework natively supports distributed execution across Kubernetes pods with automatic discovery and OpenTelemetry observability.

Containerizing Pathway Pipelines

Pathway pipelines are standard Python programs, so you package them using the same container workflow as any Python application. The official approach uses a Docker image that includes the Pathway library and your pipeline code, then invokes the spawn-from-env entrypoint to read dynamic configuration at runtime.

As documented in docs/2.developers/4.user-guide/60.deployment/32.nebius-deploy.md, this containerization strategy allows the same image to run in Nebius, Azure, AWS, or on-premise Kubernetes clusters without modification.

FROM python:3.11-slim

# Install Pathway (the library is published on PyPI)

RUN pip install pathway=={{ latest_version }}

# Copy your pipeline source

COPY my_pipeline.py /app/
WORKDIR /app

# Default command runs the Pathway CLI `spawn-from-env`

ENTRYPOINT ["pathway", "spawn-from-env"]

Configuring Runtime Arguments with spawn-from-env

The pathway spawn-from-env command, introduced in CHANGELOG.md at lines 527-531, reads the PATHWAY_SPAWN_ARGS environment variable to determine the pipeline entry point, input connectors, and output sinks. This pattern decouples configuration from the container image, allowing you to inject secrets and pipeline parameters via Kubernetes Secrets and ConfigMaps without rebuilding the image.

export PATHWAY_SPAWN_ARGS="run my_pipeline.py --source kafka --sink prometheus"

# If you have a Pathway Scale license

export PATHWAY_LICENSE_KEY="YOUR-LICENSE-KEY"

Kubernetes Deployment Architecture

Using StatefulSets for Stable Identity

For production deployments, Pathway recommends a StatefulSet rather than a Deployment. As stated in docs/2.developers/4.user-guide/60.deployment/10.cloud-deployment.md at lines 31-33, a multi-server deployment "assumes a stateful set deployment with all pods present for a successful operation." StatefulSets provide stable network identities and persistent volumes required for external state retention and consistent pod-to-pod communication.

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: pathway-pipeline
spec:
  serviceName: pathway
  replicas: 3               # number of workers

  selector:
    matchLabels:
      app: pathway
  template:
    metadata:
      labels:
        app: pathway
    spec:
      containers:
        - name: pathway
          image: ghcr.io/pathwaycom/pathway:latest
          env:
            - name: PATHWAY_SPAWN_ARGS
              value: "run my_pipeline.py --source kafka --sink prometheus"
            - name: PATHWAY_LICENSE_KEY
              valueFrom:
                secretKeyRef:
                  name: pathway-license
                  key: key
          ports:
            - containerPort: 8080   # optional HTTP endpoint

          resources:
            limits:
              memory: "4Gi"
              cpu: "2000m"
      # Optional: mount a PersistentVolume for external persistence

      volumeClaimTemplates:
        - metadata:
            name: pathway-data
          spec:
            accessModes: [ "ReadWriteOnce" ]
            resources:
              requests:
                storage: 10Gi

Enabling Distributed Execution

When running multiple replicas, Pathway automatically discovers pods via the Kubernetes API and shards the computation across the cluster. As noted in README.md at lines 360-362, Pathway "natively supports … distributed using Kubernetes." Each worker processes a partition of the data-flow graph, and the runtime handles synchronization automatically. Ensure all pods are running simultaneously, as the distributed engine expects the full StatefulSet to be present.

Observability and Monitoring

Pathway emits OpenTelemetry traces, metrics, and logs by default. To collect these signals in Kubernetes, deploy an OpenTelemetry Collector using the official Helm chart, as shown in docs/2.developers/4.user-guide/60.deployment/50.pathway-monitoring.md at lines 36-37. The collector receives OTLP data from Pathway containers and forwards it to Grafana, Prometheus, Loki, or any compatible backend.

Deploying the OpenTelemetry Collector

Configure the collector as a sidecar or independent deployment within your Helm release. The following values.yaml combines the Pathway StatefulSet with an OpenTelemetry Collector pipeline for debugging or production monitoring.

pathway:
  image: ghcr.io/pathwaycom/pathway:latest
  replicaCount: 3
  env:
    PATHWAY_SPAWN_ARGS: "run my_pipeline.py --source kafka --sink prometheus"
    PATHWAY_LICENSE_KEY: "<YOUR-LICENSE-KEY>"

otelCollector:
  enabled: true
  config:
    receivers:
      otlp:
        protocols:
          grpc:
    exporters:
      debug:
        verbosity: detailed
    service:
      pipelines:
        traces:
          receivers: [otlp]
          exporters: [debug]
        metrics:
          receivers: [otlp]
          exporters: [debug]
        logs:
          receivers: [otlp]
          exporters: [debug]

Local Testing Before Deployment

Validate your container configuration locally before applying it to Kubernetes. Build the image and run it with the same environment variables you will use in production to verify the spawn-from-env behavior and pipeline logic.


# Build the image locally

docker build -t my/pathway-pipeline .

# Run it with the same env variables you will use in K8s

docker run -e PATHWAY_SPAWN_ARGS="run my_pipeline.py --source kafka --sink prometheus" \
           -e PATHWAY_LICENSE_KEY="YOUR-LICENSE-KEY" \
           my/pathway-pipeline

Summary

  • Containerize your pipeline using a Docker image with Pathway installed from PyPI and the spawn-from-env entrypoint to enable runtime configuration.
  • Configure behavior via the PATHWAY_SPAWN_ARGS environment variable, allowing you to inject pipeline parameters and secrets through Kubernetes Secrets without image rebuilds.
  • Deploy using a Kubernetes StatefulSet to provide stable pod identities and persistent volumes required for distributed stream processing.
  • Scale horizontally by increasing the replicas count; Pathway automatically discovers pods and shards the data-flow graph across the cluster.
  • Monitor by deploying an OpenTelemetry Collector via Helm to capture traces, metrics, and logs, forwarding them to OTLP-compatible backends like Grafana or Prometheus.

Frequently Asked Questions

Why must I use a StatefulSet instead of a Deployment for Pathway?

Pathway's distributed execution engine assumes a StatefulSet with all pods present for successful operation, as documented in docs/2.developers/4.user-guide/60.deployment/10.cloud-deployment.md. StatefulSets provide stable network identities and persistent volume claims necessary for stateful stream processing, whereas Deployments are better suited for stateless workloads where pod identity is ephemeral.

How does the spawn-from-env command handle configuration?

The pathway spawn-from-env CLI command, detailed in CHANGELOG.md and the Nebius deployment guide, reads the PATHWAY_SPAWN_ARGS environment variable to determine the Python file to run and its CLI arguments. This allows you to store the pipeline entry point, connector configurations, and license keys in Kubernetes Secrets or ConfigMaps, keeping the container image generic and reusable across environments.

Does Pathway support automatic horizontal scaling in Kubernetes?

Yes. Pathway natively supports distributed execution using Kubernetes, as stated in README.md. When you increase the replica count in your StatefulSet, the runtime automatically discovers new pods via the Kubernetes API and redistributes the data-flow graph partitions across the cluster. Ensure you maintain the PATHWAY_LICENSE_KEY for Scale licenses when scaling beyond single-node limits.

How do I integrate Pathway with existing monitoring stacks?

Pathway emits OpenTelemetry signals (traces, metrics, logs) that any OTLP-compatible collector can ingest. Deploy the OpenTelemetry Collector using the official Helm chart referenced in docs/2.developers/4.user-guide/60.deployment/50.pathway-monitoring.md, and configure exporters to forward data to your existing Prometheus, Grafana, or Loki instances for unified observability of your streaming pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →