# How to Deploy and Manage Pathway Pipelines Using Kubernetes: Complete Guide

> Deploy and manage Pathway pipelines on Kubernetes. Containerize Python code, configure args via PATHWAY_SPAWN_ARGS, and orchestrate with StatefulSets for distributed execution and persistent state.

- Repository: [Pathway/pathway](https://github.com/pathwaycom/pathway)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Deploy Pathway pipelines on Kubernetes by containerizing your Python code, configuring runtime arguments via the `PATHWAY_SPAWN_ARGS` environment variable using the `spawn-from-env` CLI, and orchestrating with a StatefulSet for distributed execution and persistent state.**

Pathway is a Python stream-processing framework that compiles pipelines into in-memory data-flow graphs executed by the Pathway runtime. Because the runtime is packaged as a pure-Python library available on PyPI, you can deploy and manage Pathway pipelines using Kubernetes anywhere a container runtime exists. According to the `pathwaycom/pathway` source code, the framework natively supports distributed execution across Kubernetes pods with automatic discovery and OpenTelemetry observability.

## Containerizing Pathway Pipelines

Pathway pipelines are standard Python programs, so you package them using the same container workflow as any Python application. The official approach uses a Docker image that includes the Pathway library and your pipeline code, then invokes the `spawn-from-env` entrypoint to read dynamic configuration at runtime.

As documented in [`docs/2.developers/4.user-guide/60.deployment/32.nebius-deploy.md`](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/60.deployment/32.nebius-deploy.md), this containerization strategy allows the same image to run in Nebius, Azure, AWS, or on-premise Kubernetes clusters without modification.

```dockerfile
FROM python:3.11-slim

# Install Pathway (the library is published on PyPI)

RUN pip install pathway=={{ latest_version }}

# Copy your pipeline source

COPY my_pipeline.py /app/
WORKDIR /app

# Default command runs the Pathway CLI `spawn-from-env`

ENTRYPOINT ["pathway", "spawn-from-env"]

```

## Configuring Runtime Arguments with spawn-from-env

The `pathway spawn-from-env` command, introduced in [`CHANGELOG.md`](https://github.com/pathwaycom/pathway/blob/main/CHANGELOG.md) at lines 527-531, reads the `PATHWAY_SPAWN_ARGS` environment variable to determine the pipeline entry point, input connectors, and output sinks. This pattern decouples configuration from the container image, allowing you to inject secrets and pipeline parameters via Kubernetes Secrets and ConfigMaps without rebuilding the image.

```bash
export PATHWAY_SPAWN_ARGS="run my_pipeline.py --source kafka --sink prometheus"

# If you have a Pathway Scale license

export PATHWAY_LICENSE_KEY="YOUR-LICENSE-KEY"

```

## Kubernetes Deployment Architecture

### Using StatefulSets for Stable Identity

For production deployments, Pathway recommends a **StatefulSet** rather than a Deployment. As stated in [`docs/2.developers/4.user-guide/60.deployment/10.cloud-deployment.md`](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/60.deployment/10.cloud-deployment.md) at lines 31-33, a multi-server deployment "assumes a stateful set deployment with all pods present for a successful operation." StatefulSets provide stable network identities and persistent volumes required for external state retention and consistent pod-to-pod communication.

```yaml
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: pathway-pipeline
spec:
  serviceName: pathway
  replicas: 3               # number of workers

  selector:
    matchLabels:
      app: pathway
  template:
    metadata:
      labels:
        app: pathway
    spec:
      containers:
        - name: pathway
          image: ghcr.io/pathwaycom/pathway:latest
          env:
            - name: PATHWAY_SPAWN_ARGS
              value: "run my_pipeline.py --source kafka --sink prometheus"
            - name: PATHWAY_LICENSE_KEY
              valueFrom:
                secretKeyRef:
                  name: pathway-license
                  key: key
          ports:
            - containerPort: 8080   # optional HTTP endpoint

          resources:
            limits:
              memory: "4Gi"
              cpu: "2000m"
      # Optional: mount a PersistentVolume for external persistence

      volumeClaimTemplates:
        - metadata:
            name: pathway-data
          spec:
            accessModes: [ "ReadWriteOnce" ]
            resources:
              requests:
                storage: 10Gi

```

### Enabling Distributed Execution

When running multiple replicas, Pathway automatically discovers pods via the Kubernetes API and shards the computation across the cluster. As noted in [`README.md`](https://github.com/pathwaycom/pathway/blob/main/README.md) at lines 360-362, Pathway "natively supports … distributed using Kubernetes." Each worker processes a partition of the data-flow graph, and the runtime handles synchronization automatically. Ensure all pods are running simultaneously, as the distributed engine expects the full StatefulSet to be present.

## Observability and Monitoring

Pathway emits OpenTelemetry traces, metrics, and logs by default. To collect these signals in Kubernetes, deploy an OpenTelemetry Collector using the official Helm chart, as shown in [`docs/2.developers/4.user-guide/60.deployment/50.pathway-monitoring.md`](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/60.deployment/50.pathway-monitoring.md) at lines 36-37. The collector receives OTLP data from Pathway containers and forwards it to Grafana, Prometheus, Loki, or any compatible backend.

### Deploying the OpenTelemetry Collector

Configure the collector as a sidecar or independent deployment within your Helm release. The following [`values.yaml`](https://github.com/pathwaycom/pathway/blob/main/values.yaml) combines the Pathway StatefulSet with an OpenTelemetry Collector pipeline for debugging or production monitoring.

```yaml
pathway:
  image: ghcr.io/pathwaycom/pathway:latest
  replicaCount: 3
  env:
    PATHWAY_SPAWN_ARGS: "run my_pipeline.py --source kafka --sink prometheus"
    PATHWAY_LICENSE_KEY: "<YOUR-LICENSE-KEY>"

otelCollector:
  enabled: true
  config:
    receivers:
      otlp:
        protocols:
          grpc:
    exporters:
      debug:
        verbosity: detailed
    service:
      pipelines:
        traces:
          receivers: [otlp]
          exporters: [debug]
        metrics:
          receivers: [otlp]
          exporters: [debug]
        logs:
          receivers: [otlp]
          exporters: [debug]

```

## Local Testing Before Deployment

Validate your container configuration locally before applying it to Kubernetes. Build the image and run it with the same environment variables you will use in production to verify the `spawn-from-env` behavior and pipeline logic.

```bash

# Build the image locally

docker build -t my/pathway-pipeline .

# Run it with the same env variables you will use in K8s

docker run -e PATHWAY_SPAWN_ARGS="run my_pipeline.py --source kafka --sink prometheus" \
           -e PATHWAY_LICENSE_KEY="YOUR-LICENSE-KEY" \
           my/pathway-pipeline

```

## Summary

- **Containerize** your pipeline using a Docker image with Pathway installed from PyPI and the `spawn-from-env` entrypoint to enable runtime configuration.
- **Configure** behavior via the `PATHWAY_SPAWN_ARGS` environment variable, allowing you to inject pipeline parameters and secrets through Kubernetes Secrets without image rebuilds.
- **Deploy** using a Kubernetes **StatefulSet** to provide stable pod identities and persistent volumes required for distributed stream processing.
- **Scale** horizontally by increasing the `replicas` count; Pathway automatically discovers pods and shards the data-flow graph across the cluster.
- **Monitor** by deploying an OpenTelemetry Collector via Helm to capture traces, metrics, and logs, forwarding them to OTLP-compatible backends like Grafana or Prometheus.

## Frequently Asked Questions

### Why must I use a StatefulSet instead of a Deployment for Pathway?

Pathway's distributed execution engine assumes a **StatefulSet** with all pods present for successful operation, as documented in [`docs/2.developers/4.user-guide/60.deployment/10.cloud-deployment.md`](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/60.deployment/10.cloud-deployment.md). StatefulSets provide stable network identities and persistent volume claims necessary for stateful stream processing, whereas Deployments are better suited for stateless workloads where pod identity is ephemeral.

### How does the `spawn-from-env` command handle configuration?

The `pathway spawn-from-env` CLI command, detailed in [`CHANGELOG.md`](https://github.com/pathwaycom/pathway/blob/main/CHANGELOG.md) and the Nebius deployment guide, reads the `PATHWAY_SPAWN_ARGS` environment variable to determine the Python file to run and its CLI arguments. This allows you to store the pipeline entry point, connector configurations, and license keys in Kubernetes Secrets or ConfigMaps, keeping the container image generic and reusable across environments.

### Does Pathway support automatic horizontal scaling in Kubernetes?

Yes. Pathway natively supports distributed execution using Kubernetes, as stated in [`README.md`](https://github.com/pathwaycom/pathway/blob/main/README.md). When you increase the replica count in your StatefulSet, the runtime automatically discovers new pods via the Kubernetes API and redistributes the data-flow graph partitions across the cluster. Ensure you maintain the `PATHWAY_LICENSE_KEY` for Scale licenses when scaling beyond single-node limits.

### How do I integrate Pathway with existing monitoring stacks?

Pathway emits OpenTelemetry signals (traces, metrics, logs) that any OTLP-compatible collector can ingest. Deploy the OpenTelemetry Collector using the official Helm chart referenced in [`docs/2.developers/4.user-guide/60.deployment/50.pathway-monitoring.md`](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/60.deployment/50.pathway-monitoring.md), and configure exporters to forward data to your existing Prometheus, Grafana, or Loki instances for unified observability of your streaming pipelines.