# Key Differences Between Batch and Streaming Processing Modes in Pathway: A Complete Guide

> Explore the key differences between batch and streaming processing in Pathway. Understand how static and streaming modes impact data ingestion, processing, and output for your applications.

- Repository: [Pathway/pathway](https://github.com/pathwaycom/pathway)
- Tags: deep-dive
- Published: 2026-03-06

---

**Pathway supports two mutually exclusive execution modes—static (batch) mode and streaming mode—each with distinct connector constraints, computation semantics, and runtime behavior that determine how data is ingested, processed, and output.**

Understanding the key differences between batch and streaming processing modes in Pathway is essential for building efficient data pipelines. The `pathwaycom/pathway` repository implements these two mutually exclusive execution models at the connector level, allowing developers to switch between one-off analytics and continuous real-time processing by changing a single parameter.

## Execution Models and Data Ingestion

### Batch (Static) Mode

In batch mode, all input data is read **once** and processed as a single, finite batch. According to the [batch processing documentation](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/10.introduction/80.batch-processing.md), the engine executes the dataflow exactly once and materializes results as a complete static table. Connector calls block until the entire source is consumed, making this mode ideal for one-off analytics, testing, and daily batch jobs.

### Streaming Mode

Streaming mode maintains **continuously polling** connectors that stay open indefinitely, processing new records as they arrive. The engine performs **incremental, stateful** computations where each update triggers downstream recomputation. As implemented in the Pathway engine, this mode runs until manually terminated and automatically updates tables as new data appears, supporting real-time dashboards and event-driven pipelines.

## Connector Compatibility and Constraints

Pathway strictly separates connector types between modes. The connector mode is defined as `Literal["streaming", "static"]` throughout the I/O packages, such as in [[`python/pathway/io/kafka/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/io/kafka/__init__.py)](https://github.com/pathwaycom/pathway/blob/main/python/pathway/io/kafka/__init__.py).

- **Static mode connectors** are restricted to batch processing and defined under the *Static* column in the connectors table.
- **Streaming mode connectors** continuously poll for new data from sources like Kafka or watched directories.
- **Mixing is prohibited**: The documentation in [[`30.connectors-in-pathway.md`](https://github.com/pathwaycom/pathway/blob/main/30.connectors-in-pathway.md)](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/20.connect/30.connectors-in-pathway.md) explicitly states that static and streaming connectors cannot be combined in the same pipeline.

## Computation Semantics and Output Formats

The computation models differ fundamentally between the two modes.

**Batch Mode:**
- Executes dataflow **once** and produces a static snapshot of the final state.
- Time-related operators (e.g., `as_of`, windowed aggregations) are not applicable because logical time does not progress.
- Results materialize as complete tables suitable for debug helpers like `pw.debug.compute_and_print`.

**Streaming Mode:**
- Performs **incremental, stateful** computations with automatic recomputation of downstream operators.
- Outputs are **change-feeds** containing `time` (logical timestamp) and `diff` (`+1` for insert, `-1` for delete) columns.
- Tables update automatically as new data arrives; use `pw.debug.compute_and_print_update_stream` for streaming outputs.

## Practical Implementation: Code Examples

The following examples demonstrate how to implement both modes using the CSV connector in [[`python/pathway/io/csv/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/io/csv/__init__.py)](https://github.com/pathwaycom/pathway/blob/main/python/pathway/io/csv/__init__.py).

### Batch Processing with Static Mode

```python
import pathway as pw

# Static (batch) mode – the file is read once and the pipeline finishes.

t = pw.io.csv.read(
    "data/sales.csv",
    schema=pw.Schema(
        order_id=int,
        amount=float,
        country=str,
    ),
    mode="static",                      # <-- batch mode

)

# Simple aggregation (batch)

total = pw.select(t, total_amount=pw.sum(t.amount))
pw.debug.compute_and_print(total)       # prints a snapshot once

```

### Streaming Processing with Directory Watching

```python
import pathway as pw

# Streaming mode – the connector watches the directory and processes new files as they appear.

t = pw.io.csv.read(
    "data/streaming/",
    schema=pw.Schema(
        order_id=int,
        amount=float,
        country=str,
    ),
    mode="streaming",                   # <-- streaming mode

)

# Incremental aggregation; each new file updates the result.

total = pw.select(t, total_amount=pw.sum(t.amount))

# Use the streaming‑aware compute helper.

pw.debug.compute_and_print_update_stream(total)

```

### Switching Between Modes

```python
import pathway as pw

# Define pipeline once

def pipeline(source_mode):
    t = pw.io.csv.read(
        "data/input/",
        schema=pw.Schema(id=int, value=float),
        mode=source_mode,               # pass "static" or "streaming"

    )
    agg = pw.select(t, sum_val=pw.sum(t.value))
    return agg

# Batch execution

batch_result = pipeline("static")
pw.debug.compute_and_print(batch_result)

# Streaming execution (just change the mode string)

stream_result = pipeline("streaming")
pw.debug.compute_and_print_update_stream(stream_result)

```

## Persistence and Back-filling Behavior

Both modes support **persistence** for storing state, but the back-filling behavior differs:

- **Batch mode**: Persistence stores state between runs, enabling back-filling of only new data on subsequent executions when reprocessing the entire dataset.
- **Streaming mode**: Persistence works identically, but the engine naturally processes only **new increments** as they appear, making back-filling inherently incremental without reprocessing historical data.

## Summary

- Pathway enforces **mutually exclusive** batch and streaming modes at the connector level with no mixing allowed.
- **Static mode** processes finite datasets once and produces static snapshots, while **streaming mode** maintains continuous, stateful computations with change-feed outputs.
- Connectors declare their mode via `Literal["streaming", "static"]` type definitions as seen in the Kafka and CSV I/O packages.
- Switching execution modes requires only changing the `mode` parameter from `"static"` to `"streaming"`—pipeline logic remains unchanged.
- Debug helpers in [[`python/pathway/debug/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/debug/__init__.py)](https://github.com/pathwaycom/pathway/blob/main/python/pathway/debug/__init__.py) differ between modes: `compute_and_print` for batch snapshots, `compute_and_print_update_stream` for streaming change-feeds.

## Frequently Asked Questions

### Can I mix batch and streaming connectors in the same Pathway pipeline?

No. According to the [connector documentation](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/20.connect/30.connectors-in-pathway.md), Pathway strictly prohibits mixing static and streaming connectors. Each pipeline must use exclusively one mode type, enforced by the `Literal["streaming", "static"]` type system in the I/O package definitions.

### How do I switch an existing batch pipeline to streaming?

You only need to change the `mode` parameter from `"static"` to `"streaming"` in your connector definitions. The rest of your pipeline logic—including transformations and aggregations—remains identical. This design allows rapid prototyping on static data before deployment to production streaming environments.

### Why can't I use time-windowed aggregations in batch mode?

Time-related operators like `as_of` and windowed aggregations depend on an ever-increasing logical time that only exists in streaming mode. In batch mode, the engine processes a finite snapshot once, so there is no concept of progressing time or incremental updates to window boundaries.

### What debug helper should I use for streaming outputs?

For streaming pipelines, use `pw.debug.compute_and_print_update_stream()` instead of `pw.debug.compute_and_print()`. The streaming helper outputs change-feeds with `time` and `diff` columns showing insertions and deletions, while the batch helper prints a static snapshot of the final table state.