Key Differences Between Batch and Streaming Processing Modes in Pathway: A Complete Guide

Pathway supports two mutually exclusive execution modes—static (batch) mode and streaming mode—each with distinct connector constraints, computation semantics, and runtime behavior that determine how data is ingested, processed, and output.

Understanding the key differences between batch and streaming processing modes in Pathway is essential for building efficient data pipelines. The pathwaycom/pathway repository implements these two mutually exclusive execution models at the connector level, allowing developers to switch between one-off analytics and continuous real-time processing by changing a single parameter.

Execution Models and Data Ingestion

Batch (Static) Mode

In batch mode, all input data is read once and processed as a single, finite batch. According to the batch processing documentation, the engine executes the dataflow exactly once and materializes results as a complete static table. Connector calls block until the entire source is consumed, making this mode ideal for one-off analytics, testing, and daily batch jobs.

Streaming Mode

Streaming mode maintains continuously polling connectors that stay open indefinitely, processing new records as they arrive. The engine performs incremental, stateful computations where each update triggers downstream recomputation. As implemented in the Pathway engine, this mode runs until manually terminated and automatically updates tables as new data appears, supporting real-time dashboards and event-driven pipelines.

Connector Compatibility and Constraints

Pathway strictly separates connector types between modes. The connector mode is defined as Literal["streaming", "static"] throughout the I/O packages, such as in [python/pathway/io/kafka/__init__.py](https://github.com/pathwaycom/pathway/blob/main/python/pathway/io/kafka/__init__.py).

Computation Semantics and Output Formats

The computation models differ fundamentally between the two modes.

Batch Mode:

  • Executes dataflow once and produces a static snapshot of the final state.
  • Time-related operators (e.g., as_of, windowed aggregations) are not applicable because logical time does not progress.
  • Results materialize as complete tables suitable for debug helpers like pw.debug.compute_and_print.

Streaming Mode:

  • Performs incremental, stateful computations with automatic recomputation of downstream operators.
  • Outputs are change-feeds containing time (logical timestamp) and diff (+1 for insert, -1 for delete) columns.
  • Tables update automatically as new data arrives; use pw.debug.compute_and_print_update_stream for streaming outputs.

Practical Implementation: Code Examples

The following examples demonstrate how to implement both modes using the CSV connector in [python/pathway/io/csv/__init__.py](https://github.com/pathwaycom/pathway/blob/main/python/pathway/io/csv/__init__.py).

Batch Processing with Static Mode

import pathway as pw

# Static (batch) mode – the file is read once and the pipeline finishes.

t = pw.io.csv.read(
    "data/sales.csv",
    schema=pw.Schema(
        order_id=int,
        amount=float,
        country=str,
    ),
    mode="static",                      # <-- batch mode

)

# Simple aggregation (batch)

total = pw.select(t, total_amount=pw.sum(t.amount))
pw.debug.compute_and_print(total)       # prints a snapshot once

Streaming Processing with Directory Watching

import pathway as pw

# Streaming mode – the connector watches the directory and processes new files as they appear.

t = pw.io.csv.read(
    "data/streaming/",
    schema=pw.Schema(
        order_id=int,
        amount=float,
        country=str,
    ),
    mode="streaming",                   # <-- streaming mode

)

# Incremental aggregation; each new file updates the result.

total = pw.select(t, total_amount=pw.sum(t.amount))

# Use the streaming‑aware compute helper.

pw.debug.compute_and_print_update_stream(total)

Switching Between Modes

import pathway as pw

# Define pipeline once

def pipeline(source_mode):
    t = pw.io.csv.read(
        "data/input/",
        schema=pw.Schema(id=int, value=float),
        mode=source_mode,               # pass "static" or "streaming"

    )
    agg = pw.select(t, sum_val=pw.sum(t.value))
    return agg

# Batch execution

batch_result = pipeline("static")
pw.debug.compute_and_print(batch_result)

# Streaming execution (just change the mode string)

stream_result = pipeline("streaming")
pw.debug.compute_and_print_update_stream(stream_result)

Persistence and Back-filling Behavior

Both modes support persistence for storing state, but the back-filling behavior differs:

  • Batch mode: Persistence stores state between runs, enabling back-filling of only new data on subsequent executions when reprocessing the entire dataset.
  • Streaming mode: Persistence works identically, but the engine naturally processes only new increments as they appear, making back-filling inherently incremental without reprocessing historical data.

Summary

  • Pathway enforces mutually exclusive batch and streaming modes at the connector level with no mixing allowed.
  • Static mode processes finite datasets once and produces static snapshots, while streaming mode maintains continuous, stateful computations with change-feed outputs.
  • Connectors declare their mode via Literal["streaming", "static"] type definitions as seen in the Kafka and CSV I/O packages.
  • Switching execution modes requires only changing the mode parameter from "static" to "streaming"—pipeline logic remains unchanged.
  • Debug helpers in [python/pathway/debug/__init__.py](https://github.com/pathwaycom/pathway/blob/main/python/pathway/debug/__init__.py) differ between modes: compute_and_print for batch snapshots, compute_and_print_update_stream for streaming change-feeds.

Frequently Asked Questions

Can I mix batch and streaming connectors in the same Pathway pipeline?

No. According to the connector documentation, Pathway strictly prohibits mixing static and streaming connectors. Each pipeline must use exclusively one mode type, enforced by the Literal["streaming", "static"] type system in the I/O package definitions.

How do I switch an existing batch pipeline to streaming?

You only need to change the mode parameter from "static" to "streaming" in your connector definitions. The rest of your pipeline logic—including transformations and aggregations—remains identical. This design allows rapid prototyping on static data before deployment to production streaming environments.

Why can't I use time-windowed aggregations in batch mode?

Time-related operators like as_of and windowed aggregations depend on an ever-increasing logical time that only exists in streaming mode. In batch mode, the engine processes a finite snapshot once, so there is no concept of progressing time or incremental updates to window boundaries.

What debug helper should I use for streaming outputs?

For streaming pipelines, use pw.debug.compute_and_print_update_stream() instead of pw.debug.compute_and_print(). The streaming helper outputs change-feeds with time and diff columns showing insertions and deletions, while the batch helper prints a static snapshot of the final table state.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →