Key Differences Between Batch and Streaming Processing Modes in Pathway: A Complete Guide
Pathway supports two mutually exclusive execution modes—static (batch) mode and streaming mode—each with distinct connector constraints, computation semantics, and runtime behavior that determine how data is ingested, processed, and output.
Understanding the key differences between batch and streaming processing modes in Pathway is essential for building efficient data pipelines. The pathwaycom/pathway repository implements these two mutually exclusive execution models at the connector level, allowing developers to switch between one-off analytics and continuous real-time processing by changing a single parameter.
Execution Models and Data Ingestion
Batch (Static) Mode
In batch mode, all input data is read once and processed as a single, finite batch. According to the batch processing documentation, the engine executes the dataflow exactly once and materializes results as a complete static table. Connector calls block until the entire source is consumed, making this mode ideal for one-off analytics, testing, and daily batch jobs.
Streaming Mode
Streaming mode maintains continuously polling connectors that stay open indefinitely, processing new records as they arrive. The engine performs incremental, stateful computations where each update triggers downstream recomputation. As implemented in the Pathway engine, this mode runs until manually terminated and automatically updates tables as new data appears, supporting real-time dashboards and event-driven pipelines.
Connector Compatibility and Constraints
Pathway strictly separates connector types between modes. The connector mode is defined as Literal["streaming", "static"] throughout the I/O packages, such as in [python/pathway/io/kafka/__init__.py](https://github.com/pathwaycom/pathway/blob/main/python/pathway/io/kafka/__init__.py).
- Static mode connectors are restricted to batch processing and defined under the Static column in the connectors table.
- Streaming mode connectors continuously poll for new data from sources like Kafka or watched directories.
- Mixing is prohibited: The documentation in [
30.connectors-in-pathway.md](https://github.com/pathwaycom/pathway/blob/main/docs/2.developers/4.user-guide/20.connect/30.connectors-in-pathway.md) explicitly states that static and streaming connectors cannot be combined in the same pipeline.
Computation Semantics and Output Formats
The computation models differ fundamentally between the two modes.
Batch Mode:
- Executes dataflow once and produces a static snapshot of the final state.
- Time-related operators (e.g.,
as_of, windowed aggregations) are not applicable because logical time does not progress. - Results materialize as complete tables suitable for debug helpers like
pw.debug.compute_and_print.
Streaming Mode:
- Performs incremental, stateful computations with automatic recomputation of downstream operators.
- Outputs are change-feeds containing
time(logical timestamp) anddiff(+1for insert,-1for delete) columns. - Tables update automatically as new data arrives; use
pw.debug.compute_and_print_update_streamfor streaming outputs.
Practical Implementation: Code Examples
The following examples demonstrate how to implement both modes using the CSV connector in [python/pathway/io/csv/__init__.py](https://github.com/pathwaycom/pathway/blob/main/python/pathway/io/csv/__init__.py).
Batch Processing with Static Mode
import pathway as pw
# Static (batch) mode – the file is read once and the pipeline finishes.
t = pw.io.csv.read(
"data/sales.csv",
schema=pw.Schema(
order_id=int,
amount=float,
country=str,
),
mode="static", # <-- batch mode
)
# Simple aggregation (batch)
total = pw.select(t, total_amount=pw.sum(t.amount))
pw.debug.compute_and_print(total) # prints a snapshot once
Streaming Processing with Directory Watching
import pathway as pw
# Streaming mode – the connector watches the directory and processes new files as they appear.
t = pw.io.csv.read(
"data/streaming/",
schema=pw.Schema(
order_id=int,
amount=float,
country=str,
),
mode="streaming", # <-- streaming mode
)
# Incremental aggregation; each new file updates the result.
total = pw.select(t, total_amount=pw.sum(t.amount))
# Use the streaming‑aware compute helper.
pw.debug.compute_and_print_update_stream(total)
Switching Between Modes
import pathway as pw
# Define pipeline once
def pipeline(source_mode):
t = pw.io.csv.read(
"data/input/",
schema=pw.Schema(id=int, value=float),
mode=source_mode, # pass "static" or "streaming"
)
agg = pw.select(t, sum_val=pw.sum(t.value))
return agg
# Batch execution
batch_result = pipeline("static")
pw.debug.compute_and_print(batch_result)
# Streaming execution (just change the mode string)
stream_result = pipeline("streaming")
pw.debug.compute_and_print_update_stream(stream_result)
Persistence and Back-filling Behavior
Both modes support persistence for storing state, but the back-filling behavior differs:
- Batch mode: Persistence stores state between runs, enabling back-filling of only new data on subsequent executions when reprocessing the entire dataset.
- Streaming mode: Persistence works identically, but the engine naturally processes only new increments as they appear, making back-filling inherently incremental without reprocessing historical data.
Summary
- Pathway enforces mutually exclusive batch and streaming modes at the connector level with no mixing allowed.
- Static mode processes finite datasets once and produces static snapshots, while streaming mode maintains continuous, stateful computations with change-feed outputs.
- Connectors declare their mode via
Literal["streaming", "static"]type definitions as seen in the Kafka and CSV I/O packages. - Switching execution modes requires only changing the
modeparameter from"static"to"streaming"—pipeline logic remains unchanged. - Debug helpers in [
python/pathway/debug/__init__.py](https://github.com/pathwaycom/pathway/blob/main/python/pathway/debug/__init__.py) differ between modes:compute_and_printfor batch snapshots,compute_and_print_update_streamfor streaming change-feeds.
Frequently Asked Questions
Can I mix batch and streaming connectors in the same Pathway pipeline?
No. According to the connector documentation, Pathway strictly prohibits mixing static and streaming connectors. Each pipeline must use exclusively one mode type, enforced by the Literal["streaming", "static"] type system in the I/O package definitions.
How do I switch an existing batch pipeline to streaming?
You only need to change the mode parameter from "static" to "streaming" in your connector definitions. The rest of your pipeline logic—including transformations and aggregations—remains identical. This design allows rapid prototyping on static data before deployment to production streaming environments.
Why can't I use time-windowed aggregations in batch mode?
Time-related operators like as_of and windowed aggregations depend on an ever-increasing logical time that only exists in streaming mode. In batch mode, the engine processes a finite snapshot once, so there is no concept of progressing time or incremental updates to window boundaries.
What debug helper should I use for streaming outputs?
For streaming pipelines, use pw.debug.compute_and_print_update_stream() instead of pw.debug.compute_and_print(). The streaming helper outputs change-feeds with time and diff columns showing insertions and deletions, while the batch helper prints a static snapshot of the final table state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →