Debugging and Troubleshooting Pathway Pipelines: Best Practices and Tools

Use pw.debug.compute_and_print with tagged probes in static mode, ensure pw.run() explicitly starts streaming pipelines, and verify column schemas to isolate failures in Pathway's Rust-based dataflow engine.

Pathway is a high-performance dataflow framework built on a Rust engine and exposed through Python bindings in the pathwaycom/pathway repository. Debugging and troubleshooting Pathway pipelines requires understanding how the Python API translates into engine operations across the graph and runtime layers. Leveraging the built-in debugging utilities allows you to trace data transformations from source connectors through complex windowed aggregations.

Understanding the Pathway Debugging Architecture

Pathway’s debugging capabilities span three distinct layers that convert Python calls into observable runtime behavior.

The Python API Layer provides user-facing functions such as pw.debug.compute_and_print, pw.debug.table_from_markdown, and pw.debug.table. These wrap underlying engine calls defined in src/python_api.rs, where the debug_table method (lines 3545–3557) converts Python invocations into engine requests.

The Engine Graph Layer maintains the dataflow graph representation, tracking tables, operators, and debugging probes. In src/engine/graph.rs (lines 26–31), the debug_table trait method on the Graph struct receives a tag identifier, table handle, and column list to establish inspection points.

The Dataflow Runtime Layer executes the graph and materializes debug output. When the runtime processes a debug_table request, it iterates over table values, extracts requested columns via ColumnPath, and prints formatted rows. This implementation resides in src/engine/dataflow.rs (lines 2955–3016), producing output in the format [worker_id][tag] @timestamp diff id=key ....

Debugging Static vs. Streaming Pipelines

Pathway operates in two distinct execution modes, and applying the wrong debugging pattern to either mode is the most common source of confusion.

Static Mode Development

For batch processing and unit testing, use pw.debug.table_from_markdown or pw.debug.table_from_pandas to inject deterministic fixtures. This allows you to isolate logical errors without external dependencies.

Always invoke table.debug.compute_and_print(tag="stage_name") to force computation and emit rows to stdout. According to the best-practices documentation, this function executes only in static mode; in streaming mode, it silently produces no output.

import pathway as pw

# Create deterministic test data

sample = pw.debug.table_from_markdown("""
| id | value |
|----|-------|
| 1  | 10    |
| 2  | 20    |
""")

# Apply transformation

result = sample.with_columns(doubled=pw.this.value * 2)

# Inspect intermediate state with a tagged probe

result.debug.compute_and_print(tag="after_double")

Streaming Mode Execution

In streaming pipelines, the computation runs continuously until terminated. You must explicitly start the engine with pw.run(), or the pipeline will appear to do nothing. Add debug probes using the same compute_and_print method, but expect output only after the runtime begins ticking.

@pw.udf(deterministic=True)
def normalize(x: float) -> float:
    return x / 100.0

source = pw.io.read_csv("events.csv", mode="streaming")
windowed = source.window_by(
    pw.temporal.tumbling_window(duration=60)
).reduce(pw.this.value.apply(normalize).sum())

# Attach debug probe before starting

windowed.debug.compute_and_print(tag="windowed_sum")

# Critical: start the streaming engine

pw.run()

Resolving Common Pipeline Failures

Silent Termination and "Nothing Happens" Symptoms

If your pipeline executes but produces no output, verify three specific conditions. First, confirm you called pw.run() for streaming jobs or compute_and_print() for static jobs, as the engine defers all computation until explicitly started. Second, check that your connector is actually emitting data by inspecting connector.status(). Third, verify that source data conforms to the expected table schema, as the engine silently drops rows with mismatched column names or types.

Inspect the schema programmatically using print(table.schema) to reconcile column definitions against source data. The connector attachment logic in src/engine/graph.rs (lines 47–53) shows how input sources bind to the graph, and schema validation errors often surface at this boundary.

Low-Level Debug API

For library developers or complex debugging scenarios, use the low-level pw.debug.table API to specify exact column paths:

pw.debug.table("low_level_probe", my_table, [("col1", pw.ColumnPath("col1"))])

This invokes the Rust implementation directly, bypassing the high-level wrapper, and produces detailed timestamped diffs showing exactly when each row enters or leaves the operator.

Best Practices for Debuggable Code

Following these patterns minimizes debugging time and surface area for errors:

  • Prefer built-in transformations over Python UDFs. Native Rust operators in src/engine/dataflow.rs run faster and include automatic instrumentation that aids tracing.

  • Annotate UDF signatures with @pw.udf and mark deterministic functions with deterministic=True. This provides the compiler with result types up-front and enables result caching avoidance, as detailed in the best-practices guide.

  • Favor stateless transformations to minimize per-key state. Stateless operations reduce hidden complexity and make dataflow inspection via debug probes more intuitive.

  • Leverage multiprocessing for heavy UDFs using pathway spawn -n N. This isolates performance bottlenecks and crashes to specific worker processes, preventing total pipeline failure.

  • Enable OpenTelemetry exporters for distributed deployments. This provides telemetry across worker nodes, complementing the console-based debug_table output with centralized logging.

Summary

  • Pathway’s debugging workflow flows from the Python API through the Engine Graph to the Dataflow Runtime, with probes materializing in src/engine/dataflow.rs.

  • Use pw.debug.compute_and_print only in static mode; streaming mode requires pw.run() to activate the engine.

  • Validate schema alignment using table.schema when connectors silently drop rows.

  • Tag debug probes (e.g., tag="after_filter") to pinpoint exactly which graph stage contains the data you expect.

  • Annotate UDFs and prefer built-in transformations to maximize performance and observability.

Frequently Asked Questions

Why does my Pathway pipeline produce no output?

The pipeline likely lacks an explicit start command or contains schema mismatches. In static mode, ensure you call compute_and_print() on the final table. In streaming mode, verify that pw.run() appears at the end of your script. Additionally, check that input connector schemas match the declared table structure, as non-conforming rows are dropped silently.

How do I inspect intermediate tables in a streaming pipeline?

Insert tagged debug probes using table.debug.compute_and_print(tag="stage_name") before calling pw.run(). The runtime prints each row as it passes through that operator, prefixed with your tag and timestamp. For deeper inspection, use pw.debug.table to specify exact columns from the underlying Rust engine.

What causes schema mismatch errors in Pathway connectors?

Schema mismatches occur when the source data structure (CSV columns, Kafka message fields) does not align with the Pathway table definition. The connector layer in src/engine/graph.rs attaches sources to the graph, but runtime type checking enstricts row admission. Use print(table.schema) to compare expected column names and types against your source data, and consult the schema documentation for strict type coercion rules.

When should I use pw.debug.table versus table.debug.compute_and_print?

Use table.debug.compute_and_print() for standard debugging in Python scripts, as it handles column enumeration automatically. Use pw.debug.table() when you need low-level control over which columns to inspect or when developing library code that interacts directly with the Rust engine’s debug_table trait in src/engine/graph.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →