# How to Configure Snapshot Intervals and State Checkpointing in Pathway for Durability

> Configure Pathway snapshot intervals and state checkpointing to ensure durability. Learn how to set snapshot intervals and access modes for reliable data persistence.

- Repository: [Pathway/pathway](https://github.com/pathwaycom/pathway)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Configure snapshot intervals and state checkpointing in Pathway by instantiating `pw.persistence.Config` with a `snapshot_interval_ms` value and a `snapshot_access` mode, then passing this configuration to `pw.run()` to enable automatic periodic persistence of table states to durable backends like filesystem or S3.**

Pathway's stream processing engine provides built-in durability through its persistence subsystem, which periodically captures point-in-time snapshots of table states. When you configure snapshot intervals and state checkpointing in Pathway, you enable automatic recovery from crashes without data loss. This guide demonstrates the implementation details found in the `pathwaycom/pathway` repository, covering the `pw.persistence.Config` API and its integration with various storage backends.

## Understanding Pathway's Persistence Knobs

Pathway exposes two primary configuration parameters that control durability behavior. These settings determine how frequently the system checkpoints state and what data is actually persisted.

### snapshot_interval_ms

The **`snapshot_interval_ms`** parameter defines the minimum time between automatic snapshot writes. When set to a non-zero value (e.g., `2000` for 2 seconds), the runtime spawns a background task that triggers snapshot writes to the configured backend at the specified interval. Setting this to `0` disables periodic snapshots, causing state to be written only when the pipeline finishes, as implemented in the `Config` constructor in [`python/pathway/persistence/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/persistence/__init__.py) (lines 149-162).

### snapshot_access

The **`snapshot_access`** parameter determines the persistence mode and recovery behavior. As defined in [`python/pathway/persistence/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/persistence/__init__.py) and handled via environment variables in [`python/pathway/internals/config.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/internals/config.py) (lines 42-49), four modes are available:

- **FULL**: Stores the complete table state (default, most durable)
- **OFFSETS_ONLY**: Stores only change offsets (smaller storage, slower replay)
- **RECORD**: Records new changes without persisting previous state
- **REPLAY**: Loads existing snapshots without storing new changes

## Basic Configuration with Filesystem Backend

The most common setup uses the filesystem backend with periodic snapshots. In [`python/pathway/persistence/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/persistence/__init__.py) (lines 149-162), the `Config` constructor accepts a `Backend` instance along with the interval and access mode parameters.

```python
import pathlib
import pathway as pw

# ----------------------------------------------------------------------

# 1. Define a schema and source

# ----------------------------------------------------------------------

class Input(pw.Schema):
    key: int = pw.column_definition(primary_key=True)
    value: str

# a streaming CSV source (replace with any connector you need)

source = pw.io.csv.read("data/input.csv", schema=Input, mode="streaming")

# ----------------------------------------------------------------------

# 2. Build a simple aggregation

# ----------------------------------------------------------------------

agg = source.groupby(pw.this.key).reduce(pw.this.value, count=pw.reducers.count())

# ----------------------------------------------------------------------

# 3. Persist the result with snapshots every 2 seconds

# ----------------------------------------------------------------------

storage_path = pathlib.Path("/tmp/pw_snapshot")
persistence_cfg = pw.persistence.Config(
    pw.persistence.Backend.filesystem(storage_path),
    snapshot_interval_ms=2000,                 # 2 seconds between checkpoints

    snapshot_access=pw.api.SnapshotAccess.FULL  # store the full table state

)

# ----------------------------------------------------------------------

# 4. Run the pipeline

# ----------------------------------------------------------------------

pw.run(persistence_config=persistence_cfg)

```

Key lines: `snapshot_interval_ms=2000` and `snapshot_access=...` are highlighted in the `Config` constructor.

## Cloud Persistence with S3

For production durability, configure the S3 backend defined in [`python/pathway/persistence/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/persistence/__init__.py) (lines 48-68). This stores snapshots in cloud object storage with the same interval-based checkpointing mechanism.

```python
import pathway as pw
from pathway.internals._io_helpers import AwsS3Settings

# S3 bucket configuration (replace with your bucket details)

s3_settings = AwsS3Settings(
    region="us-east-1",
    access_key_id="YOUR_KEY",
    secret_access_key="YOUR_SECRET"
)

backend = pw.persistence.Backend.s3(
    root_path="my-pipeline-snapshots",
    bucket_settings=s3_settings,
)

persistence_cfg = pw.persistence.Config(
    backend,
    snapshot_interval_ms=5000,                # 5 seconds

    snapshot_access=pw.api.SnapshotAccess.FULL
)

# Run the same pipeline as before, passing the S3-backed config

pw.run(persistence_config=persistence_cfg)

```

Key lines: `Backend.s3` is defined in [`persistence/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/persistence/__init__.py).

## Recovery and Replay Modes

To recover from a previous state, use the **REPLAY** access mode. This loads the latest snapshot from the backend without ingesting new data, useful for batch analysis or debugging historical states.

```python
import pathlib
import pathway as pw

# The same storage path that was used for persisting

storage_path = pathlib.Path("/tmp/pw_snapshot")

# Re-run the pipeline, but only replay the snapshot (no new data ingestion)

pw.run(
    persistence_config=pw.persistence.Config(
        pw.persistence.Backend.filesystem(storage_path),
        # REPLAY mode will load the last snapshot and stop after it

        snapshot_access=pw.api.SnapshotAccess.REPLAY,
        continue_after_replay=False
    )
)

```

The `snapshot_access=REPLAY` flag tells the engine to load the latest persisted snapshot and then exit.

## How State Checkpointing Works Under the Hood

The persistence architecture in Pathway follows a multi-layered design that translates high-level configuration into runtime operations.

### Backend Abstraction

The **`pw.persistence.Backend`** class (lines 26-45 in [`python/pathway/persistence/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/persistence/__init__.py)) abstracts storage locations, converting configuration into `api.DataStorage` objects that the runtime writes to.

### Configuration Translation

When `pw.persistence.Config` is instantiated, its **`engine_config`** property (lines 209-217) constructs an `api.PersistenceConfig` that encapsulates the backend, interval, and access settings for the engine.

### Runtime Snapshot Loop

During `pw.run(persistence_config=...)`, a background task monitors elapsed time since the last snapshot. If the duration exceeds `snapshot_interval_ms`, the runtime initiates a snapshot write to the backend. The integration test at [`python/pathway/tests/test_persistence.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/tests/test_persistence.py) (lines 54-55 and 86) demonstrates this behavior with a 1000ms interval and sleep-based verification.

## Summary

- Configure snapshot intervals using `snapshot_interval_ms` in `pw.persistence.Config` to control checkpoint frequency; set to `0` to disable periodic snapshots
- Select durability levels via `snapshot_access` (**FULL** for complete state, **OFFSETS_ONLY** for minimal storage, **REPLAY** for recovery-only mode)
- Instantiate backends through `pw.persistence.Backend.filesystem()` or `Backend.s3()` defined in [`python/pathway/persistence/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/persistence/__init__.py)
- Pass the configuration to `pw.run(persistence_config=...)` to enable automatic checkpointing
- Recover state using `SnapshotAccess.REPLAY` mode to load previous snapshots without new data ingestion

## Frequently Asked Questions

### What happens if I set snapshot_interval_ms to 0?

Setting `snapshot_interval_ms=0` disables the periodic snapshot background task. The system will only write a snapshot when the pipeline explicitly finishes or when triggered by other mechanisms, meaning you lose the ability to recover from mid-run crashes and must replay from the beginning or last manual checkpoint.

### What is the difference between FULL and OFFSETS_ONLY snapshot access?

**FULL** mode persists the complete table state at each checkpoint in [`python/pathway/persistence/__init__.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/persistence/__init__.py), enabling fast recovery by loading the latest snapshot directly from the backend. **OFFSETS_ONLY** stores minimal metadata about changes, reducing storage costs but requiring the system to replay all events from the offset during recovery, which increases startup time. The [`python/pathway/tests/test_io.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/tests/test_io.py) file (lines 41-44) demonstrates usage differences between these modes.

### Can I use environment variables to configure snapshot access?

Yes. According to [`python/pathway/internals/config.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/internals/config.py) (lines 42-49), you can set the **`PATHWAY_SNAPSHOT_ACCESS`** environment variable to configure the default snapshot access mode without modifying code, allowing runtime configuration changes across different deployment environments.

### How do I verify that snapshots are being written correctly?

The test suite in [`python/pathway/tests/test_persistence.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/tests/test_persistence.py) demonstrates verification techniques: configure a short interval (e.g., 1000ms as shown on lines 54-55), run the pipeline with a sleep duration exceeding the interval (line 86), then inspect the storage backend for snapshot metadata files. For a real-world example, see [`python/pathway/integration_tests/wordcount/pw_wordcount.py`](https://github.com/pathwaycom/pathway/blob/main/python/pathway/integration_tests/wordcount/pw_wordcount.py) (lines 29-55), which sets `snapshot_interval_ms=5000` for a production word-count pipeline.