# How DocumentStore Handles Real-Time Index Updates When Source Documents Change

> Learn how DocumentStore streams changes from source documents, updating indexes in real-time without interrupting queries. Discover efficient text and vector index management.

- Repository: [Pathway/llm-app](https://github.com/pathwaycom/llm-app)
- Tags: internals
- Published: 2026-03-07

---

**Pathway's DocumentStore implements a streaming pipeline that polls connected data sources at configurable intervals, detecting file additions, modifications, and deletions to incrementally update only the changed portions of the vector and text indexes while maintaining query availability.**

The `DocumentStore` class in the [pathwaycom/llm-app](https://github.com/pathwaycom/llm-app) repository enables production-grade document retrieval without manual re-ingestion workflows. By treating document indexing as a continuous stream rather than a batch process, **real-time index updates** ensure that AI applications always retrieve the latest content from sources like Google Drive, SharePoint, S3, Kafka, or local filesystems.

## The Streaming Pipeline Architecture

Unlike traditional document stores that require manual refresh jobs, Pathway's DocumentStore operates as a persistent streaming pipeline. The architecture relies on **source connectors** that implement polling mechanisms to monitor external systems continuously.

When you instantiate a `DocumentStore`, you configure a `poll_interval` parameter that determines how frequently the system checks for changes. For example, in [`templates/slides_ai_search/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/slides_ai_search/app.py), the store is configured to poll every 5 seconds:

```python
from pathway.xpacks.llm.document_store import SlidesDocumentStore

doc_store = SlidesDocumentStore(
    source="path_to_your_folder",
    poll_interval=5,        # seconds between polls

    ...
)

```

This polling mechanism applies across all supported connectors, including cloud storage providers and message queues, ensuring that **real-time index updates** occur regardless of where the source documents reside.

## How DocumentStore Detects and Processes Changes

The real-time update flow involves three distinct phases: change detection, incremental processing, and index synchronization.

### Source Polling and Change Detection

The connector layer continuously monitors the configured source using the specified polling interval. When the poller detects a file addition, modification, or deletion, it emits a change event into the Pathway dataflow graph. According to the `question_answering_rag` template documentation, this allows the pipeline to "poll these sources at configured intervals, so when new documents appear or existing ones change, they are automatically parsed and re-indexed in real-time."

### Incremental Parsing and Chunking

Once a change event enters the pipeline, `DocumentStore` processes only the affected documents rather than rebuilding the entire index. The system re-parses the modified files, splits them into chunks, and generates fresh embeddings. As noted in the `slides_ai_search` README, "Pathway polls the changes with low latency... the corresponding change is reflected in real-time, and search results are updated accordingly."

### Index Refresh with KNNIndex

The processed chunks flow into the underlying index implementation, typically `KNNIndex` for vector search or a hybrid index combining vector and BM25 components. The index updates its internal data structures in place, inserting new vectors and removing obsolete ones. Because the index remains in memory throughout the process, queries continue to execute against the stable portions of the dataset while updates occur, ensuring zero-downtime **real-time index updates**.

## Step-by-Step Real-Time Update Flow

The complete lifecycle of a document change flowing through the system follows these stages:

| Step | Action | Implementation Detail |
|------|--------|----------------------|
| **1. Source Polling** | The connector queries the external source at the configured `poll_interval`. | Implemented in [`templates/slides_ai_search/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/slides_ai_search/app.py) via the `poll_interval` parameter. |
| **2. Change Detection** | Added, modified, or deleted files trigger events in the Pathway dataflow. | The `DocumentStore` receives events from the connector layer. |
| **3. Incremental Processing** | Only changed documents are parsed, chunked, and embedded. | Processed through Pathway's built-in parsers within the `DocumentStore` pipeline. |
| **4. Index Refresh** | The `KNNIndex` or hybrid index updates in place with new vectors. | Obsolete chunks are removed and new ones inserted without rebuilding the entire index. |
| **5. Query Availability** | Updated content is immediately available for retrieval. | As documented in [`templates/question_answering_rag/README.md`](https://github.com/pathwaycom/llm-app/blob/main/templates/question_answering_rag/README.md), changes reflect instantly in search results. |

## Summary

Pathway's `DocumentStore` eliminates the need for batch re-indexing jobs by implementing a continuous streaming architecture that delivers **real-time index updates** through the following mechanisms:

- **Continuous Polling**: Configurable `poll_interval` parameters enable low-latency detection of source changes across filesystems, cloud storage, and message queues.
- **Incremental Processing**: Only modified documents undergo parsing, chunking, and embedding, preserving computational resources.
- **In-Place Index Updates**: The underlying `KNNIndex` updates its data structures incrementally, maintaining query availability throughout the process.
- **Zero-Downtime Operation**: The index remains in memory and accessible during updates, ensuring that AI applications always retrieve the latest document versions.

## Frequently Asked Questions

### What data sources support real-time index updates in DocumentStore?

DocumentStore supports real-time updates across diverse sources including local filesystems, Google Drive, Microsoft SharePoint, Amazon S3, and Kafka streams. Each connector implements the same polling mechanism, allowing you to configure the `poll_interval` parameter to control update latency regardless of the source location.

### How does DocumentStore handle deleted documents?

When the source connector detects a file deletion during its polling cycle, DocumentStore treats this as a removal event in the dataflow graph. The system identifies all chunks associated with the deleted document and removes them from the `KNNIndex` or hybrid index immediately, ensuring that subsequent queries do not return stale or orphaned content.

### What is the performance impact of real-time indexing on query latency?

Real-time indexing occurs incrementally and asynchronously with respect to query processing. Because DocumentStore updates the index in place while keeping it resident in memory, query latency remains stable even during active indexing. Only the changed document chunks are processed, minimizing CPU and memory overhead compared to full re-indexing operations.

### Can I adjust the polling frequency for different document sources?

Yes, the `poll_interval` parameter is configurable per DocumentStore instance, allowing you to optimize for specific source characteristics. For high-velocity sources like Kafka, you might use sub-second intervals, while for stable network filesystems, intervals of 30-60 seconds may suffice. This configuration is demonstrated in [`templates/slides_ai_search/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/slides_ai_search/app.py) where `poll_interval=5` sets a 5-second polling cycle.