How DocumentStore Handles Real-Time Index Updates When Source Documents Change
Pathway's DocumentStore implements a streaming pipeline that polls connected data sources at configurable intervals, detecting file additions, modifications, and deletions to incrementally update only the changed portions of the vector and text indexes while maintaining query availability.
The DocumentStore class in the pathwaycom/llm-app repository enables production-grade document retrieval without manual re-ingestion workflows. By treating document indexing as a continuous stream rather than a batch process, real-time index updates ensure that AI applications always retrieve the latest content from sources like Google Drive, SharePoint, S3, Kafka, or local filesystems.
The Streaming Pipeline Architecture
Unlike traditional document stores that require manual refresh jobs, Pathway's DocumentStore operates as a persistent streaming pipeline. The architecture relies on source connectors that implement polling mechanisms to monitor external systems continuously.
When you instantiate a DocumentStore, you configure a poll_interval parameter that determines how frequently the system checks for changes. For example, in templates/slides_ai_search/app.py, the store is configured to poll every 5 seconds:
from pathway.xpacks.llm.document_store import SlidesDocumentStore
doc_store = SlidesDocumentStore(
source="path_to_your_folder",
poll_interval=5, # seconds between polls
...
)
This polling mechanism applies across all supported connectors, including cloud storage providers and message queues, ensuring that real-time index updates occur regardless of where the source documents reside.
How DocumentStore Detects and Processes Changes
The real-time update flow involves three distinct phases: change detection, incremental processing, and index synchronization.
Source Polling and Change Detection
The connector layer continuously monitors the configured source using the specified polling interval. When the poller detects a file addition, modification, or deletion, it emits a change event into the Pathway dataflow graph. According to the question_answering_rag template documentation, this allows the pipeline to "poll these sources at configured intervals, so when new documents appear or existing ones change, they are automatically parsed and re-indexed in real-time."
Incremental Parsing and Chunking
Once a change event enters the pipeline, DocumentStore processes only the affected documents rather than rebuilding the entire index. The system re-parses the modified files, splits them into chunks, and generates fresh embeddings. As noted in the slides_ai_search README, "Pathway polls the changes with low latency... the corresponding change is reflected in real-time, and search results are updated accordingly."
Index Refresh with KNNIndex
The processed chunks flow into the underlying index implementation, typically KNNIndex for vector search or a hybrid index combining vector and BM25 components. The index updates its internal data structures in place, inserting new vectors and removing obsolete ones. Because the index remains in memory throughout the process, queries continue to execute against the stable portions of the dataset while updates occur, ensuring zero-downtime real-time index updates.
Step-by-Step Real-Time Update Flow
The complete lifecycle of a document change flowing through the system follows these stages:
| Step | Action | Implementation Detail |
|---|---|---|
| 1. Source Polling | The connector queries the external source at the configured poll_interval. |
Implemented in templates/slides_ai_search/app.py via the poll_interval parameter. |
| 2. Change Detection | Added, modified, or deleted files trigger events in the Pathway dataflow. | The DocumentStore receives events from the connector layer. |
| 3. Incremental Processing | Only changed documents are parsed, chunked, and embedded. | Processed through Pathway's built-in parsers within the DocumentStore pipeline. |
| 4. Index Refresh | The KNNIndex or hybrid index updates in place with new vectors. |
Obsolete chunks are removed and new ones inserted without rebuilding the entire index. |
| 5. Query Availability | Updated content is immediately available for retrieval. | As documented in templates/question_answering_rag/README.md, changes reflect instantly in search results. |
Summary
Pathway's DocumentStore eliminates the need for batch re-indexing jobs by implementing a continuous streaming architecture that delivers real-time index updates through the following mechanisms:
- Continuous Polling: Configurable
poll_intervalparameters enable low-latency detection of source changes across filesystems, cloud storage, and message queues. - Incremental Processing: Only modified documents undergo parsing, chunking, and embedding, preserving computational resources.
- In-Place Index Updates: The underlying
KNNIndexupdates its data structures incrementally, maintaining query availability throughout the process. - Zero-Downtime Operation: The index remains in memory and accessible during updates, ensuring that AI applications always retrieve the latest document versions.
Frequently Asked Questions
What data sources support real-time index updates in DocumentStore?
DocumentStore supports real-time updates across diverse sources including local filesystems, Google Drive, Microsoft SharePoint, Amazon S3, and Kafka streams. Each connector implements the same polling mechanism, allowing you to configure the poll_interval parameter to control update latency regardless of the source location.
How does DocumentStore handle deleted documents?
When the source connector detects a file deletion during its polling cycle, DocumentStore treats this as a removal event in the dataflow graph. The system identifies all chunks associated with the deleted document and removes them from the KNNIndex or hybrid index immediately, ensuring that subsequent queries do not return stale or orphaned content.
What is the performance impact of real-time indexing on query latency?
Real-time indexing occurs incrementally and asynchronously with respect to query processing. Because DocumentStore updates the index in place while keeping it resident in memory, query latency remains stable even during active indexing. Only the changed document chunks are processed, minimizing CPU and memory overhead compared to full re-indexing operations.
Can I adjust the polling frequency for different document sources?
Yes, the poll_interval parameter is configurable per DocumentStore instance, allowing you to optimize for specific source characteristics. For high-velocity sources like Kafka, you might use sub-second intervals, while for stable network filesystems, intervals of 30-60 seconds may suffice. This configuration is demonstrated in templates/slides_ai_search/app.py where poll_interval=5 sets a 5-second polling cycle.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →