How to Integrate Pathway with LangChain or LlamaIndex as a Retriever Backend
To integrate Pathway with LangChain or LlamaIndex as a retriever backend, deploy the Pathway DocumentStore server from the pathwaycom/llm-app repository, then connect via PathwayVectorClient for LangChain or PathwayVectorStore for LlamaIndex to enable real-time, continuously updated retrieval.
The pathwaycom/llm-app repository provides a production-ready implementation for using Pathway as the retrieval layer in your LLM applications. This integration allows both LangChain and LlamaIndex pipelines to leverage Pathway’s real-time indexing capabilities, ensuring that your retriever always has access to the latest data without manual re-indexing.
Architecture Overview
Pathway serves as a standalone document-store service that runs independently of your LangChain or LlamaIndex application. The architecture follows a client-server pattern:
- Pathway DocumentStore Server: Continuously ingests data from sources (files, databases, streams), computes embeddings, and maintains a vector index using the
DocumentStoreclass frompathway.xpacks.llm.document_store. - REST API: Exposes HTTP endpoints (
/v2/searchfor retrieval,/v2/upsertfor updates) viaDocumentStoreServerintemplates/document_indexing/app.py. - Client Libraries: LangChain’s
PathwayVectorClientand LlamaIndex’sPathwayVectorStoretranslate framework-specific retrieval calls into HTTP requests to the Pathway server.
This decoupling allows the Pathway backend to handle live data updates while LangChain or LlamaIndex focus on orchestration and generation.
Step 1: Deploy the Pathway DocumentStore Server
Before integrating with any framework, you must run the Pathway DocumentStore server. The primary implementation resides in templates/document_indexing/app.py.
Server Implementation Details
The server uses Pydantic-based configuration and wraps a DocumentStore instance with a DocumentStoreServer:
# templates/document_indexing/app.py
import logging
from warnings import warn
import pathway as pw
from dotenv import load_dotenv
from pathway.xpacks.llm.document_store import DocumentStore
from pathway.xpacks.llm.servers import DocumentStoreServer
from pydantic import BaseModel, ConfigDict, InstanceOf
# License key (demo)
pw.set_license_key("demo-license-key-with-telemetry")
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(name)s %(levelname)s %(message)s")
load_dotenv()
class App(BaseModel):
document_store: InstanceOf[DocumentStore]
host: str = "0.0.0.0"
port: int = 8000
# … persistence settings omitted for brevity …
def run(self) -> None:
DocumentStoreServer(self.host, self.port, self.document_store)
# … persistence config omitted …
pw.run(persistence_config=None, monitoring_level=pw.MonitoringLevel.NONE)
if __name__ == "__main__":
with open("app.yaml") as f:
config = pw.load_yaml(f)
App(**config).run()
Source: [templates/document_indexing/app.py](https://github.com/pathwaycom/llm-app/blob/main/templates/document_indexing/app.py)
Deployment Options
You can run the server locally or via Docker:
# Build and run with Docker
docker build -t pathway-docstore -f templates/document_indexing/Dockerfile .
docker run -p 8000:8000 pathway-docstore
Once running, the server exposes the REST API at http://localhost:8000 (or your configured host/port).
Step 2: Integrate Pathway with LangChain as a Retriever Backend
LangChain treats Pathway as a vector store through the PathwayVectorClient class available in langchain_community.vectorstores.
LangChain Integration Code
# LangChain example – Pathway as a vector store
from langchain_community.vectorstores import PathwayVectorClient
# Connect to the running Pathway server
vectorstore_client = PathwayVectorClient(host="localhost", port=8000)
# Obtain a LangChain retriever
retriever = vectorstore_client.as_retriever(search_kwargs={"k": 5})
# Use it in a LangChain chain or agent
docs = retriever.invoke({"question": "What is the contract start date?"})
print(docs)
Source: cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb
Key Integration Points
| Component | Implementation |
|---|---|
| Import | from langchain_community.vectorstores import PathwayVectorClient |
| Connection | Instantiate with host and port of the Pathway server. |
| Retriever | Call .as_retriever() to get a LangChain BaseRetriever compatible with LCEL chains. |
| Search parameters | Pass search_kwargs={"k": N} to control the number of retrieved documents. |
Step 3: Integrate Pathway with LlamaIndex as a Retriever Backend
LlamaIndex integrates with Pathway through the PathwayVectorStore class, which acts as a vector store wrapper around the Pathway HTTP API.
LlamaIndex Integration Code
# LlamaIndex example – Pathway as a retriever
from llama_index.vector_stores import PathwayVectorStore
from llama_index import SimpleDirectoryReader, VectorStoreIndex
# 1️⃣ Create a Pathway vector store wrapper
pathway_store = PathwayVectorStore(host="localhost", port=8000)
# 2️⃣ Build a LlamaIndex index that uses the Pathway store
documents = SimpleDirectoryReader("data/").load_data()
index = VectorStoreIndex.from_vector_store(pathway_store, documents=documents)
# 3️⃣ Query via the LlamaIndex retriever
retriever = index.as_retriever(similarity_top_k=5)
response = retriever.retrieve("What are the contract terms?")
print(response)
Reference: [templates/private_rag/README.md](https://github.com/pathwaycom/llm-app/blob/main/templates/private_rag/README.md)
Key Integration Points
| Component | Implementation |
|---|---|
| Vector store wrapper | PathwayVectorStore(host, port) from llama_index.vector_stores. |
| Index construction | Use VectorStoreIndex.from_vector_store(pathway_store, documents=...). |
| Retriever | Call index.as_retriever(similarity_top_k=N) to get a LlamaIndex BaseRetriever. |
| Search results | retrieve(query) returns NodeWithScore objects containing document text and metadata. |
How the Integration Works Under the Hood
When you use Pathway as a retriever backend for LangChain or LlamaIndex, the following architecture enables real-time, continuous retrieval:
-
Continuous Ingestion and Indexing: The
DocumentStoreclass inpathway.xpacks.llm.document_storecontinuously ingests data from configured sources, computes embeddings using your specified model, and maintains a high-performance vector index (usingusearchor persistent backends). -
REST API Exposure: The
DocumentStoreServerclass (instantiated intemplates/document_indexing/app.py) exposes two critical HTTP endpoints:/v2/upsert– For adding or updating vectors in the index/v2/search– For vector similarity search (top-k retrieval)
-
Client Translation: Both
PathwayVectorClient(LangChain) andPathwayVectorStore(LlamaIndex) act as HTTP clients. They translate framework-specific method calls (likeas_retriever().invoke()orretrieve()) into POST requests to the Pathway server's/v2/searchendpoint. -
Real-Time Synchronization: Because the Pathway server runs as a standalone service with its own data connectors, any new files added to connected sources (S3, local folders, databases) are immediately indexed and become searchable through the LangChain or LlamaIndex retriever without restarting your application.
The retriever_factory pattern shown in templates/slides_ai_search/app.py demonstrates how to expose custom retriever logic that can be injected into downstream applications, further extending this architecture.
Summary
- Pathway acts as a standalone retrieval service that continuously indexes data and exposes a REST API for vector search via
DocumentStoreServerintemplates/document_indexing/app.py. - LangChain integration uses
PathwayVectorClientfromlangchain_community.vectorstoresto connect to the running server and create a retriever with.as_retriever(). - LlamaIndex integration uses
PathwayVectorStorefromllama_index.vector_storesas a vector store wrapper, building an index withVectorStoreIndex.from_vector_store()and retrieving withindex.as_retriever(). - Real-time capabilities are inherent because Pathway continuously updates its index as source data changes, making new documents immediately available to both LangChain and LlamaIndex retrievers without manual refreshes.
Frequently Asked Questions
What is the difference between PathwayVectorClient and PathwayVectorStore?
PathwayVectorClient is the LangChain-specific implementation found in langchain_community.vectorstores that conforms to LangChain’s VectorStore interface, allowing you to call .as_retriever() and use it in LCEL chains. PathwayVectorStore is the LlamaIndex implementation from llama_index.vector_stores that conforms to LlamaIndex’s BasePydanticVectorStore interface, enabling integration with VectorStoreIndex. Both act as HTTP clients communicating with the same Pathway DocumentStore server.
Does Pathway support real-time updates when used as a retriever backend?
Yes. Because Pathway runs as a continuous streaming engine, the DocumentStore in templates/document_indexing/app.py automatically indexes new documents as they appear in connected data sources (local files, S3, databases, etc.). When using PathwayVectorClient or PathwayVectorStore, your LangChain or LlamaIndex retriever immediately sees these updates without requiring manual index refreshes or application restarts.
Which file contains the server implementation for the DocumentStore?
The primary server implementation is located at templates/document_indexing/app.py. This file defines the App class that instantiates DocumentStore from pathway.xpacks.llm.document_store and wraps it with DocumentStoreServer from pathway.xpacks.llm.servers to expose the REST API endpoints required by LangChain and LlamaIndex clients.
Can I use Pathway with both LangChain and LlamaIndex simultaneously?
Yes. Because Pathway operates as a standalone service via DocumentStoreServer, you can run a single Pathway instance and connect multiple clients to it. For example, you can have a LangChain application using PathwayVectorClient and a separate LlamaIndex application using PathwayVectorStore both pointing to the same host:port endpoint, sharing the same real-time indexed data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →