How to Integrate Pathway with LangChain or LlamaIndex as a Retriever Backend

To integrate Pathway with LangChain or LlamaIndex as a retriever backend, deploy the Pathway DocumentStore server from the pathwaycom/llm-app repository, then connect via PathwayVectorClient for LangChain or PathwayVectorStore for LlamaIndex to enable real-time, continuously updated retrieval.

The pathwaycom/llm-app repository provides a production-ready implementation for using Pathway as the retrieval layer in your LLM applications. This integration allows both LangChain and LlamaIndex pipelines to leverage Pathway’s real-time indexing capabilities, ensuring that your retriever always has access to the latest data without manual re-indexing.

Architecture Overview

Pathway serves as a standalone document-store service that runs independently of your LangChain or LlamaIndex application. The architecture follows a client-server pattern:

  1. Pathway DocumentStore Server: Continuously ingests data from sources (files, databases, streams), computes embeddings, and maintains a vector index using the DocumentStore class from pathway.xpacks.llm.document_store.
  2. REST API: Exposes HTTP endpoints (/v2/search for retrieval, /v2/upsert for updates) via DocumentStoreServer in templates/document_indexing/app.py.
  3. Client Libraries: LangChain’s PathwayVectorClient and LlamaIndex’s PathwayVectorStore translate framework-specific retrieval calls into HTTP requests to the Pathway server.

This decoupling allows the Pathway backend to handle live data updates while LangChain or LlamaIndex focus on orchestration and generation.

Step 1: Deploy the Pathway DocumentStore Server

Before integrating with any framework, you must run the Pathway DocumentStore server. The primary implementation resides in templates/document_indexing/app.py.

Server Implementation Details

The server uses Pydantic-based configuration and wraps a DocumentStore instance with a DocumentStoreServer:


# templates/document_indexing/app.py

import logging
from warnings import warn

import pathway as pw
from dotenv import load_dotenv
from pathway.xpacks.llm.document_store import DocumentStore
from pathway.xpacks.llm.servers import DocumentStoreServer
from pydantic import BaseModel, ConfigDict, InstanceOf

# License key (demo)

pw.set_license_key("demo-license-key-with-telemetry")

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(name)s %(levelname)s %(message)s")

load_dotenv()

class App(BaseModel):
    document_store: InstanceOf[DocumentStore]
    host: str = "0.0.0.0"
    port: int = 8000
    # … persistence settings omitted for brevity …

    def run(self) -> None:
        DocumentStoreServer(self.host, self.port, self.document_store)
        # … persistence config omitted …

        pw.run(persistence_config=None, monitoring_level=pw.MonitoringLevel.NONE)

if __name__ == "__main__":
    with open("app.yaml") as f:
        config = pw.load_yaml(f)
    App(**config).run()

Source: [templates/document_indexing/app.py](https://github.com/pathwaycom/llm-app/blob/main/templates/document_indexing/app.py)

Deployment Options

You can run the server locally or via Docker:


# Build and run with Docker

docker build -t pathway-docstore -f templates/document_indexing/Dockerfile .
docker run -p 8000:8000 pathway-docstore

Once running, the server exposes the REST API at http://localhost:8000 (or your configured host/port).

Step 2: Integrate Pathway with LangChain as a Retriever Backend

LangChain treats Pathway as a vector store through the PathwayVectorClient class available in langchain_community.vectorstores.

LangChain Integration Code


# LangChain example – Pathway as a vector store

from langchain_community.vectorstores import PathwayVectorClient

# Connect to the running Pathway server

vectorstore_client = PathwayVectorClient(host="localhost", port=8000)

# Obtain a LangChain retriever

retriever = vectorstore_client.as_retriever(search_kwargs={"k": 5})

# Use it in a LangChain chain or agent

docs = retriever.invoke({"question": "What is the contract start date?"})
print(docs)

Source: cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb

Key Integration Points

Component Implementation
Import from langchain_community.vectorstores import PathwayVectorClient
Connection Instantiate with host and port of the Pathway server.
Retriever Call .as_retriever() to get a LangChain BaseRetriever compatible with LCEL chains.
Search parameters Pass search_kwargs={"k": N} to control the number of retrieved documents.

Step 3: Integrate Pathway with LlamaIndex as a Retriever Backend

LlamaIndex integrates with Pathway through the PathwayVectorStore class, which acts as a vector store wrapper around the Pathway HTTP API.

LlamaIndex Integration Code


# LlamaIndex example – Pathway as a retriever

from llama_index.vector_stores import PathwayVectorStore
from llama_index import SimpleDirectoryReader, VectorStoreIndex

# 1️⃣  Create a Pathway vector store wrapper

pathway_store = PathwayVectorStore(host="localhost", port=8000)

# 2️⃣  Build a LlamaIndex index that uses the Pathway store

documents = SimpleDirectoryReader("data/").load_data()
index = VectorStoreIndex.from_vector_store(pathway_store, documents=documents)

# 3️⃣  Query via the LlamaIndex retriever

retriever = index.as_retriever(similarity_top_k=5)
response = retriever.retrieve("What are the contract terms?")
print(response)

Reference: [templates/private_rag/README.md](https://github.com/pathwaycom/llm-app/blob/main/templates/private_rag/README.md)

Key Integration Points

Component Implementation
Vector store wrapper PathwayVectorStore(host, port) from llama_index.vector_stores.
Index construction Use VectorStoreIndex.from_vector_store(pathway_store, documents=...).
Retriever Call index.as_retriever(similarity_top_k=N) to get a LlamaIndex BaseRetriever.
Search results retrieve(query) returns NodeWithScore objects containing document text and metadata.

How the Integration Works Under the Hood

When you use Pathway as a retriever backend for LangChain or LlamaIndex, the following architecture enables real-time, continuous retrieval:

  1. Continuous Ingestion and Indexing: The DocumentStore class in pathway.xpacks.llm.document_store continuously ingests data from configured sources, computes embeddings using your specified model, and maintains a high-performance vector index (using usearch or persistent backends).

  2. REST API Exposure: The DocumentStoreServer class (instantiated in templates/document_indexing/app.py) exposes two critical HTTP endpoints:

    • /v2/upsert – For adding or updating vectors in the index
    • /v2/search – For vector similarity search (top-k retrieval)
  3. Client Translation: Both PathwayVectorClient (LangChain) and PathwayVectorStore (LlamaIndex) act as HTTP clients. They translate framework-specific method calls (like as_retriever().invoke() or retrieve()) into POST requests to the Pathway server's /v2/search endpoint.

  4. Real-Time Synchronization: Because the Pathway server runs as a standalone service with its own data connectors, any new files added to connected sources (S3, local folders, databases) are immediately indexed and become searchable through the LangChain or LlamaIndex retriever without restarting your application.

The retriever_factory pattern shown in templates/slides_ai_search/app.py demonstrates how to expose custom retriever logic that can be injected into downstream applications, further extending this architecture.

Summary

  • Pathway acts as a standalone retrieval service that continuously indexes data and exposes a REST API for vector search via DocumentStoreServer in templates/document_indexing/app.py.
  • LangChain integration uses PathwayVectorClient from langchain_community.vectorstores to connect to the running server and create a retriever with .as_retriever().
  • LlamaIndex integration uses PathwayVectorStore from llama_index.vector_stores as a vector store wrapper, building an index with VectorStoreIndex.from_vector_store() and retrieving with index.as_retriever().
  • Real-time capabilities are inherent because Pathway continuously updates its index as source data changes, making new documents immediately available to both LangChain and LlamaIndex retrievers without manual refreshes.

Frequently Asked Questions

What is the difference between PathwayVectorClient and PathwayVectorStore?

PathwayVectorClient is the LangChain-specific implementation found in langchain_community.vectorstores that conforms to LangChain’s VectorStore interface, allowing you to call .as_retriever() and use it in LCEL chains. PathwayVectorStore is the LlamaIndex implementation from llama_index.vector_stores that conforms to LlamaIndex’s BasePydanticVectorStore interface, enabling integration with VectorStoreIndex. Both act as HTTP clients communicating with the same Pathway DocumentStore server.

Does Pathway support real-time updates when used as a retriever backend?

Yes. Because Pathway runs as a continuous streaming engine, the DocumentStore in templates/document_indexing/app.py automatically indexes new documents as they appear in connected data sources (local files, S3, databases, etc.). When using PathwayVectorClient or PathwayVectorStore, your LangChain or LlamaIndex retriever immediately sees these updates without requiring manual index refreshes or application restarts.

Which file contains the server implementation for the DocumentStore?

The primary server implementation is located at templates/document_indexing/app.py. This file defines the App class that instantiates DocumentStore from pathway.xpacks.llm.document_store and wraps it with DocumentStoreServer from pathway.xpacks.llm.servers to expose the REST API endpoints required by LangChain and LlamaIndex clients.

Can I use Pathway with both LangChain and LlamaIndex simultaneously?

Yes. Because Pathway operates as a standalone service via DocumentStoreServer, you can run a single Pathway instance and connect multiple clients to it. For example, you can have a LangChain application using PathwayVectorClient and a separate LlamaIndex application using PathwayVectorStore both pointing to the same host:port endpoint, sharing the same real-time indexed data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →