Pathway's Built-in usearch Vector Index Architecture vs. External Vector Databases
Pathway's built-in usearch vector index is an embedded, in-process K-Nearest-Neighbor engine that eliminates network latency and operational overhead by running as a thin Python wrapper around the Rust-native usearch library, directly inside your Pathway dataflow pipeline.
Pathway's LLM application framework (pathwaycom/llm-app) provides a high-performance vector search capability without requiring external infrastructure. Unlike traditional architectures that rely on separate vector database services, Pathway embeds the search engine directly into the runtime, leveraging the usearch library's SIMD-optimized algorithms to deliver sub-millisecond query latency.
What Is Pathway's Built-in usearch Vector Index?
Pathway ships its own in-memory vector index as a thin Python wrapper around the high-performance usearch library. The wrapper is exposed through pathway.stdlib.ml.index.KNNIndex, as demonstrated in the drive-alert template:
# From templates/drive_alert/app.py
from pathway.stdlib.ml.index import KNNIndex
# Construction inside the pipeline
knn_index = KNNIndex(
data=embedded_documents,
metric="cosine",
n_neighbors=5
)
This design choice embeds the vector search capability directly into the Python process running your Pathway pipeline, eliminating the need for network calls to external services.
Architectural Comparison: Built-in usearch vs. External Vector Databases
The architectural differences between Pathway's embedded approach and traditional external vector databases impact latency, deployment complexity, and operational requirements.
| Aspect | Pathway usearch (built-in) | Typical External Vector DB (Pinecone, Weaviate, Qdrant, Milvus) |
|---|---|---|
| Engine | Rust-native, lock-free SIMD + AVX2/AVX-512 optimised search (usearch). | Often a combination of C++/Rust back-ends (e.g., Faiss, HNSWLib) wrapped in a service. |
| Deployment model | Pure Python library – runs in the same process as your Pathway pipeline, no separate service. | Separate service (cloud-hosted or self-hosted) accessed over HTTP/gRPC. |
| Persistence | In-memory only (optional snapshotting via Python pickle/Pathway checkpoint). |
Durable on-disk storage, replication, backups. |
| Scalability | Scales with the memory of the host process; suited for "millions of vectors" in a single node. | Horizontal scaling via cluster nodes, sharding, and load-balancing. |
| Latency | Sub-millisecond query latency because data lives in the same address space. | Additional network round-trip adds latency (typically 1-10 ms). |
| API surface | KNNIndex(data, metric="cosine", n_neighbors=5, ...) – direct Python calls, fully async-compatible with Pathway's dataflow. |
REST/gRPC endpoints (/query, /upsert) that require request/response handling. |
| Hybrid search | Pathway couples the usearch vector index with a Tantivy full-text index in the same pipeline. | Hybrid search is usually achieved by running two separate services (vector + full-text) and merging results in the client. |
| Operational overhead | No extra infrastructure to provision, configure, or monitor. | Requires ops for service deployment, scaling, authentication, quota management. |
Implementation Examples
Creating a KNNIndex in a Pathway Pipeline
The following example demonstrates constructing the built-in index within a Pathway dataflow, as seen in the templates/question_answering_rag implementation:
import pathway as pw
from pathway.stdlib.ml.index import KNNIndex
from pathway.xpacks.llm.embedders import OpenAIEmbedder
# Define the data schema
class Doc(pw.Schema):
text: str
embedding: float[384] # dimension of OpenAI embeddings
@pw.flow
def index_pipeline(docs: pw.DataFrame[Doc]) -> pw.DataFrame:
# Compute embeddings (sync or async)
emb = docs.select(
text=pw.col("text"),
embedding=OpenAIEmbedder()(pw.col("text"))
)
# Build the nearest-neighbor index (in-memory)
knn = KNNIndex(
data=emb,
metric="cosine",
n_neighbors=5,
index_name="doc_index", # optional identifier
)
return knn
# Run the pipeline
doc_source = pw.io.fs.read("data/*.txt", schema=Doc)
indexed = index_pipeline(doc_source)
pw.run()
Key points:
KNNIndexis instantiated directly in Python – no external service URL is required.- The index updates automatically as new rows flow through the pipeline.
Querying the Built-in Index
Querying the embedded index uses direct method calls on the index object within the same process:
@pw.flow
def query_pipeline(query_text: str) -> pw.DataFrame:
# Embed the query
q_emb = OpenAIEmbedder()(pw.to_df({"text": [query_text]}))
# Retrieve nearest neighbours
results = indexed.get_nearest_items(q_emb["embedding"], k=3)
return results
# Example usage
answers = query_pipeline("What is the refund policy?")
pw.run()
print(answers)
Contrast with External Vector Database (Pinecone)
For comparison, interacting with an external service like Pinecone requires network initialization, API keys, and HTTP/gRPC calls:
import pinecone
import openai
# Initialise Pinecone client (external service)
pinecone.init(api_key="YOUR_KEY", environment="us-west1-gcp")
index = pinecone.Index("my-index")
# Upsert vectors
emb = openai.Embedding.create(input=texts, model="text-embedding-ada-002")["data"]
vectors = [(str(i), e["embedding"]) for i, e in enumerate(emb)]
index.upsert(vectors=vectors)
# Query
q_emb = openai.Embedding.create(input=[query], model="text-embedding-ada-002")["data"][0]["embedding"]
res = index.query(vector=q_emb, top_k=3)
print(res.matches)
Notice the additional steps: client initialisation, network calls, and the need for an API key.
Key Source Files in the Repository
The following files demonstrate how Pathway implements and utilizes the built-in usearch vector index:
| File | Role | Link |
|---|---|---|
README.md |
Explains that Pathway uses the usearch library for its built-in vector index and Tantivy for hybrid full-text. |
README.md |
templates/drive_alert/app.py |
Shows the concrete import of KNNIndex and its construction (KNNIndex(). |
drive_alert/app.py |
templates/question_answering_rag/app.py |
Demonstrates a full RAG pipeline that uses the same index under the hood. | question_answering_rag/app.py |
templates/document_indexing/app.py |
Another template that exposes the index as a micro-service. | document_indexing/app.py |
These files together illustrate how Pathway embeds the usearch-based KNNIndex directly inside a Python data-flow, eliminating the need for an external vector DB.
Summary
- Pathway's built-in usearch vector index is an embedded, in-process solution that wraps the Rust-native usearch library, providing sub-millisecond query latency without network overhead.
- The index is exposed via
pathway.stdlib.ml.index.KNNIndexand operates directly within the Pathway dataflow, automatically updating as new data streams through the pipeline. - Unlike external vector databases (Pinecone, Weaviate, Qdrant, Milvus), Pathway's solution requires no separate service deployment, API keys, or network round-trips, though it scales only to the memory limits of a single node.
- Hybrid search capabilities are achieved by coupling the usearch vector index with Tantivy full-text indexing within the same pipeline, a feat that typically requires coordinating multiple external services in traditional architectures.
Frequently Asked Questions
What is the primary advantage of Pathway's built-in usearch vector index over external databases?
The primary advantage is sub-millisecond latency and zero operational overhead. Because the index runs in the same process as your Pathway pipeline, queries execute without network round-trips, authentication handshakes, or serialization overhead. This embedded architecture eliminates the need to provision, scale, or monitor separate vector database services, making it ideal for applications requiring real-time RAG (Retrieval-Augmented Generation) responses.
How does the hybrid search capability work in Pathway?
Pathway achieves hybrid search by combining the usearch vector index with a Tantivy full-text index within the same dataflow pipeline. As documents stream through the system, Pathway can simultaneously update both the vector embeddings (via KNNIndex) and the inverted text index. Queries can then perform semantic similarity searches via usearch while also filtering or boosting results based on keyword matches from Tantivy, all without coordinating multiple external services or merging results client-side.
Is Pathway's built-in vector index suitable for production-scale deployments?
Pathway's built-in index is production-ready for single-node deployments handling millions of vectors, provided the host has sufficient memory to hold the entire index. It scales vertically with available RAM and is optimized for "millions of vectors" workloads on a single node. However, for use cases requiring horizontal scaling across multiple nodes, petabyte-scale persistence, or cross-region replication, external vector databases with distributed architectures remain the appropriate choice. Pathway checkpoints can mitigate durability concerns by snapshotting the index state via Python pickle or Pathway's native checkpointing mechanisms.
What embedding models are compatible with Pathway's KNNIndex?
KNNIndex is embedding-model agnostic and accepts any fixed-dimensional vector representation. The index stores vectors as arrays of floats (e.g., float[384] or float[1536]) and supports various distance metrics including cosine similarity, Euclidean distance, and inner product. You can generate embeddings using Pathway's built-in embedders (such as OpenAIEmbedder or SentenceTransformerEmbedder) or provide your own pre-computed vectors from models like OpenAI's text-embedding-3, Cohere, or custom fine-tuned embeddings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →