Pathway's Built-in usearch Vector Index Architecture vs. External Vector Databases

Pathway's built-in usearch vector index is an embedded, in-process K-Nearest-Neighbor engine that eliminates network latency and operational overhead by running as a thin Python wrapper around the Rust-native usearch library, directly inside your Pathway dataflow pipeline.

Pathway's LLM application framework (pathwaycom/llm-app) provides a high-performance vector search capability without requiring external infrastructure. Unlike traditional architectures that rely on separate vector database services, Pathway embeds the search engine directly into the runtime, leveraging the usearch library's SIMD-optimized algorithms to deliver sub-millisecond query latency.

What Is Pathway's Built-in usearch Vector Index?

Pathway ships its own in-memory vector index as a thin Python wrapper around the high-performance usearch library. The wrapper is exposed through pathway.stdlib.ml.index.KNNIndex, as demonstrated in the drive-alert template:


# From templates/drive_alert/app.py

from pathway.stdlib.ml.index import KNNIndex

# Construction inside the pipeline

knn_index = KNNIndex(
    data=embedded_documents,
    metric="cosine",
    n_neighbors=5
)

This design choice embeds the vector search capability directly into the Python process running your Pathway pipeline, eliminating the need for network calls to external services.

Architectural Comparison: Built-in usearch vs. External Vector Databases

The architectural differences between Pathway's embedded approach and traditional external vector databases impact latency, deployment complexity, and operational requirements.

Aspect Pathway usearch (built-in) Typical External Vector DB (Pinecone, Weaviate, Qdrant, Milvus)
Engine Rust-native, lock-free SIMD + AVX2/AVX-512 optimised search (usearch). Often a combination of C++/Rust back-ends (e.g., Faiss, HNSWLib) wrapped in a service.
Deployment model Pure Python library – runs in the same process as your Pathway pipeline, no separate service. Separate service (cloud-hosted or self-hosted) accessed over HTTP/gRPC.
Persistence In-memory only (optional snapshotting via Python pickle/Pathway checkpoint). Durable on-disk storage, replication, backups.
Scalability Scales with the memory of the host process; suited for "millions of vectors" in a single node. Horizontal scaling via cluster nodes, sharding, and load-balancing.
Latency Sub-millisecond query latency because data lives in the same address space. Additional network round-trip adds latency (typically 1-10 ms).
API surface KNNIndex(data, metric="cosine", n_neighbors=5, ...) – direct Python calls, fully async-compatible with Pathway's dataflow. REST/gRPC endpoints (/query, /upsert) that require request/response handling.
Hybrid search Pathway couples the usearch vector index with a Tantivy full-text index in the same pipeline. Hybrid search is usually achieved by running two separate services (vector + full-text) and merging results in the client.
Operational overhead No extra infrastructure to provision, configure, or monitor. Requires ops for service deployment, scaling, authentication, quota management.

Implementation Examples

Creating a KNNIndex in a Pathway Pipeline

The following example demonstrates constructing the built-in index within a Pathway dataflow, as seen in the templates/question_answering_rag implementation:

import pathway as pw
from pathway.stdlib.ml.index import KNNIndex
from pathway.xpacks.llm.embedders import OpenAIEmbedder

# Define the data schema

class Doc(pw.Schema):
    text: str
    embedding: float[384]   # dimension of OpenAI embeddings

@pw.flow
def index_pipeline(docs: pw.DataFrame[Doc]) -> pw.DataFrame:
    # Compute embeddings (sync or async)

    emb = docs.select(
        text=pw.col("text"),
        embedding=OpenAIEmbedder()(pw.col("text"))
    )
    # Build the nearest-neighbor index (in-memory)

    knn = KNNIndex(
        data=emb,
        metric="cosine",
        n_neighbors=5,
        index_name="doc_index",    # optional identifier

    )
    return knn

# Run the pipeline

doc_source = pw.io.fs.read("data/*.txt", schema=Doc)
indexed = index_pipeline(doc_source)
pw.run()

Key points:

  • KNNIndex is instantiated directly in Python – no external service URL is required.
  • The index updates automatically as new rows flow through the pipeline.

Querying the Built-in Index

Querying the embedded index uses direct method calls on the index object within the same process:

@pw.flow
def query_pipeline(query_text: str) -> pw.DataFrame:
    # Embed the query

    q_emb = OpenAIEmbedder()(pw.to_df({"text": [query_text]}))
    # Retrieve nearest neighbours

    results = indexed.get_nearest_items(q_emb["embedding"], k=3)
    return results

# Example usage

answers = query_pipeline("What is the refund policy?")
pw.run()
print(answers)

Contrast with External Vector Database (Pinecone)

For comparison, interacting with an external service like Pinecone requires network initialization, API keys, and HTTP/gRPC calls:

import pinecone
import openai

# Initialise Pinecone client (external service)

pinecone.init(api_key="YOUR_KEY", environment="us-west1-gcp")
index = pinecone.Index("my-index")

# Upsert vectors

emb = openai.Embedding.create(input=texts, model="text-embedding-ada-002")["data"]
vectors = [(str(i), e["embedding"]) for i, e in enumerate(emb)]
index.upsert(vectors=vectors)

# Query

q_emb = openai.Embedding.create(input=[query], model="text-embedding-ada-002")["data"][0]["embedding"]
res = index.query(vector=q_emb, top_k=3)
print(res.matches)

Notice the additional steps: client initialisation, network calls, and the need for an API key.

Key Source Files in the Repository

The following files demonstrate how Pathway implements and utilizes the built-in usearch vector index:

File Role Link
README.md Explains that Pathway uses the usearch library for its built-in vector index and Tantivy for hybrid full-text. README.md
templates/drive_alert/app.py Shows the concrete import of KNNIndex and its construction (KNNIndex(). drive_alert/app.py
templates/question_answering_rag/app.py Demonstrates a full RAG pipeline that uses the same index under the hood. question_answering_rag/app.py
templates/document_indexing/app.py Another template that exposes the index as a micro-service. document_indexing/app.py

These files together illustrate how Pathway embeds the usearch-based KNNIndex directly inside a Python data-flow, eliminating the need for an external vector DB.

Summary

  • Pathway's built-in usearch vector index is an embedded, in-process solution that wraps the Rust-native usearch library, providing sub-millisecond query latency without network overhead.
  • The index is exposed via pathway.stdlib.ml.index.KNNIndex and operates directly within the Pathway dataflow, automatically updating as new data streams through the pipeline.
  • Unlike external vector databases (Pinecone, Weaviate, Qdrant, Milvus), Pathway's solution requires no separate service deployment, API keys, or network round-trips, though it scales only to the memory limits of a single node.
  • Hybrid search capabilities are achieved by coupling the usearch vector index with Tantivy full-text indexing within the same pipeline, a feat that typically requires coordinating multiple external services in traditional architectures.

Frequently Asked Questions

What is the primary advantage of Pathway's built-in usearch vector index over external databases?

The primary advantage is sub-millisecond latency and zero operational overhead. Because the index runs in the same process as your Pathway pipeline, queries execute without network round-trips, authentication handshakes, or serialization overhead. This embedded architecture eliminates the need to provision, scale, or monitor separate vector database services, making it ideal for applications requiring real-time RAG (Retrieval-Augmented Generation) responses.

How does the hybrid search capability work in Pathway?

Pathway achieves hybrid search by combining the usearch vector index with a Tantivy full-text index within the same dataflow pipeline. As documents stream through the system, Pathway can simultaneously update both the vector embeddings (via KNNIndex) and the inverted text index. Queries can then perform semantic similarity searches via usearch while also filtering or boosting results based on keyword matches from Tantivy, all without coordinating multiple external services or merging results client-side.

Is Pathway's built-in vector index suitable for production-scale deployments?

Pathway's built-in index is production-ready for single-node deployments handling millions of vectors, provided the host has sufficient memory to hold the entire index. It scales vertically with available RAM and is optimized for "millions of vectors" workloads on a single node. However, for use cases requiring horizontal scaling across multiple nodes, petabyte-scale persistence, or cross-region replication, external vector databases with distributed architectures remain the appropriate choice. Pathway checkpoints can mitigate durability concerns by snapshotting the index state via Python pickle or Pathway's native checkpointing mechanisms.

What embedding models are compatible with Pathway's KNNIndex?

KNNIndex is embedding-model agnostic and accepts any fixed-dimensional vector representation. The index stores vectors as arrays of floats (e.g., float[384] or float[1536]) and supports various distance metrics including cosine similarity, Euclidean distance, and inner product. You can generate embeddings using Pathway's built-in embedders (such as OpenAIEmbedder or SentenceTransformerEmbedder) or provide your own pre-computed vectors from models like OpenAI's text-embedding-3, Cohere, or custom fine-tuned embeddings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →