# How to Migrate Existing Pinecone, Weaviate, or Qdrant Vector Database Implementations to Pathway

> Easily migrate Pinecone, Weaviate, or Qdrant vector databases to Pathway. Replace SDK calls with Pathway's VectorStoreServer and VectorStoreClient for faster ANN and hybrid search.

- Repository: [Pathway/llm-app](https://github.com/pathwaycom/llm-app)
- Tags: migration-guide
- Published: 2026-03-07

---

**Migrate your existing vector database implementation to Pathway by replacing vendor-specific SDK calls with Pathway's built-in `VectorStoreServer` and `VectorStoreClient`, which use the ultra-fast usearch library for ANN search and Tantivy for hybrid full-text indexing.**

Pathway's `llm-app` repository (pathwaycom/llm-app) provides a self-hosted, zero-maintenance alternative to external vector databases. This guide maps your existing Pinecone, Weaviate, or Qdrant components to their Pathway equivalents, implements the migration using actual source patterns from the codebase, and ensures your semantic search capabilities remain intact without external dependencies.

## Architectural Mapping: Vendor SDK to Pathway

Understanding the component mapping is essential before modifying code. Each external service maps to a native Pathway class or connector:

**Data Ingestion**: Replace manual file loading or cloud SDK calls with **Pathway I/O connectors**. Use `pw.io.fs.read` for local files or `pw.io.gdrive.read` for cloud storage, as defined in [`templates/document_indexing/app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/templates/document_indexing/app.yaml).

**Embedding Generation**: Swap direct API calls (e.g., `openai.Embedding.create`) for **Pathway embedder classes**. The `embedders.OpenAIEmbedder` or `embedders.SentenceTransformerEmbedder` handle batching and caching automatically, referenced in `cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb`.

**Document Processing**: Replace pre-processing scripts with **Pathway parsers and splitters**. Use `parsers.UnstructuredParser` or `parsers.DoclingParser` combined with `splitters.TokenCountSplitter` to chunk documents before embedding.

**Vector Storage**: Substitute the remote index (Pinecone pod, Weaviate instance, or Qdrant collection) with the **Pathway vector store server** (`VectorStoreServer`). This class bundles the usearch ANN index and Tantivy full-text index into a single process.

**Query Interface**: Replace the vendor's query SDK with the **Pathway client** (`VectorStoreClient`). The `query` method maintains a compatible signature with `top_k` and `include_metadata` parameters.

## Step 1: Install Pathway

Remove your existing vector database client libraries (`pinecone-client`, `weaviate-client`, `qdrant-client`) and install Pathway with all LLM extensions:

```bash
pip install "pathway[all]"

```

Pathway bundles the required `usearch` and `tantivy` binaries, eliminating the need for separate vector database installations or container orchestration.

## Step 2: Create the Vector Store Server

The server replaces your external vector database instance. This self-contained script mirrors a typical Pinecone upsert pipeline but runs entirely on-premise:

```python

# vector_store_server.py

import pathway as pw
from pathway.xpacks.llm.vector_store import VectorStoreServer
from pathway.xpacks.llm import embedders, parsers, splitters

# 1. Data source – replaces manual file loading before Pinecone upsert

folder = pw.io.fs.read(
    path="data/*.txt",
    format="binary",
    with_metadata=True,
)

# 2. Document processing pipeline

parser = parsers.UnstructuredParser()
splitter = splitters.TokenCountSplitter(min_tokens=150, max_tokens=450)

# 3. Embedding – uses the same OpenAI model as your existing implementation

embedder = embedders.OpenAIEmbedder()  # reads OPENAI_API_KEY from environment

# 4. Assemble the server

vector_server = VectorStoreServer(
    folder,  # supports multiple sources via *sources

    embedder=embedder,
    splitter=splitter,
    parser=parser,
)

# 5. Run the server (default: host 0.0.0.0, port 8000)

vector_server.run_server(host="0.0.0.0", port=8000, threaded=True)

```

Key implementation details from `cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb`:

- The `VectorStoreServer` automatically builds the usearch ANN index and Tantivy full-text index, providing hybrid retrieval without separate configuration.
- Data flows continuously from the connector through the embedding pipeline; there is no explicit "upsert" step as the server indexes live data automatically.

## Step 3: Query with the Vector Store Client

Replace your existing query logic (e.g., `index.query()`) with the Pathway client. The response format mirrors Pinecone's structure to minimize refactoring:

```python

# vector_store_client.py

from pathway.xpacks.llm.vector_store import VectorStoreClient

# Connect to the server

client = VectorStoreClient(host="localhost", port=8000)

# Query – same shape as Pinecone's query method

results = client.query(
    query="What are the key risks in the contract?",
    top_k=5,
    include_metadata=True,
)

# Process results

for hit in results["hits"]:
    print(f"Score: {hit['score']:.2f}, ID: {hit['id']}")
    print(hit["metadata"]["source_path"])

```

As implemented in `cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb`, the `VectorStoreClient` returns a dictionary containing `hits` with `score`, `id`, and metadata fields, ensuring semantic equivalence with your existing implementation.

## Step 4: Configure Persistence (Optional)

For production workloads requiring index survival across restarts, enable Pathway persistence as shown in [`templates/document_indexing/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/document_indexing/app.py):

```python

# Add to vector_store_server.py

vector_server.run_server(
    host="0.0.0.0",
    port=8000,
    threaded=True,
    persistence_mode=pw.PersistenceMode.PERSISTING,
    persistence_backend=pw.persistence.Backend.filesystem("./Cache")
)

```

This configuration writes index state to disk, providing durability comparable to managed vector database snapshots without external storage services.

## Migration Checklist

Follow this sequence to ensure a complete transition:

1. **Remove vendor SDK dependencies** (uninstall `pinecone-client`, `weaviate-client`, or `qdrant-client`).
2. **Add Pathway imports** (`pathway as pw`, `pathway.xpacks.llm.vector_store`).
3. **Replace data ingestion** with Pathway I/O connectors (`pw.io.fs.read`, `pw.io.gdrive.read`).
4. **Swap embedding calls** for Pathway embedder objects (`embedders.OpenAIEmbedder`).
5. **Instantiate `VectorStoreServer`** instead of connecting to a remote index.
6. **Start the server** using `run_server` (local process, Docker, or Kubernetes).
7. **Replace query calls** with `VectorStoreClient.query`.
8. **Enable persistence** if needed by setting `persistence_mode` and `persistence_backend`.
9. **Update configuration** in [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml) to reflect host, port, and connector credentials.
10. **Test end-to-end** to verify semantic search equivalence and metadata retrieval.

## Deployment Options

Deploy the [`vector_store_server.py`](https://github.com/pathwaycom/llm-app/blob/main/vector_store_server.py) script using standard containerization:

- **Local Development**: Run directly with Python 3.10+ or use `docker run -p 8000:8000` with a base Python image.
- **Kubernetes**: Use the official `pathwaycom/pathway:latest` image, expose port 8000, and mount a PersistentVolumeClaim at `/Cache` when using persistence mode.
- **Cloud Containers**: For AWS Fargate or Azure Container Instances, pass environment variables (e.g., `OPENAI_API_KEY`) via your orchestrator's secret management.

Reference the Dockerfiles in `templates/document_indexing/` and `templates/private_rag/` for production-ready container templates.

## Summary

- **Pathway's `VectorStoreServer`** provides a drop-in replacement for Pinecone, Weaviate, and Qdrant using usearch and Tantivy for high-performance search.
- **Migration requires** replacing vendor SDKs with Pathway connectors, embedders, and the server/client architecture.
- **Hybrid search** (semantic + full-text) is available out-of-the-box without additional services.
- **Persistence and scaling** are handled through Pathway's native `PersistenceMode` and standard container orchestration, eliminating external database management overhead.

## Frequently Asked Questions

### Does Pathway's vector store support hybrid search like Weaviate?

Yes. The `VectorStoreServer` automatically maintains both a **usearch** ANN index for dense vector similarity and a **Tantivy** full-text index for sparse keyword search. You can perform hybrid queries without configuring separate services or index types, as the server handles both retrieval methods internally according to the implementation in `cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb`.

### Can I reuse my existing OpenAI embeddings when migrating from Pinecone?

Absolutely. Use the `embedders.OpenAIEmbedder` class with the same model name (e.g., `text-embedding-ada-002`) you used with Pinecone. Pathway reads the `OPENAI_API_KEY` environment variable and manages batching, ensuring your embedding logic remains identical while gaining Pathway's pipeline automation.

### How does Pathway handle data persistence compared to Qdrant?

Pathway offers explicit persistence modes via `pw.PersistenceMode.PERSISTING` and backends like `pw.persistence.Backend.filesystem`. Unlike Qdrant's internal storage engine, Pathway writes state to your specified filesystem path (e.g., `./Cache`), allowing you to mount network volumes or object storage mounts for durability, as demonstrated in [`templates/document_indexing/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/document_indexing/app.py).

### Is the VectorStoreClient API compatible with LangChain?

Yes. The `VectorStoreClient` provides a standard query interface that integrates seamlessly with LangChain's retriever patterns. The response format (`hits` containing `score`, `id`, and `metadata`) aligns with expectations of LangChain vector store wrappers, making it straightforward to substitute into existing RAG chains without rewriting your application logic.