How to Migrate Existing Pinecone, Weaviate, or Qdrant Vector Database Implementations to Pathway
Migrate your existing vector database implementation to Pathway by replacing vendor-specific SDK calls with Pathway's built-in VectorStoreServer and VectorStoreClient, which use the ultra-fast usearch library for ANN search and Tantivy for hybrid full-text indexing.
Pathway's llm-app repository (pathwaycom/llm-app) provides a self-hosted, zero-maintenance alternative to external vector databases. This guide maps your existing Pinecone, Weaviate, or Qdrant components to their Pathway equivalents, implements the migration using actual source patterns from the codebase, and ensures your semantic search capabilities remain intact without external dependencies.
Architectural Mapping: Vendor SDK to Pathway
Understanding the component mapping is essential before modifying code. Each external service maps to a native Pathway class or connector:
Data Ingestion: Replace manual file loading or cloud SDK calls with Pathway I/O connectors. Use pw.io.fs.read for local files or pw.io.gdrive.read for cloud storage, as defined in templates/document_indexing/app.yaml.
Embedding Generation: Swap direct API calls (e.g., openai.Embedding.create) for Pathway embedder classes. The embedders.OpenAIEmbedder or embedders.SentenceTransformerEmbedder handle batching and caching automatically, referenced in cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb.
Document Processing: Replace pre-processing scripts with Pathway parsers and splitters. Use parsers.UnstructuredParser or parsers.DoclingParser combined with splitters.TokenCountSplitter to chunk documents before embedding.
Vector Storage: Substitute the remote index (Pinecone pod, Weaviate instance, or Qdrant collection) with the Pathway vector store server (VectorStoreServer). This class bundles the usearch ANN index and Tantivy full-text index into a single process.
Query Interface: Replace the vendor's query SDK with the Pathway client (VectorStoreClient). The query method maintains a compatible signature with top_k and include_metadata parameters.
Step 1: Install Pathway
Remove your existing vector database client libraries (pinecone-client, weaviate-client, qdrant-client) and install Pathway with all LLM extensions:
pip install "pathway[all]"
Pathway bundles the required usearch and tantivy binaries, eliminating the need for separate vector database installations or container orchestration.
Step 2: Create the Vector Store Server
The server replaces your external vector database instance. This self-contained script mirrors a typical Pinecone upsert pipeline but runs entirely on-premise:
# vector_store_server.py
import pathway as pw
from pathway.xpacks.llm.vector_store import VectorStoreServer
from pathway.xpacks.llm import embedders, parsers, splitters
# 1. Data source – replaces manual file loading before Pinecone upsert
folder = pw.io.fs.read(
path="data/*.txt",
format="binary",
with_metadata=True,
)
# 2. Document processing pipeline
parser = parsers.UnstructuredParser()
splitter = splitters.TokenCountSplitter(min_tokens=150, max_tokens=450)
# 3. Embedding – uses the same OpenAI model as your existing implementation
embedder = embedders.OpenAIEmbedder() # reads OPENAI_API_KEY from environment
# 4. Assemble the server
vector_server = VectorStoreServer(
folder, # supports multiple sources via *sources
embedder=embedder,
splitter=splitter,
parser=parser,
)
# 5. Run the server (default: host 0.0.0.0, port 8000)
vector_server.run_server(host="0.0.0.0", port=8000, threaded=True)
Key implementation details from cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb:
- The
VectorStoreServerautomatically builds the usearch ANN index and Tantivy full-text index, providing hybrid retrieval without separate configuration. - Data flows continuously from the connector through the embedding pipeline; there is no explicit "upsert" step as the server indexes live data automatically.
Step 3: Query with the Vector Store Client
Replace your existing query logic (e.g., index.query()) with the Pathway client. The response format mirrors Pinecone's structure to minimize refactoring:
# vector_store_client.py
from pathway.xpacks.llm.vector_store import VectorStoreClient
# Connect to the server
client = VectorStoreClient(host="localhost", port=8000)
# Query – same shape as Pinecone's query method
results = client.query(
query="What are the key risks in the contract?",
top_k=5,
include_metadata=True,
)
# Process results
for hit in results["hits"]:
print(f"Score: {hit['score']:.2f}, ID: {hit['id']}")
print(hit["metadata"]["source_path"])
As implemented in cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb, the VectorStoreClient returns a dictionary containing hits with score, id, and metadata fields, ensuring semantic equivalence with your existing implementation.
Step 4: Configure Persistence (Optional)
For production workloads requiring index survival across restarts, enable Pathway persistence as shown in templates/document_indexing/app.py:
# Add to vector_store_server.py
vector_server.run_server(
host="0.0.0.0",
port=8000,
threaded=True,
persistence_mode=pw.PersistenceMode.PERSISTING,
persistence_backend=pw.persistence.Backend.filesystem("./Cache")
)
This configuration writes index state to disk, providing durability comparable to managed vector database snapshots without external storage services.
Migration Checklist
Follow this sequence to ensure a complete transition:
- Remove vendor SDK dependencies (uninstall
pinecone-client,weaviate-client, orqdrant-client). - Add Pathway imports (
pathway as pw,pathway.xpacks.llm.vector_store). - Replace data ingestion with Pathway I/O connectors (
pw.io.fs.read,pw.io.gdrive.read). - Swap embedding calls for Pathway embedder objects (
embedders.OpenAIEmbedder). - Instantiate
VectorStoreServerinstead of connecting to a remote index. - Start the server using
run_server(local process, Docker, or Kubernetes). - Replace query calls with
VectorStoreClient.query. - Enable persistence if needed by setting
persistence_modeandpersistence_backend. - Update configuration in
app.yamlto reflect host, port, and connector credentials. - Test end-to-end to verify semantic search equivalence and metadata retrieval.
Deployment Options
Deploy the vector_store_server.py script using standard containerization:
- Local Development: Run directly with Python 3.10+ or use
docker run -p 8000:8000with a base Python image. - Kubernetes: Use the official
pathwaycom/pathway:latestimage, expose port 8000, and mount a PersistentVolumeClaim at/Cachewhen using persistence mode. - Cloud Containers: For AWS Fargate or Azure Container Instances, pass environment variables (e.g.,
OPENAI_API_KEY) via your orchestrator's secret management.
Reference the Dockerfiles in templates/document_indexing/ and templates/private_rag/ for production-ready container templates.
Summary
- Pathway's
VectorStoreServerprovides a drop-in replacement for Pinecone, Weaviate, and Qdrant using usearch and Tantivy for high-performance search. - Migration requires replacing vendor SDKs with Pathway connectors, embedders, and the server/client architecture.
- Hybrid search (semantic + full-text) is available out-of-the-box without additional services.
- Persistence and scaling are handled through Pathway's native
PersistenceModeand standard container orchestration, eliminating external database management overhead.
Frequently Asked Questions
Does Pathway's vector store support hybrid search like Weaviate?
Yes. The VectorStoreServer automatically maintains both a usearch ANN index for dense vector similarity and a Tantivy full-text index for sparse keyword search. You can perform hybrid queries without configuring separate services or index types, as the server handles both retrieval methods internally according to the implementation in cookbooks/self-rag-agents/pathway_langgraph_agentic_rag.ipynb.
Can I reuse my existing OpenAI embeddings when migrating from Pinecone?
Absolutely. Use the embedders.OpenAIEmbedder class with the same model name (e.g., text-embedding-ada-002) you used with Pinecone. Pathway reads the OPENAI_API_KEY environment variable and manages batching, ensuring your embedding logic remains identical while gaining Pathway's pipeline automation.
How does Pathway handle data persistence compared to Qdrant?
Pathway offers explicit persistence modes via pw.PersistenceMode.PERSISTING and backends like pw.persistence.Backend.filesystem. Unlike Qdrant's internal storage engine, Pathway writes state to your specified filesystem path (e.g., ./Cache), allowing you to mount network volumes or object storage mounts for durability, as demonstrated in templates/document_indexing/app.py.
Is the VectorStoreClient API compatible with LangChain?
Yes. The VectorStoreClient provides a standard query interface that integrates seamlessly with LangChain's retriever patterns. The response format (hits containing score, id, and metadata) aligns with expectations of LangChain vector store wrappers, making it straightforward to substitute into existing RAG chains without rewriting your application logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →