How to Migrate Between Different Vector Database Providers in AnythingLLM

Migrating between vector database providers in AnythingLLM requires updating the VECTOR_DB environment variable, resetting existing vector stores via resetAllVectorStores, and allowing the system to automatically re-index your documents.

AnythingLLM uses a provider-pattern abstraction to support multiple vector databases including Pinecone, Weaviate, Qdrant, and LanceDB. Changing your vector store involves configuration updates followed by a complete reset and re-indexing process managed by the Mintplex-Labs/anything-llm codebase.

Understanding the Vector Database Provider Pattern

AnythingLLM abstracts vector storage behind a unified interface defined in server/utils/vectorDbProviders/base.js. The environment variable VECTOR_DB determines which concrete implementation gets instantiated at runtime.

In server/utils/helpers/index.js, the getVectorDbClass function reads process.env.VECTOR_DB and returns the appropriate provider instance:

function getVectorDbClass(getExactly = null) {
  const vectorSelection = getExactly ?? process.env.VECTOR_DB ?? "lancedb";
  switch (vectorSelection) {
    case "pinecone":
      const { Pinecone } = require("../vectorDbProviders/pinecone");
      return new Pinecone();
    case "pgvector":
      const { PGVector } = require("../vectorDbProviders/pgvector");
      return new PGVector();
    default:
      const { LanceDb: DefaultLanceDb } = require("../vectorDbProviders/lance");
      return new DefaultLanceDb();
  }
}

When you change providers, the system must reset all existing vectors because embedding dimensions and schemas vary between stores.

Step 1: Install and Configure the New Provider

Before updating environment variables, install any required dependencies for your target provider:

Provider NPM Package Setup Documentation
Pinecone @pinecone-database/pinecone server/utils/vectorDbProviders/pinecone/PINECONE_SETUP.md
Weaviate weaviate-ts-client server/utils/vectorDbProviders/weaviate/WEAVIATE_SETUP.md
Qdrant @qdrant/qdrant-client server/utils/vectorDbProviders/qdrant/QDRANT_SETUP.md
Milvus @zilliz/milvus2-sdk-node server/utils/vectorDbProviders/milvus/MILVUS_SETUP.md
AstraDB @astradb/astrodb server/utils/vectorDbProviders/astra/ASTRA_SETUP.md
PGVector pg + pgvector Built-in (Postgres-based)
LanceDB vector-db-lancedb Default (no extra steps)

Add the package to your server package.json and follow the provider-specific setup documentation for API keys and endpoint configuration.

Step 2: Update Environment Variables

Set the VECTOR_DB variable and provider-specific credentials in your .env file or via the UI Settings page:

VECTOR_DB=qdrant
QDRANT_URL=https://my-qdrant-instance:6333
QDRANT_API_KEY=your-api-key-here

Available options for VECTOR_DB include: pinecone, weaviate, qdrant, milvus, zilliz, astra, pgvector, and lancedb.

Step 3: Reset Existing Vector Stores

When VECTOR_DB changes, AnythingLLM automatically triggers resetAllVectorStores via the updateENV flow. To manually reset before switching, use:

const { resetAllVectorStores } = require("./server/utils/vectorStore/resetAllVectorStores");

// Specify the provider you are leaving
await resetAllVectorStores({ vectorDbKey: "pinecone" });

The reset implementation in server/utils/vectorStore/resetAllVectorStores.js (lines 18-34) handles cleanup differently based on the provider:

async function resetAllVectorStores({ vectorDbKey }) {
  const workspaces = await Workspace.where();
  purgeEntireVectorCache();
  await DocumentVectors.delete();
  await Document.delete();
  
  const VectorDb = getVectorDbClass(vectorDbKey);
  if (vectorDbKey === "pgvector") {
    await VectorDb.reset();  // Drops entire table
  } else {
    for (const workspace of workspaces) {
      await VectorDb["delete-namespace"]({ namespace: workspace.slug });
    }
  }
}

Resetting removes incompatible vectors from the previous store before re-indexing begins.

Step 4: Trigger Re-indexing

After resetting, the system automatically re-indexes documents when you perform queries. The re-indexing flow in server/models/documents.js fetches the current provider and stores new embeddings:

async function addDocument({ workspaceSlug, filePath, ...metadata }) {
  const VectorDb = getVectorDbClass();  // Uses current VECTOR_DB
  // ... embedding generation ...
  await VectorDb.addDocumentToNamespace(workspaceSlug, docData, filePath);
}

Simply open a workspace or run a search to trigger the embedding process. The system detects missing vectors and populates the new store automatically.

Step 5: Verify the Migration

Confirm the new provider is active using the heartbeat check:

node -e "require('./server/utils/helpers').getVectorDbClass().heartbeat().then(console.log)"

You should receive a JSON response with a recent timestamp indicating successful connectivity.

Additionally, query the namespace statistics endpoint to verify vector counts:

GET /api/v1/vector-db/namespace-stats

This endpoint delegates to the concrete provider's namespace-stats implementation.

Rollback Procedure

If issues arise, revert VECTOR_DB to the previous value, run resetAllVectorStores for that provider key, and allow the system to re-index again.

Summary

  • Update VECTOR_DB in environment variables to select the new provider (e.g., weaviate, qdrant).
  • Install dependencies and configure provider-specific credentials per the *_SETUP.md files.
  • Reset vector stores using resetAllVectorStores to clear old, incompatible embeddings.
  • Re-index automatically by querying workspaces; getVectorDbClass instantiates the new provider on-demand.
  • Verify connectivity via the heartbeat() method and namespace statistics endpoints.

Frequently Asked Questions

Does AnythingLLM support zero-downtime vector database migration?

No, migrating between vector database providers requires a full reset and re-index of all documents. The system calls resetAllVectorStores to purge old vectors because embedding schemas differ between providers. Plan for brief downtime while documents re-embed into the new store.

Will my existing embeddings transfer to the new vector database?

No, existing embeddings do not transfer between providers. The resetAllVectorStores function in server/utils/vectorStore/resetAllVectorStores.js explicitly deletes old DocumentVectors and purges caches. Documents must be re-embedded using the new provider's connection, which happens automatically when you query workspaces after migration.

How do I know which VECTOR_DB values are valid?

Valid values correspond to the switch cases in server/utils/helpers/index.js: pinecone, weaviate, qdrant, milvus, zilliz, astra, pgvector, and lancedb. LanceDB serves as the default when VECTOR_DB is unset. Each provider has a dedicated directory under server/utils/vectorDbProviders/ with setup documentation.

What happens if I change VECTOR_DB without resetting?

The system automatically triggers resetAllVectorStores when the environment variable changes via the updateENV flow. However, if you manually edit .env without restarting properly, you may encounter errors because getVectorDbClass will attempt to query a non-existent namespace in the new store. Always ensure the reset function runs to clear stale metadata.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →