How to Configure a Custom Vector Database Provider in AnythingLLM

To configure a custom vector database provider in AnythingLLM, extend the VectorDatabase base class, implement required methods such as connect and addDocumentToNamespace, register your provider in the getVectorDbClass factory, and set the VECTOR_DB environment variable to your custom key.

Mintplex-Labs/anything-llm abstracts vector storage operations behind a unified VectorDatabase interface, enabling seamless integration of proprietary or specialized databases. By implementing the contract defined in server/utils/vectorDbProviders/base.js and registering your provider in the factory at server/utils/helpers/index.js, you can configure a custom vector database provider that works transparently with document ingestion, chat retrieval, and the admin UI.

Understanding the VectorDatabase Architecture

AnythingLLM uses a provider pattern to isolate vector storage logic. The abstract VectorDatabase class in server/utils/vectorDbProviders/base.js defines the contract that all vector databases must follow, including methods for connection management, health checks, document insertion, and similarity search.

At runtime, the getVectorDbClass function in server/utils/helpers/index.js reads the VECTOR_DB environment variable and instantiates the matching provider via a switch statement. This factory pattern ensures that downstream components—such as document ingestion pipelines and chat retrieval systems—remain agnostic to the underlying storage implementation.

Implementing a Custom Vector Database Provider

Step 1: Create the Provider Module

Create a new directory under server/utils/vectorDbProviders/ and implement a class that extends VectorDatabase. The following skeleton demonstrates the required structure for a hypothetical mydb provider:

// server/utils/vectorDbProviders/mydb/index.js
const { VectorDatabase } = require("../base");

class MyDb extends VectorDatabase {
  constructor() {
    super();
  }

  get name() {
    return "MyDb";
  }

  async connect() {
    if (process.env.VECTOR_DB !== "mydb") {
      throw new Error(`${this.name}::Invalid ENV settings`);
    }

    const client = new MyDbClient({
      host: process.env.MYDB_HOST,
      apiKey: process.env.MYDB_API_KEY,
    });

    if (!(await client.isAlive())) {
      throw new Error(`${this.name}::Unable to reach the service`);
    }
    return { client };
  }

  async heartbeat() {
    const { client } = await this.connect();
    return { heartbeat: Date.now() };
  }

  async totalVectors() {
    const { client } = await this.connect();
    return await client.countAllVectors();
  }

  async namespaceCount(namespace) {
    const { client } = await this.connect();
    return await client.countVectorsInNamespace(namespace);
  }

  async addDocumentToNamespace(namespace, documentData, fullFilePath = null, skipCache = false) {
    const { client } = await this.connect();
    await client.upsert(namespace, documentData);
    return { vectorized: true, error: null };
  }

  // Implement similarityResponse, deleteDocumentFromNamespace, etc.
}

module.exports.MyDb = MyDb;

Step 2: Register in the Factory

Add a case to the switch statement in server/utils/helpers/index.js inside the getVectorDbClass function:

// Inside server/utils/helpers/index.js -> getVectorDbClass()
case "mydb":
  const { MyDb } = require("../vectorDbProviders/mydb");
  return new MyDb();

Step 3: Configure Environment Variables

Set the VECTOR_DB variable to your custom key and define any provider-specific configuration:


# .env

VECTOR_DB=mydb
MYDB_HOST=https://mydb.example.com
MYDB_API_KEY=your_api_key_here

Step 4: Verify the Integration

Restart the server and programmatically verify that your provider loads correctly:

const { getVectorDbClass } = require("./utils/helpers");
const vectorDb = getVectorDbClass();

await vectorDb.connect();
const count = await vectorDb.totalVectors();
console.log(`Connected to ${vectorDb.name}. Total vectors: ${count}`);

Required Methods and Interface Contract

According to the source code in server/utils/vectorDbProviders/base.js, a functional custom vector database provider must implement:

  • connect() – Establishes the client connection and returns an object containing the client instance.
  • heartbeat() – Returns a health check object to verify connectivity.
  • totalVectors() – Returns the total count of vectors across all namespaces.
  • namespaceCount(namespace) – Returns the vector count for a specific namespace.
  • addDocumentToNamespace(namespace, documentData, fullFilePath, skipCache) – Handles document embedding insertion and caching logic.
  • similarityResponse() – Performs vector similarity search for chat retrieval (required for RAG functionality).
  • deleteDocumentFromNamespace() – Removes documents and their vectors from a namespace.

Reference implementations in server/utils/vectorDbProviders/weaviate/index.js and server/utils/vectorDbProviders/zilliz/index.js demonstrate production-ready patterns for handling these operations with specific database clients.

Summary

  • Extend VectorDatabase from server/utils/vectorDbProviders/base.js to ensure interface compatibility.
  • Implement required methods including connect, heartbeat, totalVectors, namespaceCount, and addDocumentToNamespace.
  • Register your provider in server/utils/helpers/index.js by adding a case to the getVectorDbClass factory function.
  • Set VECTOR_DB to your custom key in the environment configuration.
  • Reference existing providers such as Weaviate (weaviate/index.js) and Zilliz (zilliz/index.js) for implementation patterns.

Frequently Asked Questions

What methods must I implement for a custom vector database provider?

You must implement the core interface methods defined in server/utils/vectorDbProviders/base.js: connect for initialization, heartbeat for health checks, totalVectors and namespaceCount for statistics, addDocumentToNamespace for ingestion, and similarityResponse for retrieval. Additional methods like deleteDocumentFromNamespace are required for full document lifecycle management.

How does AnythingLLM select which vector database to use?

The getVectorDbClass factory in server/utils/helpers/index.js reads the VECTOR_DB environment variable and instantiates the corresponding provider via a switch statement. When VECTOR_DB matches your registered key (e.g., mydb), the factory returns an instance of your custom class.

Can I use environment variables to configure my custom vector database?

Yes. The standard pattern, as shown in the connect() method implementations, checks process.env.VECTOR_DB for validation and reads provider-specific variables (such as MYDB_HOST or MYDB_API_KEY) to configure the client connection.

Where should I place my custom vector database provider files?

Create a new subdirectory under server/utils/vectorDbProviders/ (e.g., server/utils/vectorDbProviders/mydb/) containing an index.js file that exports your class. This location keeps your code colocated with built-in providers like Weaviate and Zilliz, ensuring clean imports in the factory registration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →