How to Configure Vector Database Providers in Cognee: A Complete Guide

Configure vector database providers in Cognee by setting the VECTOR_DB_PROVIDER environment variable or programmatically saving a VectorDBConfig, which the factory at create_vector_engine() uses to instantiate the appropriate adapter.

Cognee is an open-source knowledge graph library that stores embeddings in a pluggable vector database layer. Understanding how to configure vector database providers allows you to switch between LanceDB, PGVector, ChromaDB, or Neptune Analytics—or even integrate custom storage backends—without modifying core application logic.

Understanding the Vector Database Configuration Pipeline

The configuration flow follows a provider-based factory pattern. First, VectorConfig in cognee/infrastructure/databases/vector/config.py reads environment variables and caches the settings via get_vectordb_config() (lines 78-92). When your code requests a vector engine, create_vector_engine() in cognee/infrastructure/databases/vector/create_vector_engine.py (lines 10-30) inspects the provider name and instantiates the matching adapter class.

The system checks supported_databases (populated by external plugins) first, then falls through to built-in elif blocks for pgvector, chromadb, neptune_analytics, and lancedb. Each adapter receives the database URL, optional API key, and a shared embedding engine from get_embedding_engine(). Global access throughout the codebase happens through get_vector_engine() in cognee/infrastructure/databases/vector/get_vector_engine.py, which reads context-aware config via get_vectordb_context_config().

Supported Vector Database Providers

Cognee ships with four built-in vector database providers. Each requires specific configuration parameters and instantiates differently within the factory.

LanceDB (Default)

LanceDB is the default provider when no configuration is specified. It requires vector_db_url and optionally accepts vector_db_key.

In cognee/infrastructure/databases/vector/create_vector_engine.py (lines 84-91), the factory constructs:

LanceDBAdapter(
    url=config.vector_db_url,
    api_key=config.vector_db_key,
    embedding_engine=embedding_engine
)

If VECTOR_DB_URL is unset, Cognee defaults to a local directory under <system_root>/databases/cognee.lancedb.

PGVector

PGVector requires PostgreSQL credentials. You can provide explicit variables (VECTOR_DB_USERNAME, VECTOR_DB_PASSWORD, VECTOR_DB_HOST, VECTOR_DB_PORT, VECTOR_DB_NAME) or rely on the relational database configuration (DB_* environment variables).

The factory builds an asyncpg connection string (lines 92-125) and instantiates:

PGVectorAdapter(
    connection_string=connection_string,
    api_key=config.vector_db_key,
    embedding_engine=embedding_engine
)

ChromaDB

ChromaDB requires vector_db_url and optionally vector_db_key. The factory imports chromadb directly and constructs ChromaDBAdapter (lines 48-54):

ChromaDBAdapter(
    url=config.vector_db_url,
    api_key=config.vector_db_key,
    embedding_engine=embedding_engine
)

Neptune Analytics

Neptune Analytics requires vector_db_url containing the Neptune endpoint. The factory validates the URL prefix, extracts the graph ID (lines 66-82), and creates:

NeptuneAnalyticsAdapter(
    graph_id=graph_id,
    embedding_engine=embedding_engine
)

Configuration Methods

You can configure vector database providers through environment variables or programmatically at runtime.

Environment Variables

Create a .env file at your project root with the following variables:

VECTOR_DB_PROVIDER=lancedb
VECTOR_DB_URL=file:///tmp/cognee_lancedb
VECTOR_DB_KEY=your_optional_key

# For PGVector only

VECTOR_DB_USERNAME=postgres
VECTOR_DB_PASSWORD=secret
VECTOR_DB_HOST=localhost
VECTOR_DB_PORT=5432
VECTOR_DB_NAME=cognee_db

VectorConfig automatically loads these values when the library initializes. If VECTOR_DB_PROVIDER is missing, it defaults to lancedb.

Programmatic Configuration

To switch providers dynamically, use save_vector_db_config() from cognee/modules/settings/save_vector_db_config.py (lines 12-19):

from cognee.modules.settings.save_vector_db_config import VectorDBConfig, save_vector_db_config
from cognee.infrastructure.databases.vector.get_vector_engine import get_vector_engine

async def switch_to_pgvector():
    new_cfg = VectorDBConfig(
        url="",  # Not used for pgvector; connection details come from relational config

        api_key="",
        provider="pgvector",
    )
    await save_vector_db_config(new_cfg)
    
    # Retrieve the new engine instance

    vector_engine = await get_vector_engine()
    print(type(vector_engine))  # <class 'cognee.infrastructure.databases.vector.pgvector.PGVectorAdapter'>

Implementing Custom Vector Database Adapters

For unsupported databases, implement the VectorDBInterface and register your adapter using use_vector_adapter() from cognee/infrastructure/databases/vector/use_vector_adapter.py.

First, create your adapter class:


# my_custom_adapter.py

from cognee.infrastructure.databases.vector.vector_db_interface import VectorDBInterface

class CustomVectorAdapter(VectorDBInterface):
    def __init__(self, url: str, api_key: str, embedding_engine, **_):
        self.url = url
        self.api_key = api_key
        self.embedding_engine = embedding_engine
    
    async def add(self, collection_name, vectors, payloads):
        # Implementation here

        pass
    
    async def retrieve(self, collection_name, query_vector, limit=10):
        # Implementation here

        pass

Then register it at runtime:

from cognee.infrastructure.databases.vector.use_vector_adapter import use_vector_adapter
from my_custom_adapter import CustomVectorAdapter

use_vector_adapter("customdb", CustomVectorAdapter)

Now setting VECTOR_DB_PROVIDER=customdb resolves to your CustomVectorAdapter. The registry lives in supported_databases (cognee/infrastructure/databases/vector/supported_databases.py).

Accessing the Vector Engine in Application Code

Once configured, retrieve the vector engine anywhere in your application using the global accessor:

from cognee.infrastructure.databases.vector.get_vector_engine import get_vector_engine

async def upsert_documents(docs: list[dict]):
    vector_engine = await get_vector_engine()
    
    # vector_engine implements VectorDBInterface (add, retrieve, delete)

    await vector_engine.add(
        collection_name="my_docs",
        vectors=[doc["embedding"] for doc in docs],
        payloads=[{"text": doc["text"]} for doc in docs],
    )

The engine returned is a singleton instance cached by get_vector_engine(), ensuring consistent database connections across your application.

Summary

  • Configuration entry point: VectorConfig in cognee/infrastructure/databases/vector/config.py reads environment variables and caches settings.
  • Factory pattern: create_vector_engine() instantiates the correct adapter based on VECTOR_DB_PROVIDER.
  • Built-in providers: LanceDB (default), PGVector, ChromaDB, and Neptune Analytics each have specific connection requirements.
  • Custom extensions: Implement VectorDBInterface and register via use_vector_adapter() to support additional vector databases.
  • Global access: Use get_vector_engine() to obtain the configured engine instance throughout your codebase.

Frequently Asked Questions

How do I switch from LanceDB to PGVector in an existing Cognee project?

Set the VECTOR_DB_PROVIDER environment variable to pgvector and provide PostgreSQL credentials via VECTOR_DB_USERNAME, VECTOR_DB_PASSWORD, VECTOR_DB_HOST, VECTOR_DB_PORT, and VECTOR_DB_NAME. Alternatively, call save_vector_db_config() with provider="pgvector" programmatically. The next call to get_vector_engine() returns a PGVectorAdapter instance connected to your PostgreSQL database.

Can I use multiple vector databases simultaneously in the same application?

Cognee's current architecture uses a singleton pattern via get_vector_engine(), which returns a single global instance based on the active configuration. To use multiple databases simultaneously, you would need to instantiate adapter classes directly (such as LanceDBAdapter or PGVectorAdapter) rather than using the global accessor, passing the required connection parameters manually.

What happens if I don't configure any vector database provider?

If no provider is specified, Cognee defaults to LanceDB with a local file path under <system_root>/databases/cognee.lancedb. This default is set in VectorConfig and ensures the library works out-of-the-box without external dependencies, though it stores data locally on disk.

How do I add authentication for cloud-based vector databases like Neptune Analytics?

For Neptune Analytics, include the authentication token in the VECTOR_DB_URL or handle authentication within your network configuration, as the adapter extracts the graph ID from the endpoint URL (lines 66-82 in create_vector_engine.py). For other providers like ChromaDB or custom adapters, pass API keys via the VECTOR_DB_KEY environment variable or the api_key parameter in VectorDBConfig, which the factory passes directly to the adapter constructor.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →