Understanding metadata_store and vector_index Database Configurations in memU

The metadata_store configuration manages structured metadata persistence (categories, timestamps, tags) while vector_index handles vector embeddings for similarity search, with each supporting distinct backend providers and intelligent auto-defaulting logic.

In the NevaMind-AI/memU repository, the persistence layer is architecturally split between two distinct configuration domains: metadata_store and vector_index. Understanding the difference between these database configurations is essential for optimizing memory storage, retrieval performance, and deployment architecture in production environments.

The Architectural Split: Structured Data vs. Vector Embeddings

The fundamental distinction lies in what each system persists and how it is queried:

  • metadata_store: Stores structured metadata about memories—including categories, items, timestamps, tags, and relational attributes. This represents traditional tabular data requiring ACID compliance and structured querying capabilities.

  • vector_index: Stores vector embeddings used for similarity search and nearest-neighbor retrieval. This enables semantic search capabilities where memories are retrieved based on embedding distance rather than exact keyword matches.

Configuration Models and Provider Options

Both configurations are defined in src/memu/app/settings.py using distinct Pydantic models with different provider ecosystems.

MetadataStoreConfig Structure

The MetadataStoreConfig class (lines 299-303) controls structured persistence with three supported providers:

class MetadataStoreConfig(BaseModel):
    provider: Literal["inmemory", "postgres", "sqlite"] = "inmemory"
    dsn: Optional[str] = None
    ddl_mode: Literal["create", "validate"] = "create"

Supported backends:

  • "inmemory": Pure Python dictionary storage for testing and prototyping
  • "postgres": Production PostgreSQL backend with full SQL support
  • "sqlite": File-based SQLite for lightweight deployments

VectorIndexConfig Structure

The VectorIndexConfig class (lines 305-308) manages embedding storage with specialized vector search providers:

class VectorIndexConfig(BaseModel):
    provider: Literal["bruteforce", "pgvector", "none"] = "bruteforce"
    dsn: Optional[str] = None

Supported backends:

  • "bruteforce": In-memory brute-force cosine similarity (exact but computationally expensive for large datasets)
  • "pgvector": PostgreSQL extension for approximate nearest neighbor (ANN) search
  • "none": Disables vector similarity search entirely

Auto-Defaulting and Provider Selection Logic

A critical difference lies in how each configuration handles defaults. The DatabaseConfig.model_post_init method (lines 314-321) implements intelligent auto-configuration that couples the two systems:

When vector_index is omitted from the configuration, the system automatically selects the appropriate provider based on the metadata store choice:

  • If metadata_store.provider="postgres", the vector_index defaults to pgvector using the same DSN
  • For all other metadata providers (inmemory, sqlite), vector_index defaults to bruteforce

This coupling ensures that production deployments using PostgreSQL automatically gain vector capabilities without explicit configuration, while lightweight setups remain dependency-free.

Practical Configuration Examples

Production Deployment: PostgreSQL with pgvector

For production workloads requiring durable storage and fast semantic search:

from memu.app.settings import DatabaseConfig, MetadataStoreConfig

config = DatabaseConfig(
    metadata_store=MetadataStoreConfig(
        provider="postgres",
        dsn="postgresql://user:pwd@localhost/memu",
        ddl_mode="create",
    )
    # vector_index auto-defaults to pgvector with the same DSN

)

The factory in src/memu/database/factory.py (lines 28-43) instantiates PostgresStore, which receives both the metadata configuration and the vector provider setting from src/memu/database/postgres/__init__.py (line 20).

Lightweight Prototyping: In-Memory with Brute Force

For testing or development environments requiring zero external dependencies:

from memu.app.settings import DatabaseConfig, MetadataStoreConfig, VectorIndexConfig

config = DatabaseConfig(
    metadata_store=MetadataStoreConfig(provider="inmemory"),
    vector_index=VectorIndexConfig(provider="bruteforce")
)

This configuration requires no connection strings or external services, making it ideal for unit tests and CI pipelines.

Key Implementation Files

Understanding these configurations requires familiarity with specific source files in the NevaMind-AI/memU repository:

  • src/memu/app/settings.py: Defines MetadataStoreConfig (lines 299-303), VectorIndexConfig (lines 305-308), and the auto-defaulting logic in DatabaseConfig.model_post_init (lines 314-321).

  • src/memu/database/factory.py: Contains the factory method (lines 28-43) that reads config.metadata_store.provider to instantiate the appropriate Database implementation.

  • src/memu/database/postgres/__init__.py: Shows how the Postgres backend consumes both configurations, receiving the vector provider at line 20 to determine whether to initialize pgvector columns.

  • src/memu/database/sqlite/__init__.py: Demonstrates the SQLite backend implementation, which uses only the metadata store configuration while defaulting the vector index to bruteforce.

Summary

The distinction between metadata_store and vector_index configurations in memU reflects a clean architectural separation between structured data persistence and semantic search capabilities:

  • metadata_store handles ACID-compliant storage of memory attributes (categories, timestamps, tags) via MetadataStoreConfig, supporting PostgreSQL, SQLite, or in-memory backends.
  • vector_index manages embedding storage for similarity search via VectorIndexConfig, offering brute-force in-memory scanning, pgvector PostgreSQL extension, or disabled support.
  • Auto-defaulting logic couples the two configurations when vector_index is omitted, automatically selecting pgvector for PostgreSQL metadata stores and bruteforce for others.
  • Factory pattern in src/memu/database/factory.py instantiates the appropriate backend based on metadata_store.provider, while vector capabilities are injected via vector_index.provider.

Frequently Asked Questions

Can I use PostgreSQL for metadata but disable vector search entirely?

Yes. Set vector_index.provider="none" in your configuration. This stores all memory metadata in PostgreSQL while disabling embedding storage and similarity search, effectively removing semantic retrieval capabilities from your memU instance while maintaining structured data persistence.

Why does the vector index default to bruteforce when using SQLite?

The bruteforce provider performs exact cosine similarity calculations in memory without external dependencies. SQLite lacks native vector extension support comparable to pgvector, so memU defaults to the lightweight bruteforce implementation to maintain zero-dependency operation while still enabling semantic search functionality for smaller datasets.

Is the DSN required for both configurations when using PostgreSQL?

Only the metadata_store requires an explicit DSN. When vector_index.provider="pgvector" and the DSN is omitted, DatabaseConfig.model_post_init automatically inherits the DSN from metadata_store.dsn. This design ensures both structured data and vectors reside in the same database instance by default, simplifying connection management.

Can I mix different providers, such as SQLite for metadata and pgvector for vectors?

No. The pgvector provider requires a PostgreSQL backend because it relies on the pgvector PostgreSQL extension. If you configure metadata_store.provider="sqlite", you cannot use vector_index.provider="pgvector"; the system will either default to bruteforce or require you to explicitly set it. The vector index provider must be compatible with the underlying database technology when using SQL-backed storage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →