Understanding metadata_store and vector_index Database Configurations in memU
The metadata_store configuration manages structured metadata persistence (categories, timestamps, tags) while vector_index handles vector embeddings for similarity search, with each supporting distinct backend providers and intelligent auto-defaulting logic.
In the NevaMind-AI/memU repository, the persistence layer is architecturally split between two distinct configuration domains: metadata_store and vector_index. Understanding the difference between these database configurations is essential for optimizing memory storage, retrieval performance, and deployment architecture in production environments.
The Architectural Split: Structured Data vs. Vector Embeddings
The fundamental distinction lies in what each system persists and how it is queried:
-
metadata_store: Stores structured metadata about memories—including categories, items, timestamps, tags, and relational attributes. This represents traditional tabular data requiring ACID compliance and structured querying capabilities. -
vector_index: Stores vector embeddings used for similarity search and nearest-neighbor retrieval. This enables semantic search capabilities where memories are retrieved based on embedding distance rather than exact keyword matches.
Configuration Models and Provider Options
Both configurations are defined in src/memu/app/settings.py using distinct Pydantic models with different provider ecosystems.
MetadataStoreConfig Structure
The MetadataStoreConfig class (lines 299-303) controls structured persistence with three supported providers:
class MetadataStoreConfig(BaseModel):
provider: Literal["inmemory", "postgres", "sqlite"] = "inmemory"
dsn: Optional[str] = None
ddl_mode: Literal["create", "validate"] = "create"
Supported backends:
"inmemory": Pure Python dictionary storage for testing and prototyping"postgres": Production PostgreSQL backend with full SQL support"sqlite": File-based SQLite for lightweight deployments
VectorIndexConfig Structure
The VectorIndexConfig class (lines 305-308) manages embedding storage with specialized vector search providers:
class VectorIndexConfig(BaseModel):
provider: Literal["bruteforce", "pgvector", "none"] = "bruteforce"
dsn: Optional[str] = None
Supported backends:
"bruteforce": In-memory brute-force cosine similarity (exact but computationally expensive for large datasets)"pgvector": PostgreSQL extension for approximate nearest neighbor (ANN) search"none": Disables vector similarity search entirely
Auto-Defaulting and Provider Selection Logic
A critical difference lies in how each configuration handles defaults. The DatabaseConfig.model_post_init method (lines 314-321) implements intelligent auto-configuration that couples the two systems:
When vector_index is omitted from the configuration, the system automatically selects the appropriate provider based on the metadata store choice:
- If
metadata_store.provider="postgres", thevector_indexdefaults topgvectorusing the same DSN - For all other metadata providers (
inmemory,sqlite),vector_indexdefaults tobruteforce
This coupling ensures that production deployments using PostgreSQL automatically gain vector capabilities without explicit configuration, while lightweight setups remain dependency-free.
Practical Configuration Examples
Production Deployment: PostgreSQL with pgvector
For production workloads requiring durable storage and fast semantic search:
from memu.app.settings import DatabaseConfig, MetadataStoreConfig
config = DatabaseConfig(
metadata_store=MetadataStoreConfig(
provider="postgres",
dsn="postgresql://user:pwd@localhost/memu",
ddl_mode="create",
)
# vector_index auto-defaults to pgvector with the same DSN
)
The factory in src/memu/database/factory.py (lines 28-43) instantiates PostgresStore, which receives both the metadata configuration and the vector provider setting from src/memu/database/postgres/__init__.py (line 20).
Lightweight Prototyping: In-Memory with Brute Force
For testing or development environments requiring zero external dependencies:
from memu.app.settings import DatabaseConfig, MetadataStoreConfig, VectorIndexConfig
config = DatabaseConfig(
metadata_store=MetadataStoreConfig(provider="inmemory"),
vector_index=VectorIndexConfig(provider="bruteforce")
)
This configuration requires no connection strings or external services, making it ideal for unit tests and CI pipelines.
Key Implementation Files
Understanding these configurations requires familiarity with specific source files in the NevaMind-AI/memU repository:
-
src/memu/app/settings.py: DefinesMetadataStoreConfig(lines 299-303),VectorIndexConfig(lines 305-308), and the auto-defaulting logic inDatabaseConfig.model_post_init(lines 314-321). -
src/memu/database/factory.py: Contains the factory method (lines 28-43) that readsconfig.metadata_store.providerto instantiate the appropriateDatabaseimplementation. -
src/memu/database/postgres/__init__.py: Shows how the Postgres backend consumes both configurations, receiving the vector provider at line 20 to determine whether to initialize pgvector columns. -
src/memu/database/sqlite/__init__.py: Demonstrates the SQLite backend implementation, which uses only the metadata store configuration while defaulting the vector index to bruteforce.
Summary
The distinction between metadata_store and vector_index configurations in memU reflects a clean architectural separation between structured data persistence and semantic search capabilities:
metadata_storehandles ACID-compliant storage of memory attributes (categories, timestamps, tags) viaMetadataStoreConfig, supporting PostgreSQL, SQLite, or in-memory backends.vector_indexmanages embedding storage for similarity search viaVectorIndexConfig, offering brute-force in-memory scanning, pgvector PostgreSQL extension, or disabled support.- Auto-defaulting logic couples the two configurations when
vector_indexis omitted, automatically selectingpgvectorfor PostgreSQL metadata stores andbruteforcefor others. - Factory pattern in
src/memu/database/factory.pyinstantiates the appropriate backend based onmetadata_store.provider, while vector capabilities are injected viavector_index.provider.
Frequently Asked Questions
Can I use PostgreSQL for metadata but disable vector search entirely?
Yes. Set vector_index.provider="none" in your configuration. This stores all memory metadata in PostgreSQL while disabling embedding storage and similarity search, effectively removing semantic retrieval capabilities from your memU instance while maintaining structured data persistence.
Why does the vector index default to bruteforce when using SQLite?
The bruteforce provider performs exact cosine similarity calculations in memory without external dependencies. SQLite lacks native vector extension support comparable to pgvector, so memU defaults to the lightweight bruteforce implementation to maintain zero-dependency operation while still enabling semantic search functionality for smaller datasets.
Is the DSN required for both configurations when using PostgreSQL?
Only the metadata_store requires an explicit DSN. When vector_index.provider="pgvector" and the DSN is omitted, DatabaseConfig.model_post_init automatically inherits the DSN from metadata_store.dsn. This design ensures both structured data and vectors reside in the same database instance by default, simplifying connection management.
Can I mix different providers, such as SQLite for metadata and pgvector for vectors?
No. The pgvector provider requires a PostgreSQL backend because it relies on the pgvector PostgreSQL extension. If you configure metadata_store.provider="sqlite", you cannot use vector_index.provider="pgvector"; the system will either default to bruteforce or require you to explicitly set it. The vector index provider must be compatible with the underlying database technology when using SQL-backed storage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →