# Understanding metadata_store and vector_index Database Configurations in memU

> Learn the difference between metadata_store and vector_index database configurations in memU. Understand how each handles metadata and vector embeddings for distinct use cases and backend providers.

- Repository: [NevaMind AI/memU](https://github.com/nevamind-ai/memu)
- Tags: api-reference
- Published: 2026-02-19

---

**The `metadata_store` configuration manages structured metadata persistence (categories, timestamps, tags) while `vector_index` handles vector embeddings for similarity search, with each supporting distinct backend providers and intelligent auto-defaulting logic.**

In the **NevaMind-AI/memU** repository, the persistence layer is architecturally split between two distinct configuration domains: `metadata_store` and `vector_index`. Understanding the difference between these database configurations is essential for optimizing memory storage, retrieval performance, and deployment architecture in production environments.

## The Architectural Split: Structured Data vs. Vector Embeddings

The fundamental distinction lies in what each system persists and how it is queried:

- **`metadata_store`**: Stores *structured* metadata about memories—including categories, items, timestamps, tags, and relational attributes. This represents traditional tabular data requiring ACID compliance and structured querying capabilities.

- **`vector_index`**: Stores *vector embeddings* used for similarity search and nearest-neighbor retrieval. This enables semantic search capabilities where memories are retrieved based on embedding distance rather than exact keyword matches.

## Configuration Models and Provider Options

Both configurations are defined in [`src/memu/app/settings.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/settings.py) using distinct Pydantic models with different provider ecosystems.

### MetadataStoreConfig Structure

The `MetadataStoreConfig` class (lines 299-303) controls structured persistence with three supported providers:

```python
class MetadataStoreConfig(BaseModel):
    provider: Literal["inmemory", "postgres", "sqlite"] = "inmemory"
    dsn: Optional[str] = None
    ddl_mode: Literal["create", "validate"] = "create"

```

**Supported backends:**
- `"inmemory"`: Pure Python dictionary storage for testing and prototyping
- `"postgres"`: Production PostgreSQL backend with full SQL support
- `"sqlite"`: File-based SQLite for lightweight deployments

### VectorIndexConfig Structure

The `VectorIndexConfig` class (lines 305-308) manages embedding storage with specialized vector search providers:

```python
class VectorIndexConfig(BaseModel):
    provider: Literal["bruteforce", "pgvector", "none"] = "bruteforce"
    dsn: Optional[str] = None

```

**Supported backends:**
- `"bruteforce"`: In-memory brute-force cosine similarity (exact but computationally expensive for large datasets)
- `"pgvector"`: PostgreSQL extension for approximate nearest neighbor (ANN) search
- `"none"`: Disables vector similarity search entirely

## Auto-Defaulting and Provider Selection Logic

A critical difference lies in how each configuration handles defaults. The `DatabaseConfig.model_post_init` method (lines 314-321) implements intelligent auto-configuration that couples the two systems:

When `vector_index` is omitted from the configuration, the system automatically selects the appropriate provider based on the metadata store choice:

- If `metadata_store.provider="postgres"`, the `vector_index` defaults to `pgvector` using the same DSN
- For all other metadata providers (`inmemory`, `sqlite`), `vector_index` defaults to `bruteforce`

This coupling ensures that production deployments using PostgreSQL automatically gain vector capabilities without explicit configuration, while lightweight setups remain dependency-free.

## Practical Configuration Examples

### Production Deployment: PostgreSQL with pgvector

For production workloads requiring durable storage and fast semantic search:

```python
from memu.app.settings import DatabaseConfig, MetadataStoreConfig

config = DatabaseConfig(
    metadata_store=MetadataStoreConfig(
        provider="postgres",
        dsn="postgresql://user:pwd@localhost/memu",
        ddl_mode="create",
    )
    # vector_index auto-defaults to pgvector with the same DSN

)

```

The factory in [`src/memu/database/factory.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/factory.py) (lines 28-43) instantiates `PostgresStore`, which receives both the metadata configuration and the vector provider setting from [`src/memu/database/postgres/__init__.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/postgres/__init__.py) (line 20).

### Lightweight Prototyping: In-Memory with Brute Force

For testing or development environments requiring zero external dependencies:

```python
from memu.app.settings import DatabaseConfig, MetadataStoreConfig, VectorIndexConfig

config = DatabaseConfig(
    metadata_store=MetadataStoreConfig(provider="inmemory"),
    vector_index=VectorIndexConfig(provider="bruteforce")
)

```

This configuration requires no connection strings or external services, making it ideal for unit tests and CI pipelines.

## Key Implementation Files

Understanding these configurations requires familiarity with specific source files in the NevaMind-AI/memU repository:

- **[`src/memu/app/settings.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/settings.py)**: Defines `MetadataStoreConfig` (lines 299-303), `VectorIndexConfig` (lines 305-308), and the auto-defaulting logic in `DatabaseConfig.model_post_init` (lines 314-321).

- **[`src/memu/database/factory.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/factory.py)**: Contains the factory method (lines 28-43) that reads `config.metadata_store.provider` to instantiate the appropriate `Database` implementation.

- **[`src/memu/database/postgres/__init__.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/postgres/__init__.py)**: Shows how the Postgres backend consumes both configurations, receiving the vector provider at line 20 to determine whether to initialize pgvector columns.

- **[`src/memu/database/sqlite/__init__.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/sqlite/__init__.py)**: Demonstrates the SQLite backend implementation, which uses only the metadata store configuration while defaulting the vector index to bruteforce.

## Summary

The distinction between `metadata_store` and `vector_index` configurations in memU reflects a clean architectural separation between structured data persistence and semantic search capabilities:

- **`metadata_store`** handles ACID-compliant storage of memory attributes (categories, timestamps, tags) via `MetadataStoreConfig`, supporting PostgreSQL, SQLite, or in-memory backends.
- **`vector_index`** manages embedding storage for similarity search via `VectorIndexConfig`, offering brute-force in-memory scanning, pgvector PostgreSQL extension, or disabled support.
- **Auto-defaulting logic** couples the two configurations when `vector_index` is omitted, automatically selecting `pgvector` for PostgreSQL metadata stores and `bruteforce` for others.
- **Factory pattern** in [`src/memu/database/factory.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/factory.py) instantiates the appropriate backend based on `metadata_store.provider`, while vector capabilities are injected via `vector_index.provider`.

## Frequently Asked Questions

### Can I use PostgreSQL for metadata but disable vector search entirely?

Yes. Set `vector_index.provider="none"` in your configuration. This stores all memory metadata in PostgreSQL while disabling embedding storage and similarity search, effectively removing semantic retrieval capabilities from your memU instance while maintaining structured data persistence.

### Why does the vector index default to bruteforce when using SQLite?

The bruteforce provider performs exact cosine similarity calculations in memory without external dependencies. SQLite lacks native vector extension support comparable to pgvector, so memU defaults to the lightweight bruteforce implementation to maintain zero-dependency operation while still enabling semantic search functionality for smaller datasets.

### Is the DSN required for both configurations when using PostgreSQL?

Only the `metadata_store` requires an explicit DSN. When `vector_index.provider="pgvector"` and the DSN is omitted, `DatabaseConfig.model_post_init` automatically inherits the DSN from `metadata_store.dsn`. This design ensures both structured data and vectors reside in the same database instance by default, simplifying connection management.

### Can I mix different providers, such as SQLite for metadata and pgvector for vectors?

No. The pgvector provider requires a PostgreSQL backend because it relies on the pgvector PostgreSQL extension. If you configure `metadata_store.provider="sqlite"`, you cannot use `vector_index.provider="pgvector"`; the system will either default to `bruteforce` or require you to explicitly set it. The vector index provider must be compatible with the underlying database technology when using SQL-backed storage.