How to Set Up PostgreSQL with pgvector for Semantic Search in memU
Configure memU with metadata_store.provider="postgres" and the framework automatically enables pgvector for vector storage, runs migrations to create the extension, and performs cosine-similarity search via PostgresMemoryItemRepo.
Setting up PostgreSQL with pgvector for semantic search in memU requires only a few configuration changes to leverage native vector operations. The NevaMind-AI/memU repository abstracts storage behind a Database interface that automatically detects PostgreSQL and enables pgvector support. By using the official pgvector Docker image and setting the correct DSN, you can store embeddings in a native VECTOR column and perform fast similarity searches without manual SQL.
Start PostgreSQL with the pgvector Extension
The fastest way to run a compatible database is using the official Docker image that ships with pgvector pre-installed.
docker run -d \
--name memu-postgres \
-e POSTGRES_USER=postgres \
-e POSTGRES_PASSWORD=postgres \
-e POSTGRES_DB=memu \
-p 5432:5432 \
pgvector/pgvector:pg16
This exposes PostgreSQL on port 5432 with the vector extension available. You must provide a DSN (Data Source Name) so memU can connect. Set the environment variable or pass it directly in your configuration:
export POSTGRES_DSN="postgresql+psycopg://postgres:postgres@127.0.0.1:5432/memu"
Configure memU for pgvector
memU uses DatabaseConfig to determine which backend to instantiate. When you set metadata_store.provider to "postgres", the framework automatically configures pgvector for vector storage.
In src/memu/app/settings.py, the DatabaseConfig.model_post_init method (lines 14-22) detects the postgres provider and auto-populates a VectorIndexConfig with provider="pgvector". This means you only need to specify the metadata store:
from memu.app import MemoryService
import os
service = MemoryService(
llm_profiles={"default": {"api_key": os.getenv("OPENAI_API_KEY")}},
database_config={
"metadata_store": {
"provider": "postgres",
"dsn": os.getenv("POSTGRES_DSN"),
"ddl_mode": "create", # runs migrations & creates pgvector extension
}
# vector_index defaults to pgvector automatically
},
retrieve_config={"method": "rag"},
)
The ddl_mode="create" flag is essential on first run. It triggers the migration system to initialize the schema and enable the pgvector extension.
How memU Implements pgvector Storage
Understanding the internal flow helps debug issues and optimize performance. The implementation spans several modules in src/memu/database/postgres/.
Factory and Store Initialization
The entry point is build_postgres_database in src/memu/database/postgres/__init__.py, which is called by the factory in src/memu/database/factory.py when provider="postgres" is detected. The PostgresStore class in src/memu/database/postgres/postgres.py handles the connection. On line 51, it checks vector_provider == "pgvector" and sets self._use_vector_type = True, registering the vector type with SQLAlchemy.
Migrations and Schema
The migration logic lives in src/memu/database/postgres/migration.py. When run_migrations executes (lines 44-60), it performs two critical actions:
- Enables the extension: Executes
CREATE EXTENSION IF NOT EXISTS vector(lines 46-50). - Creates tables: Calls
metadata.create_all(engine)(line 63) to build tables defined insrc/memu/database/postgres/schema.py.
The schema imports VECTOR from pgvector.sqlalchemy (line 21 of schema.py) and declares embedding columns using this type, allowing native vector storage.
Retrieval with Vector Similarity
During retrieval, PostgresMemoryItemRepo in src/memu/database/postgres/repositories/memory_item_repo.py constructs the query. If use_vector=True, it orders results by pgvector distance. The repository includes a fallback: if the extension is unavailable, it transparently switches to a local brute-force cosine similarity calculation (see the logic around line 290).
Complete Working Example
The following script demonstrates the full flow: connecting to PostgreSQL, memorizing a resource, and retrieving results using pgvector similarity.
import os
import asyncio
from memu.app import MemoryService
async def main() -> None:
# Connection string
dsn = os.getenv(
"POSTGRES_DSN",
"postgresql+psycopg://postgres:postgres@127.0.0.1:5432/memu",
)
# Initialize service with postgres + pgvector
service = MemoryService(
llm_profiles={"default": {"api_key": os.getenv("OPENAI_API_KEY")}},
database_config={
"metadata_store": {
"provider": "postgres",
"dsn": dsn,
"ddl_mode": "create",
}
},
retrieve_config={"method": "rag"},
)
# Store data
await service.memorize(
resource_url="tests/example/example_conversation.json",
modality="conversation",
user={"user_id": "123"},
)
# Retrieve using semantic search
result = await service.retrieve(
queries=[{"role": "user", "content": {"text": "Tell me about preferences"}}],
where={"user_id": "123"},
)
print("Top results via pgvector:")
for item in result["items"][:3]:
print(f"- [{item['memory_type']}] {item['summary'][:120]}...")
if __name__ == "__main__":
asyncio.run(main())
Manual Migration (Optional)
If you need to run migrations outside the service lifecycle, call the helper directly:
from memu.database.postgres.migration import run_migrations, DDLMode
dsn = "postgresql+psycopg://postgres:postgres@localhost:5432/memu"
run_migrations(dsn=dsn, scope_model=object, ddl_mode=DDLMode("create"))
Summary
- Use the official image: Run
pgvector/pgvector:pg16via Docker to get a pre-configured PostgreSQL instance with the extension installed. - Set the provider: Configure
metadata_store.provider="postgres"inDatabaseConfig; memU automatically setsvector_index.provider="pgvector"viamodel_post_initinsrc/memu/app/settings.py. - Run migrations: Use
ddl_mode="create"to triggerrun_migrationsinsrc/memu/database/postgres/migration.py, which executesCREATE EXTENSION IF NOT EXISTS vectorand builds the schema. - Store and search: The
PostgresStoreregisters the vector type, whilePostgresMemoryItemRepoperforms cosine-similarity ordering; if pgvector is unavailable, the system falls back to brute-force search.
Frequently Asked Questions
What Docker image should I use for pgvector with memU?
Use the official pgvector/pgvector:pg16 image. According to the memU source code and README, this image ships with the pgvector extension pre-installed, eliminating the need to manually compile or enable extensions inside the container.
How does memU automatically configure pgvector?
When you set metadata_store.provider="postgres", the DatabaseConfig.model_post_init method in src/memu/app/settings.py (lines 14-22) automatically instantiates a VectorIndexConfig with provider="pgvector". The PostgresStore in src/memu/database/postgres/postgres.py then sets self._use_vector_type = True on line 51, activating pgvector-specific handling for embeddings.
What happens if the pgvector extension is not available on the database?
The PostgresMemoryItemRepo in src/memu/database/postgres/repositories/memory_item_repo.py includes a fallback mechanism (around line 290). If the extension is missing or use_vector is disabled, the repository performs a brute-force cosine similarity calculation in Python rather than using the native PostgreSQL vector operators.
Can I use pgvector with an existing PostgreSQL instance?
Yes, as long as the pgvector extension is installed on your existing server. Provide the DSN to your instance in the configuration, ensure the vector extension is available, and set ddl_mode="create" on the first run to allow memU to run migrations and create the necessary tables with VECTOR columns defined in src/memu/database/postgres/schema.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →