# How MemPalace Performs Semantic Search with ChromaDB Vectors: A Hybrid Retrieval Pipeline

> Discover how MemPalace achieves powerful semantic search using ChromaDB vectors. Explore its hybrid retrieval pipeline combining vector search and BM25 keyword scoring for superior results.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: deep-dive
- Published: 2026-06-06

---

**MemPalace implements semantic search with ChromaDB vectors through a hybrid retrieval pipeline that combines cosine similarity vector search with Okapi-BM25 keyword scoring, re-ranking candidates using a configurable weighted combination of both signals.**

MemPalace is an open-source memory management system that stores information in a structured "palace" architecture. The `mempalace.searcher` module orchestrates **semantic search with ChromaDB vectors** alongside traditional lexical retrieval, enabling robust information recall that balances semantic meaning with keyword precision. This implementation leverages ChromaDB's HNSW index configured for cosine similarity, while applying BM25 term relevance as a secondary ranking signal to refine results.

## Opening ChromaDB Collections and Building Filters

The search process begins by resolving the palace directory and establishing a connection to the vector store. The `_open_collection_or_explain` function in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) (lines 10-13, 26-31) initializes the ChromaDB collection that stores embeddings alongside document metadata.

Before executing the vector query, the system constructs optional metadata filters. The `build_where_filter` function (lines 80-88) creates a ChromaDB **where** clause from optional `wing` and `room` arguments, enabling scoped searches within specific project segments or temporal rooms. This filtering happens at the database level, reducing the candidate pool before vector comparison.

## Executing Vector Queries with Cosine Similarity

The core semantic retrieval uses ChromaDB's `query` method to perform nearest-neighbor search against the HNSW index. The implementation in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) (lines 22-33) executes:

```python
results = col.query(
    query_texts=[query],
    n_results=n_results,
    include=["documents", "metadatas", "distances"],
    where=where,
)

```

Because the palace initializes collections with `hnsw:space=cosine`, ChromaDB returns **cosine distances** rather than similarities. The system requests the raw documents, their metadata, and the distance values for downstream processing.

## Converting Distance to Similarity Scores

Raw cosine distances require transformation into comparable similarity scores. The searcher converts each distance `d` into an absolute cosine similarity using the formula `max(0.0, 1.0 - d)` as implemented in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) (lines 68-71). This normalization ensures that:
- Perfect matches (distance 0.0) yield similarity 1.0
- Orthogonal vectors (distance 1.0) yield similarity 0.0
- Negative similarities clamp to 0.0 to avoid counterintuitive scores

## Implementing BM25 Keyword Scoring

Parallel to vector retrieval, the system computes lexical relevance using the Okapi-BM25 algorithm. The `_bm25_scores` function (lines 74-105) tokenizes the query and retrieved documents, builds term-frequency statistics, and applies the BM25 formula to each candidate. This classical IR approach captures exact keyword matches and term frequency patterns that dense vector embeddings might overlook.

## Hybrid Re-ranking with Configurable Weights

The `_hybrid_rank` function (lines 123-147) normalizes BM25 scores and combines them with vector similarities using a convex combination. The default configuration applies:
- **60% weight to vector similarity** (semantic meaning)
- **40% weight to BM25 score** (keyword relevance)

This hybrid scoring allows the system to surface documents that match the query conceptually while boosting those containing specific terminology.

## Advanced Retrieval Strategies

### Union Candidate Strategy

When `candidate_strategy="union"` is specified, the searcher expands the candidate pool by pulling additional lexical matches via the backend's `lexical_search` method ([`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) lines 274-304). This strategy merges vector-derived candidates with BM25-only candidates before hybrid ranking, increasing recall for rare terms or domain-specific jargon that might not align well with learned embeddings.

### Closet-Based Boosting

The system implements a secondary "closet" collection containing regex-extracted pointers to related drawers. Matches from this collection apply a small, rank-based boost to the final distance calculation (lines 480-527), improving relevance for narrative content that spans multiple related memory chunks.

### Context Enrichment

For results boosted by closet matches, the module optionally expands hits with neighboring chunks (±1 chunk) to provide better context windows. This enrichment (lines 530-578) assembles coherent passages from fragmented storage units while preserving the underlying semantic relationships.

## Fallback to SQLite FTS5 Index

If the HNSW vector index becomes unavailable due to corruption or storage issues, the searcher falls back to a pure-BM25 mode. The `_bm25_only_via_sqlite` function (lines 88-105, 1075-1100) queries the SQLite-backed FTS5 index directly, bypassing vector retrieval entirely. This resilience ensures search functionality persists even when embedding infrastructure fails.

## Implementation Examples

### Basic CLI Search

```python
from mempalace.searcher import search

# Search the default palace located at "./my_palace"

search(
    query="how to backup my notes",
    palace_path="./my_palace",
    wing=None,          # optional project filter

    room=None,          # optional day/room filter

    n_results=5
)

```

This function prints formatted results including cosine similarity and BM25 scores (lines 162-176).

### Programmatic Search with Hybrid Options

```python
from mempalace.searcher import search_memories

result = search_memories(
    query="best practices for LLM prompting",
    palace_path="./my_palace",
    wing="proj_llm",
    room="2024-05-01",
    n_results=10,
    max_distance=0.4,               # only accept results with distance ≤ 0.4

    vector_disabled=False,
    candidate_strategy="union",     # also pull lexical BM25 candidates

)

# `result` is a dict:

# {

#   "query": "...",

#   "filters": {"wing": "...", "room": "..."},

#   "total_before_filter": 123,

#   "results": [

#       {

#           "text": "...",

#           "wing": "...",

#           "room": "...",

#           "source_file": "...",

#           "similarity": 0.81,

#           "distance": 0.19,

#           "bm25_score": 0.42,

#           "closet_boost": 0.0,

#           "matched_via": "drawer"

#       },

#       …

#   ]

# }

```

The `search_memories` function returns fully structured dictionaries including hybrid-scored results (lines 1570-1603).

### Forcing BM25-Only Fallback

```python
from mempalace.searcher import search_memories

fallback = search_memories(
    query="install mempalace on macOS",
    palace_path="./my_palace",
    vector_disabled=True,   # skip vector search entirely

    n_results=5,
)

print(fallback["results"][0]["bm25_score"])

```

Setting `vector_disabled=True` routes directly to `_bm25_only_via_sqlite`, using the FTS5 index for keyword-only retrieval (lines 1667-1700).

## Summary

- **MemPalace** implements **semantic search with ChromaDB vectors** using a hybrid architecture defined in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py).
- The pipeline executes **cosine similarity queries** against HNSW indexes, converts distances to similarities using `1.0 - d`, and combines these with **BM25 lexical scores**.
- Default weighting assigns **60% priority to vector similarity** and **40% to keyword relevance**, configurable via the `_hybrid_rank` function.
- **Advanced strategies** include union candidate retrieval, closet-based boosting, and context enrichment for narrative coherence.
- A **resilient fallback path** routes to SQLite FTS5 when vector indices are unavailable, ensuring continuous search functionality.
- The `search()` and `search_memories()` APIs expose these capabilities while handling collection management via [`mempalace/backends/chroma.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/chroma.py).

## Frequently Asked Questions

### How does MemPalace convert ChromaDB distances to similarity scores?

MemPalace converts ChromaDB's cosine distance values to absolute similarity scores using the formula `max(0.0, 1.0 - d)` in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) (lines 68-71). This clamps negative values to zero and ensures that smaller distances (closer vectors) produce higher similarity scores, with perfect matches yielding 1.0 and orthogonal vectors yielding 0.0.

### Can I disable semantic search and use only keyword matching?

Yes. Set `vector_disabled=True` when calling `search_memories()`. This routes the query to `_bm25_only_via_sqlite`, which queries the SQLite FTS5 index directly without accessing the ChromaDB vector store. This mode is useful when the HNSW index is corrupted or when you require exact keyword matching without semantic interpretation.

### What is the union candidate strategy in MemPalace search?

The union candidate strategy (`candidate_strategy="union"`) expands the result pool by combining vector-based candidates with additional lexical matches from BM25. The searcher pulls candidates via both `col.query()` and the backend's `lexical_search` method, merging them before applying hybrid re-ranking. This increases recall for specialized terminology that might not align well with dense embeddings.

### How are wing and room filters implemented in the vector search?

Wing and room filters translate into ChromaDB **where** clauses via the `build_where_filter` function (lines 80-88). These clauses filter the underlying metadata before vector similarity calculation, allowing scoped searches within specific project wings or temporal rooms without post-processing the results in Python.