# How MemPalace Hybrid BM25 and Vector Search Ranking Works

> Discover how MemPalace hybrid ranking merges BM25 and vector search. Achieve balanced results with semantic awareness and keyword precision. Learn more about its configurable weights and normalization.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: deep-dive
- Published: 2026-06-07

---

**MemPalace combines Okapi-BM25 lexical scoring with cosine-similarity vector search using configurable weights (default 60% vector, 40% BM25) and min-max normalization to deliver results that balance semantic awareness with keyword precision.**

The MemPalace repository implements a sophisticated hybrid retrieval pipeline in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) that merges traditional inverted-index techniques with modern embedding-based similarity. This architecture ensures that queries match both exact terminology and conceptual meaning, addressing the limitations of pure lexical or semantic search approaches.

## Tokenization and BM25 Scoring

The lexical foundation begins with the `_tokenize` helper function, which extracts alphanumeric tokens of at least two characters from query and document text. In `_bm25_scores` (lines 74-94), the system calculates Okapi-BM25 scores using an **IDF derived exclusively from the current candidate set** rather than the global corpus. This design choice ensures that term discrimination reflects the specific retrieved context, making the BM25 component responsive to the actual result pool.

### Local IDF Calculation

By computing inverse document frequency only from candidates passed to `_bm25_scores`, MemPalace adapts lexical relevance to the subset of documents under consideration. This prevents common terms in the broader corpus from diluting scores when they appear frequently within a specific result set.

## Vector Similarity Processing

For the semantic component, MemPalace queries a **Chroma HNSW index** where each hit returns a cosine distance value. The system converts this to a similarity score using the formula `max(0.0, 1 - distance)`, mapping the raw distance range `[0, 2]` to a normalized similarity scale of `[1, 0]`. This transformation occurs in the hybrid ranking pipeline before combination with lexical scores.

## The Hybrid Re-Ranking Algorithm

The core blending logic resides in `_hybrid_rank` (lines 33-55). This function implements a weighted combination strategy that normalizes disparate score scales and merges signals into a unified ranking.

### Score Normalization and Weighting

Raw BM25 scores undergo **min-max normalization** to the `[0, 1]` range before combination. This ensures that dissimilar scales between BM25 (typically unbounded) and cosine similarity (0-1) do not bias the final ranking. The default configuration applies a `vector_weight` of 0.6 and `bm25_weight` of 0.4, though these parameters remain fully configurable. The final hybrid score computes as: `(normalized_bm25 * bm25_weight) + (vector_similarity * vector_weight)`.

### Handling Missing Embeddings

When using the `"union"` candidate strategy, some results may lack vector embeddings. These entries receive a vector component of 0, allowing the normalized BM25 score to dominate the final ranking for that specific document while preserving the hybrid scoring framework.

## Context-Aware Ranking with Closet Boosts

Beyond the base hybrid calculation, `search_memories` (lines 88-119) implements a secondary **"closet" collection** query that provides context-aware boosting. The best matching closet per source file contributes an ordinal boost (`CLOSET_RANK_BOOSTS`) that is subtracted from the raw cosine distance, producing an effective distance: `effective_dist = max(0, min(2, dist - boost))`.

This boost applies **before** final sorting, allowing documents with strong lexical matches and relevant closet pointers to rise higher in the result list. The subtraction reduces the effective distance, thereby increasing the final similarity score for closet-associated memories.

## Candidate Retrieval Strategies

The pipeline supports two distinct retrieval modes controlled by the `candidate_strategy` parameter, affecting which documents enter the hybrid re-ranking phase.

### Vector-Only Retrieval

The default `candidate_strategy="vector"` retrieves candidates exclusively from the Chroma HNSW index. This mode provides strict semantic similarity guarantees and fully respects distance thresholds configured in the query.

### Union Strategy for Lexical Expansion

Setting `candidate_strategy="union"` merges vector results with additional lexical hits from the backend's `lexical_search` capability. When a strict `max_distance` threshold is configured, the union mode automatically skips BM25-only candidates to preserve the distance guarantee for vector results. This strategy improves recall by ensuring keyword matches that might fall outside the vector index's approximate nearest-neighbor search still appear in final results.

## Practical Implementation

The following examples demonstrate how to invoke MemPalace's hybrid ranking from both the command line and Python:

```bash

# CLI search with hybrid ranking (default behavior)

$ mempalace search "how to prune a bonsai tree" --wing gardening --n-results 10

```

```python

# Programmatic hybrid search with vector and BM25 blending

from mempalace.searcher import search_memories

results = search_memories(
    query="how to prune a bonsai tree",
    palace_path="/path/to/palace",
    wing="gardening",
    n_results=10,
    candidate_strategy="union",   # Also pull lexical candidates

    vector_weight=0.6,
    bm25_weight=0.4,
)

for hit in results["results"]:
    print(f"Similarity: {hit['similarity']}, "
          f"BM25: {hit['bm25_score']}, "
          f"Text: {hit['text'][:200]}")

```

## Summary

- **MemPalace hybrid BM25 and vector search ranking** combines lexical precision with semantic understanding through a weighted blend of Okapi-BM25 and cosine similarity scores.
- The `_bm25_scores` function calculates IDF locally from candidate sets only, ensuring term relevance reflects the retrieved context.
- **Min-max normalization** standardizes BM25 scores to `[0, 1]` before combination with vector similarities using configurable weights (default 60/40 vector/BM25).
- The **closet boost** mechanism subtracts ordinal boosts from cosine distances, allowing contextually related documents to rank higher.
- The `"union"` **candidate strategy** expands recall by merging lexical search results with vector hits, while respecting distance thresholds when configured.

## Frequently Asked Questions

### How does MemPalace normalize BM25 and vector scores for combination?

MemPalace applies **min-max normalization** to raw BM25 scores within the `_hybrid_rank` function, compressing them to the `[0, 1]` range to match the vector similarity scale. Vector similarities derive from `max(0.0, 1 - distance)` calculations on Chroma HNSW cosine distances. These normalized values then blend using the configured `vector_weight` and `bm25_weight` parameters.

### What happens if a document has no vector embedding in hybrid search?

When the `"union"` strategy retrieves BM25-only candidates lacking embeddings, the system assigns them a vector similarity component of 0. The normalized BM25 score then dominates the final hybrid calculation, allowing purely lexical matches to compete in the ranked results based on keyword relevance alone.

### How does the closet boost affect final rankings?

The closet boost operates as an **ordinal distance reduction** applied before final sorting. When a document associates with a matching closet, the system calculates `effective_dist = max(0, min(2, dist - boost))`, potentially lowering the distance to 0 for strong closet matches. This subtraction increases the effective similarity score, pushing contextually relevant documents higher in the output list.

### When should I use the union candidate strategy versus vector-only?

Use **`candidate_strategy="union"`** when you need to capture exact keyword matches that might not surface in the approximate nearest-neighbor vector search, particularly for rare terminology or specific proper nouns. Use the default **`"vector"`** strategy when you require strict adherence to semantic distance thresholds and prefer higher precision over lexical recall.