# How to Integrate LLM Reranking into the MemPalace Search Pipeline: A Complete Guide

> Learn to integrate LLM reranking into your MemPalace search pipeline. Enhance relevance scoring with a semantic layer for superior search results. Follow our complete guide.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: how-to-guide
- Published: 2026-06-07

---

**Integrating LLM reranking into the MemPalace search pipeline involves injecting a semantic scoring layer after the hybrid vector+BM25 retrieval in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py), utilizing the `LLMProvider` abstraction from [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py) to evaluate candidate relevance before final sorting.**

MemPalace is a local-first, verbatim-preserving memory system that stores text as **drawers** (chunks) and organizes them via **closets** (topic pointers). The core retrieval pipeline combines ChromaDB vector similarity with BM25 lexical scoring, but you can significantly improve result quality by **integrating LLM reranking** to capture nuanced semantic relevance. This guide demonstrates exactly where to hook into the existing architecture and how to implement the reranking logic without compromising the system's privacy-first guarantees.

## Understanding the MemPalace Retrieval Architecture

Before adding LLM reranking, you must understand the two-stage retrieval process implemented in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py).

### The Hybrid Baseline

The baseline search performs hybrid ranking that merges dense and sparse signals. The `search_memories` function first retrieves candidates via vector similarity from ChromaDB, then the `_hybrid_rank` helper (lines 33-41) combines cosine similarity (`1-distance`) with BM25 scores computed on the candidate set. This hybrid approach balances embedding proximity with token overlap, but lacks deep semantic understanding of query intent and can miss conceptually relevant drawers that share few surface-level tokens with the query.

### The Extension Point for Reranking

The `candidate_strategy` parameter in `search_memories` (lines 475-486) provides the architectural hook for LLM integration. When set to `"union"`, the function merges vector candidates with lexical candidates from `lexical_search`, creating an enlarged pool before ranking. This same merge point is where you inject LLM-based relevance scoring, passing each candidate through an LLM judge to compute a semantic relevance signal that complements the existing vector and BM25 scores.

## Architectural Flow with LLM Reranking

When LLM reranking is activated, the pipeline follows this sequence:

1. **Initial Retrieval**: Vector search via `drawers_col.query` generates candidate list A, while optional `lexical_search` generates candidate list B.
2. **Pool Merging**: The candidate lists merge and deduplicate into pool P.
3. **Hybrid Scoring**: The `_hybrid_rank` function computes baseline scores combining vector similarity and BM25.
4. **LLM Evaluation**: Each candidate passes through `LLMProvider.classify()` with a relevance prompting template.
5. **Score Blending**: The system combines vector, BM25, and LLM signals using configurable weights.
6. **Final Sort**: The reordered results return to the caller or CLI.

The `LLMProvider` class (defined in [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py), lines 21-31) exposes a unified `classify(system, user, json_mode=True)` method that returns an `LLMResponse` object containing the generated text and raw JSON. This abstraction supports Ollama, OpenAI-compatible, and Anthropic APIs through a single interface.

## Implementing LLM Reranking

### Configure the LLMProvider

First, instantiate the provider using environment variables or explicit configuration. The system validates the endpoint URL and checks `_endpoint_is_local` (lines 35-44) to emit privacy warnings when connecting to external services.

```python
from mempalace.llm_client import get_provider
import os

provider = get_provider(
    name="ollama",          # or "openai-compat" / "anthropic"

    model="llama3:8b",
    endpoint=os.getenv("LLM_ENDPOINT"),
)

```

### Create the Scoring Prompt

Define a system prompt that instructs the LLM to evaluate relevance, and a user prompt containing the query and drawer text. The `json_mode=True` parameter ensures parseable output.

```python
system_prompt = (
    "You are a relevance judge. Return a JSON object with a single field "
    "`score` between 0 and 1 indicating how well the drawer matches the query."
)

user_prompt = f"Query: {query}\n\nDrawer:\n{drawer_text}"
response = provider.classify(system_prompt, user_prompt, json_mode=True)
llm_score = float(response.text["score"])

```

### Blend the Relevance Signals

Modify the ranking logic to combine the three signals. You can either replace the BM25 component entirely or use a weighted linear combination that preserves all three signals.

```python

# Inside _hybrid_rank or a wrapper function

final_score = (
    vector_weight * vec_sim + 
    bm25_weight * bm25_norm + 
    llm_weight * llm_score
)

```

Alternatively, replace BM25 entirely by setting its weight to zero and relying solely on vector plus LLM signals for applications where semantic nuance outweighs lexical precision.

## Complete Code Examples

### Minimal LLM-Augmented Search Function

This implementation wraps `search_memories` to add LLM scoring as a post-processing step, maintaining compatibility with the existing API.

```python
from pathlib import Path
from mempalace.searcher import search_memories
from mempalace.llm_client import get_provider
import json
import os

def llm_enhanced_search(
    query: str, 
    palace_path: str, 
    *,
    llm_name: str = "ollama",
    llm_model: str = "llama3:8b",
    llm_endpoint: str = None,
    llm_weight: float = 0.3,
    vector_weight: float = 0.5,
    bm25_weight: float = 0.2,
    **search_kwargs
):
    # 1️⃣ Build LLM provider

    provider = get_provider(
        name=llm_name, 
        model=llm_model, 
        endpoint=llm_endpoint
    )

    # 2️⃣ Run the normal search to get the candidate pool

    result = search_memories(query, palace_path, **search_kwargs)

    # 3️⃣ Score each hit with the LLM

    for hit in result["results"]:
        sys_prompt = (
            "You are a relevance judge. Return a JSON object with a single field "
            "`score` between 0 and 1 indicating how well the drawer matches the query."
        )
        user_prompt = f"Query: {query}\n\nDrawer:\n{hit['text']}"
        llm_resp = provider.classify(sys_prompt, user_prompt, json_mode=True)
        
        try:
            llm_score = float(json.loads(llm_resp.text).get("score", 0.0))
        except Exception:
            llm_score = 0.0
        hit["_llm_score"] = llm_score

    # 4️⃣ Blend the scores (simple linear combination)

    for hit in result["results"]:
        vec_sim = hit.get("similarity", 0.0)
        bm25 = hit.get("bm25_score", 0.0) / 10.0   # normalize roughly

        llm = hit["_llm_score"]
        hit["final_score"] = (
            vector_weight * vec_sim +
            bm25_weight * bm25 +
            llm_weight * llm
        )

    # 5️⃣ Return hits sorted by the new composite score

    result["results"] = sorted(
        result["results"],
        key=lambda h: h["final_score"],
        reverse=True
    )
    return result

```

### CLI Integration

Add a `--llm-rerank` flag to the command-line interface in [`mempalace/cli.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/cli.py) to activate the LLM scoring path without breaking existing workflows.

```python

# In mempalace/cli.py (excerpt)

if args.llm_rerank:
    from mempalace.llm_client import get_provider
    
    llm = get_provider(
        name=args.llm_provider,
        model=args.llm_model,
        endpoint=args.llm_endpoint
    )
    
    # Pass the provider into the search call

    results = search_memories(
        query=args.query,
        palace_path=args.palace,
        candidate_strategy=args.candidate_strategy,
        llm_provider=llm,
        llm_weight=args.llm_weight
    )
else:
    results = search_memories(...)

```

### Extending search_memories with LLM Support

To integrate natively, modify the `search_memories` signature to accept optional LLM parameters and insert the scoring logic after hybrid ranking but before final sorting.

```python
def search_memories(
    query: str,
    palace_path: str,
    llm_provider=None,
    llm_weight: float = 0.0,
    # ... other existing params

):
    # Existing retrieval and hybrid ranking...

    hits = _hybrid_rank(candidates, query)
    
    # After hybrid ranking but before final sort:

    if llm_provider and llm_weight > 0:
        for hit in hits:
            resp = llm_provider.classify(
                system_prompt, 
                user_prompt, 
                json_mode=True
            )
            llm_score = float(json.loads(resp.text).get("score", 0.0))
            hit["_llm_score"] = llm_score
        
        # Blend scores

        for hit in hits:
            hit["composite"] = (
                hit["effective_distance"] * -1 * (1 - llm_weight) +
                hit["_llm_score"] * llm_weight
            )
        hits.sort(key=lambda h: h["composite"], reverse=True)
    
    return hits

```

## Privacy and Safety Considerations

The LLM integration is **opt-in** by design. According to the source code in [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py) (lines 35-44), the system validates that the endpoint URL uses HTTP/HTTPS and checks `_endpoint_is_local` to emit warnings when connecting to external services. Core operations remain local-first unless you explicitly configure `LLM_ENDPOINT`, `LLM_MODEL`, and optional `LLM_KEY`. When using local endpoints like Ollama, all text processing stays on your machine, preserving MemPalace's verbatim-preserving privacy guarantees.

## Summary

- The `_hybrid_rank` function in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) (lines 33-41) provides the baseline vector+BM25 scoring that LLM reranking augments.
- The `LLMProvider` class in [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py) exposes a `classify()` method supporting Ollama, OpenAI-compatible, and Anthropic APIs.
- LLM reranking fits as a second-stage evaluator on the candidate pool after `candidate_strategy` merges vector and lexical results.
- Implementation requires constructing relevance prompts, parsing JSON responses, and tuning weights for vector, BM25, and LLM signal blending.
- The architecture preserves local-first operation when using local LLM endpoints, with explicit privacy warnings for external connections.

## Frequently Asked Questions

### Where does LLM reranking fit in the MemPalace search pipeline?

LLM reranking integrates as a **post-processing step** after the initial candidate retrieval and hybrid ranking. According to [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py), the `search_memories` function retrieves candidates via vector similarity and lexical search, merges them using `candidate_strategy`, then applies `_hybrid_rank`. LLM scoring occurs immediately after this hybrid ranking, evaluating each drawer for semantic relevance before the final sort returns results to the caller.

### Can I use cloud LLM providers like OpenAI or Anthropic with MemPalace?

Yes. The `LLMProvider` abstraction in [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py) supports three provider types: `"ollama"` for local inference, `"openai-compat"` for OpenAI-compatible endpoints, and `"anthropic"` for Claude models. The system validates external endpoints using `_endpoint_is_local` (lines 35-44) and emits privacy warnings, but never transmits data unless you explicitly configure the `LLM_ENDPOINT` environment variable or pass the endpoint parameter.

### How do I balance LLM scores with existing BM25 and vector signals?

Use a **weighted linear combination** in your ranking function. For example: `final_score = (vector_weight * vec_sim) + (bm25_weight * bm25_norm) + (llm_weight * llm_score)`. Expose these weights—`vector_weight`, `bm25_weight`, and `llm_weight`—as CLI flags or configuration parameters to allow runtime tuning. You can also set `bm25_weight` to zero to replace lexical scoring entirely with LLM evaluation for tasks requiring deep semantic understanding over token overlap.

### Will adding LLM reranking break existing local-first guarantees?

No. The LLM integration is **opt-in and non-invasive**. The core `search_memories` function in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) operates entirely locally unless you pass an `llm_provider` instance. When using local endpoints like Ollama (via `LLM_ENDPOINT=http://localhost:11434`), all data remains on your machine. The system only contacts external services when explicitly configured, maintaining MemPalace's design principle of verbatim, privacy-preserving storage.