How to Integrate LLM Reranking into the MemPalace Search Pipeline: A Complete Guide

Integrating LLM reranking into the MemPalace search pipeline involves injecting a semantic scoring layer after the hybrid vector+BM25 retrieval in mempalace/searcher.py, utilizing the LLMProvider abstraction from mempalace/llm_client.py to evaluate candidate relevance before final sorting.

MemPalace is a local-first, verbatim-preserving memory system that stores text as drawers (chunks) and organizes them via closets (topic pointers). The core retrieval pipeline combines ChromaDB vector similarity with BM25 lexical scoring, but you can significantly improve result quality by integrating LLM reranking to capture nuanced semantic relevance. This guide demonstrates exactly where to hook into the existing architecture and how to implement the reranking logic without compromising the system's privacy-first guarantees.

Understanding the MemPalace Retrieval Architecture

Before adding LLM reranking, you must understand the two-stage retrieval process implemented in mempalace/searcher.py.

The Hybrid Baseline

The baseline search performs hybrid ranking that merges dense and sparse signals. The search_memories function first retrieves candidates via vector similarity from ChromaDB, then the _hybrid_rank helper (lines 33-41) combines cosine similarity (1-distance) with BM25 scores computed on the candidate set. This hybrid approach balances embedding proximity with token overlap, but lacks deep semantic understanding of query intent and can miss conceptually relevant drawers that share few surface-level tokens with the query.

The Extension Point for Reranking

The candidate_strategy parameter in search_memories (lines 475-486) provides the architectural hook for LLM integration. When set to "union", the function merges vector candidates with lexical candidates from lexical_search, creating an enlarged pool before ranking. This same merge point is where you inject LLM-based relevance scoring, passing each candidate through an LLM judge to compute a semantic relevance signal that complements the existing vector and BM25 scores.

Architectural Flow with LLM Reranking

When LLM reranking is activated, the pipeline follows this sequence:

  1. Initial Retrieval: Vector search via drawers_col.query generates candidate list A, while optional lexical_search generates candidate list B.
  2. Pool Merging: The candidate lists merge and deduplicate into pool P.
  3. Hybrid Scoring: The _hybrid_rank function computes baseline scores combining vector similarity and BM25.
  4. LLM Evaluation: Each candidate passes through LLMProvider.classify() with a relevance prompting template.
  5. Score Blending: The system combines vector, BM25, and LLM signals using configurable weights.
  6. Final Sort: The reordered results return to the caller or CLI.

The LLMProvider class (defined in mempalace/llm_client.py, lines 21-31) exposes a unified classify(system, user, json_mode=True) method that returns an LLMResponse object containing the generated text and raw JSON. This abstraction supports Ollama, OpenAI-compatible, and Anthropic APIs through a single interface.

Implementing LLM Reranking

Configure the LLMProvider

First, instantiate the provider using environment variables or explicit configuration. The system validates the endpoint URL and checks _endpoint_is_local (lines 35-44) to emit privacy warnings when connecting to external services.

from mempalace.llm_client import get_provider
import os

provider = get_provider(
    name="ollama",          # or "openai-compat" / "anthropic"

    model="llama3:8b",
    endpoint=os.getenv("LLM_ENDPOINT"),
)

Create the Scoring Prompt

Define a system prompt that instructs the LLM to evaluate relevance, and a user prompt containing the query and drawer text. The json_mode=True parameter ensures parseable output.

system_prompt = (
    "You are a relevance judge. Return a JSON object with a single field "
    "`score` between 0 and 1 indicating how well the drawer matches the query."
)

user_prompt = f"Query: {query}\n\nDrawer:\n{drawer_text}"
response = provider.classify(system_prompt, user_prompt, json_mode=True)
llm_score = float(response.text["score"])

Blend the Relevance Signals

Modify the ranking logic to combine the three signals. You can either replace the BM25 component entirely or use a weighted linear combination that preserves all three signals.


# Inside _hybrid_rank or a wrapper function

final_score = (
    vector_weight * vec_sim + 
    bm25_weight * bm25_norm + 
    llm_weight * llm_score
)

Alternatively, replace BM25 entirely by setting its weight to zero and relying solely on vector plus LLM signals for applications where semantic nuance outweighs lexical precision.

Complete Code Examples

Minimal LLM-Augmented Search Function

This implementation wraps search_memories to add LLM scoring as a post-processing step, maintaining compatibility with the existing API.

from pathlib import Path
from mempalace.searcher import search_memories
from mempalace.llm_client import get_provider
import json
import os

def llm_enhanced_search(
    query: str, 
    palace_path: str, 
    *,
    llm_name: str = "ollama",
    llm_model: str = "llama3:8b",
    llm_endpoint: str = None,
    llm_weight: float = 0.3,
    vector_weight: float = 0.5,
    bm25_weight: float = 0.2,
    **search_kwargs
):
    # 1️⃣ Build LLM provider

    provider = get_provider(
        name=llm_name, 
        model=llm_model, 
        endpoint=llm_endpoint
    )

    # 2️⃣ Run the normal search to get the candidate pool

    result = search_memories(query, palace_path, **search_kwargs)

    # 3️⃣ Score each hit with the LLM

    for hit in result["results"]:
        sys_prompt = (
            "You are a relevance judge. Return a JSON object with a single field "
            "`score` between 0 and 1 indicating how well the drawer matches the query."
        )
        user_prompt = f"Query: {query}\n\nDrawer:\n{hit['text']}"
        llm_resp = provider.classify(sys_prompt, user_prompt, json_mode=True)
        
        try:
            llm_score = float(json.loads(llm_resp.text).get("score", 0.0))
        except Exception:
            llm_score = 0.0
        hit["_llm_score"] = llm_score

    # 4️⃣ Blend the scores (simple linear combination)

    for hit in result["results"]:
        vec_sim = hit.get("similarity", 0.0)
        bm25 = hit.get("bm25_score", 0.0) / 10.0   # normalize roughly

        llm = hit["_llm_score"]
        hit["final_score"] = (
            vector_weight * vec_sim +
            bm25_weight * bm25 +
            llm_weight * llm
        )

    # 5️⃣ Return hits sorted by the new composite score

    result["results"] = sorted(
        result["results"],
        key=lambda h: h["final_score"],
        reverse=True
    )
    return result

CLI Integration

Add a --llm-rerank flag to the command-line interface in mempalace/cli.py to activate the LLM scoring path without breaking existing workflows.


# In mempalace/cli.py (excerpt)

if args.llm_rerank:
    from mempalace.llm_client import get_provider
    
    llm = get_provider(
        name=args.llm_provider,
        model=args.llm_model,
        endpoint=args.llm_endpoint
    )
    
    # Pass the provider into the search call

    results = search_memories(
        query=args.query,
        palace_path=args.palace,
        candidate_strategy=args.candidate_strategy,
        llm_provider=llm,
        llm_weight=args.llm_weight
    )
else:
    results = search_memories(...)

Extending search_memories with LLM Support

To integrate natively, modify the search_memories signature to accept optional LLM parameters and insert the scoring logic after hybrid ranking but before final sorting.

def search_memories(
    query: str,
    palace_path: str,
    llm_provider=None,
    llm_weight: float = 0.0,
    # ... other existing params

):
    # Existing retrieval and hybrid ranking...

    hits = _hybrid_rank(candidates, query)
    
    # After hybrid ranking but before final sort:

    if llm_provider and llm_weight > 0:
        for hit in hits:
            resp = llm_provider.classify(
                system_prompt, 
                user_prompt, 
                json_mode=True
            )
            llm_score = float(json.loads(resp.text).get("score", 0.0))
            hit["_llm_score"] = llm_score
        
        # Blend scores

        for hit in hits:
            hit["composite"] = (
                hit["effective_distance"] * -1 * (1 - llm_weight) +
                hit["_llm_score"] * llm_weight
            )
        hits.sort(key=lambda h: h["composite"], reverse=True)
    
    return hits

Privacy and Safety Considerations

The LLM integration is opt-in by design. According to the source code in mempalace/llm_client.py (lines 35-44), the system validates that the endpoint URL uses HTTP/HTTPS and checks _endpoint_is_local to emit warnings when connecting to external services. Core operations remain local-first unless you explicitly configure LLM_ENDPOINT, LLM_MODEL, and optional LLM_KEY. When using local endpoints like Ollama, all text processing stays on your machine, preserving MemPalace's verbatim-preserving privacy guarantees.

Summary

  • The _hybrid_rank function in mempalace/searcher.py (lines 33-41) provides the baseline vector+BM25 scoring that LLM reranking augments.
  • The LLMProvider class in mempalace/llm_client.py exposes a classify() method supporting Ollama, OpenAI-compatible, and Anthropic APIs.
  • LLM reranking fits as a second-stage evaluator on the candidate pool after candidate_strategy merges vector and lexical results.
  • Implementation requires constructing relevance prompts, parsing JSON responses, and tuning weights for vector, BM25, and LLM signal blending.
  • The architecture preserves local-first operation when using local LLM endpoints, with explicit privacy warnings for external connections.

Frequently Asked Questions

Where does LLM reranking fit in the MemPalace search pipeline?

LLM reranking integrates as a post-processing step after the initial candidate retrieval and hybrid ranking. According to mempalace/searcher.py, the search_memories function retrieves candidates via vector similarity and lexical search, merges them using candidate_strategy, then applies _hybrid_rank. LLM scoring occurs immediately after this hybrid ranking, evaluating each drawer for semantic relevance before the final sort returns results to the caller.

Can I use cloud LLM providers like OpenAI or Anthropic with MemPalace?

Yes. The LLMProvider abstraction in mempalace/llm_client.py supports three provider types: "ollama" for local inference, "openai-compat" for OpenAI-compatible endpoints, and "anthropic" for Claude models. The system validates external endpoints using _endpoint_is_local (lines 35-44) and emits privacy warnings, but never transmits data unless you explicitly configure the LLM_ENDPOINT environment variable or pass the endpoint parameter.

How do I balance LLM scores with existing BM25 and vector signals?

Use a weighted linear combination in your ranking function. For example: final_score = (vector_weight * vec_sim) + (bm25_weight * bm25_norm) + (llm_weight * llm_score). Expose these weights—vector_weight, bm25_weight, and llm_weight—as CLI flags or configuration parameters to allow runtime tuning. You can also set bm25_weight to zero to replace lexical scoring entirely with LLM evaluation for tasks requiring deep semantic understanding over token overlap.

Will adding LLM reranking break existing local-first guarantees?

No. The LLM integration is opt-in and non-invasive. The core search_memories function in mempalace/searcher.py operates entirely locally unless you pass an llm_provider instance. When using local endpoints like Ollama (via LLM_ENDPOINT=http://localhost:11434), all data remains on your machine. The system only contacts external services when explicitly configured, maintaining MemPalace's design principle of verbatim, privacy-preserving storage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →