How to Integrate LLM Reranking into the MemPalace Search Pipeline: A Complete Guide
Integrating LLM reranking into the MemPalace search pipeline involves injecting a semantic scoring layer after the hybrid vector+BM25 retrieval in mempalace/searcher.py, utilizing the LLMProvider abstraction from mempalace/llm_client.py to evaluate candidate relevance before final sorting.
MemPalace is a local-first, verbatim-preserving memory system that stores text as drawers (chunks) and organizes them via closets (topic pointers). The core retrieval pipeline combines ChromaDB vector similarity with BM25 lexical scoring, but you can significantly improve result quality by integrating LLM reranking to capture nuanced semantic relevance. This guide demonstrates exactly where to hook into the existing architecture and how to implement the reranking logic without compromising the system's privacy-first guarantees.
Understanding the MemPalace Retrieval Architecture
Before adding LLM reranking, you must understand the two-stage retrieval process implemented in mempalace/searcher.py.
The Hybrid Baseline
The baseline search performs hybrid ranking that merges dense and sparse signals. The search_memories function first retrieves candidates via vector similarity from ChromaDB, then the _hybrid_rank helper (lines 33-41) combines cosine similarity (1-distance) with BM25 scores computed on the candidate set. This hybrid approach balances embedding proximity with token overlap, but lacks deep semantic understanding of query intent and can miss conceptually relevant drawers that share few surface-level tokens with the query.
The Extension Point for Reranking
The candidate_strategy parameter in search_memories (lines 475-486) provides the architectural hook for LLM integration. When set to "union", the function merges vector candidates with lexical candidates from lexical_search, creating an enlarged pool before ranking. This same merge point is where you inject LLM-based relevance scoring, passing each candidate through an LLM judge to compute a semantic relevance signal that complements the existing vector and BM25 scores.
Architectural Flow with LLM Reranking
When LLM reranking is activated, the pipeline follows this sequence:
- Initial Retrieval: Vector search via
drawers_col.querygenerates candidate list A, while optionallexical_searchgenerates candidate list B. - Pool Merging: The candidate lists merge and deduplicate into pool P.
- Hybrid Scoring: The
_hybrid_rankfunction computes baseline scores combining vector similarity and BM25. - LLM Evaluation: Each candidate passes through
LLMProvider.classify()with a relevance prompting template. - Score Blending: The system combines vector, BM25, and LLM signals using configurable weights.
- Final Sort: The reordered results return to the caller or CLI.
The LLMProvider class (defined in mempalace/llm_client.py, lines 21-31) exposes a unified classify(system, user, json_mode=True) method that returns an LLMResponse object containing the generated text and raw JSON. This abstraction supports Ollama, OpenAI-compatible, and Anthropic APIs through a single interface.
Implementing LLM Reranking
Configure the LLMProvider
First, instantiate the provider using environment variables or explicit configuration. The system validates the endpoint URL and checks _endpoint_is_local (lines 35-44) to emit privacy warnings when connecting to external services.
from mempalace.llm_client import get_provider
import os
provider = get_provider(
name="ollama", # or "openai-compat" / "anthropic"
model="llama3:8b",
endpoint=os.getenv("LLM_ENDPOINT"),
)
Create the Scoring Prompt
Define a system prompt that instructs the LLM to evaluate relevance, and a user prompt containing the query and drawer text. The json_mode=True parameter ensures parseable output.
system_prompt = (
"You are a relevance judge. Return a JSON object with a single field "
"`score` between 0 and 1 indicating how well the drawer matches the query."
)
user_prompt = f"Query: {query}\n\nDrawer:\n{drawer_text}"
response = provider.classify(system_prompt, user_prompt, json_mode=True)
llm_score = float(response.text["score"])
Blend the Relevance Signals
Modify the ranking logic to combine the three signals. You can either replace the BM25 component entirely or use a weighted linear combination that preserves all three signals.
# Inside _hybrid_rank or a wrapper function
final_score = (
vector_weight * vec_sim +
bm25_weight * bm25_norm +
llm_weight * llm_score
)
Alternatively, replace BM25 entirely by setting its weight to zero and relying solely on vector plus LLM signals for applications where semantic nuance outweighs lexical precision.
Complete Code Examples
Minimal LLM-Augmented Search Function
This implementation wraps search_memories to add LLM scoring as a post-processing step, maintaining compatibility with the existing API.
from pathlib import Path
from mempalace.searcher import search_memories
from mempalace.llm_client import get_provider
import json
import os
def llm_enhanced_search(
query: str,
palace_path: str,
*,
llm_name: str = "ollama",
llm_model: str = "llama3:8b",
llm_endpoint: str = None,
llm_weight: float = 0.3,
vector_weight: float = 0.5,
bm25_weight: float = 0.2,
**search_kwargs
):
# 1️⃣ Build LLM provider
provider = get_provider(
name=llm_name,
model=llm_model,
endpoint=llm_endpoint
)
# 2️⃣ Run the normal search to get the candidate pool
result = search_memories(query, palace_path, **search_kwargs)
# 3️⃣ Score each hit with the LLM
for hit in result["results"]:
sys_prompt = (
"You are a relevance judge. Return a JSON object with a single field "
"`score` between 0 and 1 indicating how well the drawer matches the query."
)
user_prompt = f"Query: {query}\n\nDrawer:\n{hit['text']}"
llm_resp = provider.classify(sys_prompt, user_prompt, json_mode=True)
try:
llm_score = float(json.loads(llm_resp.text).get("score", 0.0))
except Exception:
llm_score = 0.0
hit["_llm_score"] = llm_score
# 4️⃣ Blend the scores (simple linear combination)
for hit in result["results"]:
vec_sim = hit.get("similarity", 0.0)
bm25 = hit.get("bm25_score", 0.0) / 10.0 # normalize roughly
llm = hit["_llm_score"]
hit["final_score"] = (
vector_weight * vec_sim +
bm25_weight * bm25 +
llm_weight * llm
)
# 5️⃣ Return hits sorted by the new composite score
result["results"] = sorted(
result["results"],
key=lambda h: h["final_score"],
reverse=True
)
return result
CLI Integration
Add a --llm-rerank flag to the command-line interface in mempalace/cli.py to activate the LLM scoring path without breaking existing workflows.
# In mempalace/cli.py (excerpt)
if args.llm_rerank:
from mempalace.llm_client import get_provider
llm = get_provider(
name=args.llm_provider,
model=args.llm_model,
endpoint=args.llm_endpoint
)
# Pass the provider into the search call
results = search_memories(
query=args.query,
palace_path=args.palace,
candidate_strategy=args.candidate_strategy,
llm_provider=llm,
llm_weight=args.llm_weight
)
else:
results = search_memories(...)
Extending search_memories with LLM Support
To integrate natively, modify the search_memories signature to accept optional LLM parameters and insert the scoring logic after hybrid ranking but before final sorting.
def search_memories(
query: str,
palace_path: str,
llm_provider=None,
llm_weight: float = 0.0,
# ... other existing params
):
# Existing retrieval and hybrid ranking...
hits = _hybrid_rank(candidates, query)
# After hybrid ranking but before final sort:
if llm_provider and llm_weight > 0:
for hit in hits:
resp = llm_provider.classify(
system_prompt,
user_prompt,
json_mode=True
)
llm_score = float(json.loads(resp.text).get("score", 0.0))
hit["_llm_score"] = llm_score
# Blend scores
for hit in hits:
hit["composite"] = (
hit["effective_distance"] * -1 * (1 - llm_weight) +
hit["_llm_score"] * llm_weight
)
hits.sort(key=lambda h: h["composite"], reverse=True)
return hits
Privacy and Safety Considerations
The LLM integration is opt-in by design. According to the source code in mempalace/llm_client.py (lines 35-44), the system validates that the endpoint URL uses HTTP/HTTPS and checks _endpoint_is_local to emit warnings when connecting to external services. Core operations remain local-first unless you explicitly configure LLM_ENDPOINT, LLM_MODEL, and optional LLM_KEY. When using local endpoints like Ollama, all text processing stays on your machine, preserving MemPalace's verbatim-preserving privacy guarantees.
Summary
- The
_hybrid_rankfunction inmempalace/searcher.py(lines 33-41) provides the baseline vector+BM25 scoring that LLM reranking augments. - The
LLMProviderclass inmempalace/llm_client.pyexposes aclassify()method supporting Ollama, OpenAI-compatible, and Anthropic APIs. - LLM reranking fits as a second-stage evaluator on the candidate pool after
candidate_strategymerges vector and lexical results. - Implementation requires constructing relevance prompts, parsing JSON responses, and tuning weights for vector, BM25, and LLM signal blending.
- The architecture preserves local-first operation when using local LLM endpoints, with explicit privacy warnings for external connections.
Frequently Asked Questions
Where does LLM reranking fit in the MemPalace search pipeline?
LLM reranking integrates as a post-processing step after the initial candidate retrieval and hybrid ranking. According to mempalace/searcher.py, the search_memories function retrieves candidates via vector similarity and lexical search, merges them using candidate_strategy, then applies _hybrid_rank. LLM scoring occurs immediately after this hybrid ranking, evaluating each drawer for semantic relevance before the final sort returns results to the caller.
Can I use cloud LLM providers like OpenAI or Anthropic with MemPalace?
Yes. The LLMProvider abstraction in mempalace/llm_client.py supports three provider types: "ollama" for local inference, "openai-compat" for OpenAI-compatible endpoints, and "anthropic" for Claude models. The system validates external endpoints using _endpoint_is_local (lines 35-44) and emits privacy warnings, but never transmits data unless you explicitly configure the LLM_ENDPOINT environment variable or pass the endpoint parameter.
How do I balance LLM scores with existing BM25 and vector signals?
Use a weighted linear combination in your ranking function. For example: final_score = (vector_weight * vec_sim) + (bm25_weight * bm25_norm) + (llm_weight * llm_score). Expose these weights—vector_weight, bm25_weight, and llm_weight—as CLI flags or configuration parameters to allow runtime tuning. You can also set bm25_weight to zero to replace lexical scoring entirely with LLM evaluation for tasks requiring deep semantic understanding over token overlap.
Will adding LLM reranking break existing local-first guarantees?
No. The LLM integration is opt-in and non-invasive. The core search_memories function in mempalace/searcher.py operates entirely locally unless you pass an llm_provider instance. When using local endpoints like Ollama (via LLM_ENDPOINT=http://localhost:11434), all data remains on your machine. The system only contacts external services when explicitly configured, maintaining MemPalace's design principle of verbatim, privacy-preserving storage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →