Preventing Prompt Contamination with MemPalace query_sanitizer: A 4-Step Recovery Strategy

The sanitize_query function in mempalace/query_sanitizer.py automatically extracts genuine user questions from LLM-contaminated prompts using a four-step fallback strategy that restores retrieval accuracy from 1% to nearly 90% without breaking MemPalace's verbatim storage guarantees.

MemPalace stores every user utterance verbatim to ensure perfect recall, but LLM-driven agents often prefix search requests with lengthy system instructions that poison vector embeddings. The query sanitizer in the MemPalace/mempalace repository solves this by implementing a robust extraction pipeline that isolates the true query intent before indexing.

The Prompt Contamination Problem

When an LLM-driven agent generates a search request for MemPalace, it frequently prefixes the real question with a massive system prompt. This contamination forces the embedder to process a thousand-character string where the actual query comprises only a few dozen characters, causing retrieval performance to collapse to approximately 1% R@10 (Recall@10).

The query sanitizer mitigates this by processing the raw LLM output before it reaches the embedding model, ensuring the vector store receives only the semantically relevant portion of the text.

The Four-Step Sanitization Strategy

The sanitizer implements a cascading fallback strategy defined in mempalace/query_sanitizer.py (lines 28-33). Each step triggers only if the previous one fails, maximizing recovery of retrieval accuracy:

Step 1: Passthrough for Safe Short Queries

If the raw query is ≤ SAFE_QUERY_LENGTH (200 characters), the sanitizer returns the string unchanged. This fast-path preserves the roughly 89% R@10 baseline for clean inputs and avoids unnecessary processing overhead (lines 105-113).

Step 2: Question Extraction

When the input exceeds the safe length, the sanitizer searches for a sentence ending with ? or ? using the _QUESTION_MARK regex. If found, it returns that specific question, trimmed to MAX_QUERY_LENGTH (250 characters) if necessary. This method achieves 85-89% R@10 recovery by isolating the explicit interrogative intent (lines 115-156).

Step 3: Tail-Sentence Extraction

If no explicit question mark exists, the algorithm walks backwards through newline-separated segments to locate the last meaningful sentence ≥ MIN_QUERY_LENGTH (10 characters). This captures implied questions or statements that serve as the actual search intent, yielding 80-89% R@10 (lines 158-178).

Step 4: Tail Truncation Fallback

For all remaining cases—typically malformed or extremely verbose outputs—the sanitizer returns the final MAX_QUERY_LENGTH (250 characters) of the string. This brute-force truncation prevents embedding failures and still recovers 70-80% R@10 (lines 180-191).

Core Implementation Details

The entry point sanitize_query(raw_query: str) -> dict (lines 41-59) orchestrates the pipeline and returns a dictionary containing clean_query, was_sanitized, original_length, clean_length, and method.

Pre-Processing Helpers

Before detection begins, the sanitizer performs three normalization steps:

  1. UTF-8 Safety: strip_lone_surrogates (imported from mempalace/config.py) removes lone surrogate characters to prevent encoding crashes (lines 70-73).
  2. Quote Stripping: _strip_wrapping_quotes removes stray surrounding quotes that appear when LLMs format the query as a string literal (lines 75-88).
  3. Length Guard: _trim_candidate re-splits long candidates on sentence boundaries and selects the longest fragment fitting within MAX_QUERY_LENGTH; if none fit, it falls back to the last MAX_QUERY_LENGTH characters (lines 89-103).

Logging and Observability

Each branch logs a warning containing the length reduction and the sanitization method used, giving operators visibility into how frequently contamination occurs and which recovery paths activate.

Integration with the MemPalace Search Pipeline

The sanitizer sits between the LLM agent and the embedding model. In mempalace/searcher.py, the search pipeline calls sanitize_query immediately before feeding text to the vector store. This ensures that only purified queries enter the retrieval system while the original, verbatim utterance remains stored in MemPalace's memory layer for audit purposes.

Practical Usage and Code Examples

Import and invoke the sanitizer directly when processing raw LLM outputs:

from mempalace.query_sanitizer import sanitize_query

# Example 1: Raw LLM output with system prompt contamination

raw = (
    "You are a helpful assistant specialized in household statistics. "
    "You have access to a knowledge base. Please answer the user's question "
    "based on the provided context. Be concise and factual.\n\n"
    "User: How many cats does the average household have?"
)

result = sanitize_query(raw)
print(result["clean_query"])

# Output: "How many cats does the average household have?"

print(result["method"])

# Output: "question_extraction"

print(result["was_sanitized"])

# Output: True

# Example 2: Clean, short query requires no sanitization

raw = "What is the capital of France?"
result = sanitize_query(raw)
print(result["was_sanitized"])

# Output: False

print(result["clean_query"])

# Output: "What is the capital of France?"

The returned dictionary includes analytics-friendly fields that allow developers to track contamination rates and sanitizer effectiveness in production environments.

Summary

  • Prompt contamination occurs when LLM agents prepend system instructions to search queries, collapsing MemPalace retrieval accuracy to approximately 1%.
  • The sanitize_query function in mempalace/query_sanitizer.py implements a four-step cascade: passthrough, question extraction, tail-sentence extraction, and tail truncation.
  • Key constants define the thresholds: SAFE_QUERY_LENGTH (200), MAX_QUERY_LENGTH (250), and MIN_QUERY_LENGTH (10).
  • The sanitizer recovers 70-89% R@10 depending on the extraction method used, while preserving MemPalace's "verbatim always" storage guarantee for the original data.
  • Integration occurs in mempalace/searcher.py, with preprocessing utilities provided by mempalace/config.py.

Frequently Asked Questions

What is prompt contamination in MemPalace?

Prompt contamination refers to the phenomenon where LLM-generated search requests include lengthy system prompts or instructions alongside the actual user query. Because MemPalace stores utterances verbatim, these bloated strings enter the embedding pipeline and dilute the semantic signal, causing vector search to fail.

How does the query sanitizer affect retrieval metrics?

According to the MemPalace source code, unsanitized contaminated queries achieve only ~1% R@10. After sanitization, recovery rates range from 70% to 89% R@10 depending on which of the four extraction steps succeeds, with passthrough and question extraction performing best.

Can I adjust the length thresholds for my deployment?

Yes. The constants SAFE_QUERY_LENGTH, MAX_QUERY_LENGTH, and MIN_QUERY_LENGTH are defined at lines 28-33 of mempalace/query_sanitizer.py. Modifying these values allows you to tune the trade-off between aggressive truncation and preservation of longer, potentially valid queries.

Where exactly is sanitize_query invoked in the codebase?

The primary integration point is in mempalace/searcher.py, where the function is called immediately before text is passed to the embedding model. The test suite in tests/test_query_sanitizer.py provides comprehensive coverage of all four sanitization branches and edge cases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →