How QuerySanitizer Prevents Prompt Contamination in MemPalace

The QuerySanitizer extracts genuine user intent from AI-contaminated strings using a deterministic four-step heuristic that removes prepended system prompts before embedding, ensuring the actual query dominates the vector representation rather than instructional text.

The mempalace/query_sanitizer.py module in the MemPalace repository implements a lightweight, O(n) defense against prompt contamination, a critical failure mode where AI assistants prepend lengthy system prompts (>2000 characters) to user questions. If sent directly to an embedding model, these instructions dominate the vector representation and cause the actual query to be virtually ignored, destroying retrieval accuracy. By applying a multi-stage extraction algorithm, the sanitizer ensures that only the user's true intent is embedded, preventing the catastrophic retrieval collapse that occurs when system text overshadows semantic content.

The Prompt Contamination Threat

When AI assistants generate search requests programmatically, they often wrap the real user question inside elaborate instructions, role definitions, and formatting rules. If this concatenated string is passed directly to a embedding model, the resulting vector reflects the AI's system prompt rather than the user's information need. MemPalace's sanitizer intercepts these contaminated strings at the search boundary defined in mempalace/searcher.py, applying a recovery pipeline that operates entirely within Python's standard library to avoid external latency.

The Four-Step Sanitization Heuristic

The sanitize_query function orchestrates a cascading fallback strategy that prioritizes linguistic coherence while guaranteeing a usable result. Each step is guarded by length checks against MIN_QUERY_LENGTH (10 characters) to prevent returning empty or nonsensical fragments.

Step 1: Passthrough for Short Queries

If the incoming text is already brief, it is almost certainly clean. The sanitizer returns the original query unchanged when its length is ≤ SAFE_QUERY_LENGTH (200 characters). This fast-path handles roughly 90% of traffic without modification.

Step 2: Question Extraction

For longer inputs, the algorithm searches backwards through sentence and newline boundaries using the _SENTENCE_SPLIT regex for the last segment ending with a question mark (? or fullwidth ?). If found, the candidate is passed to _trim_candidate which enforces the MAX_QUERY_LENGTH limit (250 characters) and strips surrounding quotes defined in QUOTE_CHARS (', "). This method achieves near-full recovery (≈85-89%) by capturing explicit interrogative intent that typically signals the true user query.

Step 3: Tail-Sentence Recovery

When no explicit question mark exists, the sanitizer walks backwards through newline-separated segments, selecting the first fragment that satisfies MIN_QUERY_LENGTH after trimming. This tail-sentence extraction handles statements or commands that lack punctuation but still contain the user's semantic goal, yielding moderate recovery rates of ≈80-89%.

Step 4: Tail Truncation Fallback

As a final safeguard, the algorithm slices the final MAX_QUERY_LENGTH (250) characters from the input. This tail truncation runs when all linguistic heuristics fail, providing minimal viable recovery (≈70-80%) that still prevents the embedding model from drowning in thousands of characters of instruction text.

Core Implementation Details

The sanitization pipeline relies on deterministic string operations defined in mempalace/query_sanitizer.py, with preprocessing support from mempalace/config.py.

UTF-16 Surrogate Handling

Before any extraction logic runs, the input passes through strip_lone_surrogates (imported from mempalace/config.py) to remove malformed UTF-16 surrogate pairs. This prevents downstream crashes in embedding models that expect valid Unicode.

Quote Normalization

The private helper _strip_wrapping_quotes removes matching leading and trailing quote characters, handling nested quote scenarios that often appear when AI assistants format queries as JSON or code blocks.

Length Enforcement and Metadata

The _trim_candidate function applies the final length constraints and returns a dictionary containing:

  • clean_query: The extracted, sanitized string
  • method: The recovery strategy used (passthrough, question_extraction, tail_sentence, or tail_truncation)
  • original_length: Character count of the contaminated input
  • clean_length: Character count of the output
  • was_sanitized: Boolean indicating whether modification occurred

Each sanitization event logs a warning with the original and final lengths and the method employed, making the behavior observable in production without exposing user data.

Practical Usage Examples

Extracting Explicit Questions

When the contaminated text ends with a clear question, the sanitizer isolates it precisely:

from mempalace.query_sanitizer import sanitize_query

raw = """
You are an AI assistant. Follow the guidelines strictly.
...
Now answer the following question: How does QuerySanitizer prevent prompt contamination?
"""

result = sanitize_query(raw)
print(result["clean_query"])

# → "How does QuerySanitizer prevent prompt contamination?"

print(result["method"])

# → "question_extraction"

Handling Implicit Queries

For statements without question marks, the algorithm selects the final meaningful sentence:

raw = """
System prompt... (very long)
User wants to retrieve the latest log entries from the server.
"""

result = sanitize_query(raw)
print(result["clean_query"])

# → "User wants to retrieve the latest log entries from the server."

print(result["method"])

# → "tail_sentence"

Fallback Recovery

When no clear linguistic boundaries exist, the sanitizer falls back to tail truncation:

raw = "A" * 3000  # massive prompt with no delimiters

result = sanitize_query(raw)
print(result["clean_query"])

# → last 250 characters of the prompt

print(result["method"])

# → "tail_truncation"

Inspecting Sanitization Metadata

The response dictionary provides transparency into the transformation:

info = sanitize_query(raw)
print(f"Original length: {info['original_length']}")
print(f"Clean length:    {info['clean_length']}")
print(f"Sanitized?      {info['was_sanitized']}")

Summary

  • Prompt contamination occurs when AI system prompts (>2000 characters) prepended to queries dominate embedding vectors, destroying retrieval quality.
  • The QuerySanitizer in mempalace/query_sanitizer.py implements a four-step cascade: passthrough for short texts, question extraction, tail-sentence recovery, and tail truncation fallback.
  • All operations run in O(n) time using pure-Python regex and string manipulation, with no external dependencies.
  • Constants SAFE_QUERY_LENGTH (200), MAX_QUERY_LENGTH (250), and MIN_QUERY_LENGTH (10) guard against edge cases.
  • The module integrates with mempalace/searcher.py to clean every search request before embedding, ensuring vector representations reflect user intent rather than system instructions.

Frequently Asked Questions

Prompt contamination happens when an AI assistant prepends lengthy instructions or role definitions to a user's actual query. When this combined text is converted to an embedding vector, the system prompt's semantic content overwhelms the user's specific question, causing the retrieval system to match against the AI's instructions rather than the information need. This typically reduces recall@10 metrics to nearly 1%, effectively breaking the search pipeline.

How does QuerySanitizer handle malformed UTF-16 input?

Before applying any extraction heuristics, sanitize_query calls strip_lone_surrogates from mempalace/config.py to sanitize malformed UTF-16 surrogate pairs. This preprocessing step prevents UnicodeEncodeError exceptions or undefined behavior in downstream embedding models that require valid Unicode strings, ensuring the sanitization pipeline never crashes on corrupted input.

What happens when no question mark or clear sentence boundary exists?

When the heuristic cannot locate a question mark (step 2) or a valid sentence fragment (step 3), the sanitizer executes tail truncation (step 4), grabbing the final 250 characters of the input. While this provides minimal linguistic coherence (≈70-80% recovery), it guarantees that the embedding model receives a short, user-focused fragment rather than thousands of characters of system instructions, preventing total retrieval failure.

What are the performance characteristics of the sanitization process?

The entire algorithm runs in O(n) time relative to the input string length and uses only constant auxiliary space. Because it relies on compiled regex patterns (_SENTENCE_SPLIT, _QUESTION_MARK) and standard string slicing rather than machine learning models or external API calls, it adds negligible latency to the search pipeline—typically microseconds per request—making it safe to run on every incoming query in high-throughput environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →