# How Query Rewriting and Intent Prediction Work in LLM-Based Retrieval: A Deep Dive into memU

> Discover how memU optimizes LLM retrieval with query rewriting and intent prediction. Learn how XML prompts enhance memory lookup accuracy and direct answers.

- Repository: [NevaMind AI/memU](https://github.com/nevamind-ai/memu)
- Tags: deep-dive
- Published: 2026-02-19

---

**memU implements query rewriting and intent prediction through a structured LLM workflow that first routes queries between direct answers (`NO_RETRIEVE`) and memory lookups (`RETRIEVE`), then rewrites queries using XML-tagged prompts to optimize retrieval accuracy.**

The `memU` framework employs a sophisticated two-stage mechanism for query rewriting and intent prediction in LLM-based retrieval systems. By leveraging structured prompt templates and deterministic parsing logic in [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py), the system determines whether user queries require external memory access and reformulates them for maximum relevance before vector search execution.

## The Two-Stage Architecture for Intent Prediction and Query Rewriting

The retrieval workflow centers on two LLM-powered components that operate sequentially to optimize query processing.

### Intent Routing: Deciding When to Retrieve

The **intent routing** layer determines whether the current query requires a retrieval step (`RETRIEVE`) or can be answered directly (`NO_RETRIEVE`). Implemented in [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py) within the `RetrieveMixin` class, this layer exposes two primary methods:

- **`_rag_route_intention`** – Handles retrieval-augmented generation (RAG) mode, processing queries through the full decision pipeline.
- **`_llm_route_intention`** – Handles LLM-only mode when vector search is disabled.

When the `route_intention` configuration is disabled, the system bypasses this layer and passes the original query through unchanged. When enabled, the methods instantiate an LLM client and invoke the decision engine.

### The Decision Engine: `_decide_if_retrieval_needed`

At the core of the system lies the `_decide_if_retrieval_needed` method, also defined in [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py). This method orchestrates the interaction between conversation context, retrieved content, and the LLM to produce structured decisions.

The method constructs two critical inputs:

- **`history_text`** – Formatted previous messages via `_format_query_context`.
- **`content_text`** – Already-retrieved material or "No content retrieved yet."

These values populate the `PRE_RETRIEVAL_USER_PROMPT` template from [`src/memu/prompts/retrieve/pre_retrieval_decision.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/retrieve/pre_retrieval_decision.py), which is sent to the LLM alongside the system prompt.

## Step-by-Step RAG Workflow in memU

The retrieval process follows a deterministic five-stage pipeline that iteratively refines queries:

1. **Initial State Construction** – The `retrieve()` method initializes a `WorkflowState` object containing `original_query` (the user's last message), `context_queries` (previous messages), and `skip_rewrite` (set to true for single-turn conversations).

2. **Intent Routing** – The `_rag_route_intention` method evaluates the query. If routing is disabled, it passes the original query verbatim. Otherwise, it proceeds to the decision engine.

3. **Decision and Rewriting** – Inside `_decide_if_retrieval_needed`, the system formats the user prompt using the conversation history and retrieved content. The LLM receives this via `client.chat` with the system prompt and returns structured XML containing `<decision>` and `<rewritten_query>` tags. The `_extract_decision` method parses the `RETRIEVE` or `NO_RETRIEVE` value, while `_extract_rewritten_query` extracts the reformulated query.

4. **State Update** – The workflow state updates `needs_retrieval` based on the decision. The `rewritten_query` field stores the extracted rewrite (or falls back to the original), and `active_query` is set to drive the next retrieval tier (category, item, or resource).

5. **Sufficiency Checking** – After each retrieval tier, the system reuses `_decide_if_retrieval_needed` to evaluate whether additional retrieval is necessary. If the LLM returns `RETRIEVE`, the query undergoes further rewriting and proceeds to the next tier.

## Prompt Engineering for Structured Decisions

The reliability of intent prediction depends on carefully engineered prompts that enforce structured output formats.

### System Prompts and XML Tag Extraction

The [`src/memu/prompts/retrieve/pre_retrieval_decision.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/retrieve/pre_retrieval_decision.py) file defines the **system prompt** and **user prompt** templates that govern LLM behavior. The system prompt encodes strict rules for the decision boundary:

- **`NO_RETRIEVE`** – Triggered for casual chat, pure conversation context, generic knowledge, clarifications, or meta-questions.
- **`RETRIEVE`** – Required for requests referencing past events, user-specific data, preferences, or explicit recalls of stored information.

The LLM embeds its response within explicit XML-like tags:

```xml
<decision>RETRIEVE</decision>
<rewritten_query>
What were my favorite movies last year? Please include any stored ratings and watch dates.
</rewritten_query>

```

The parser methods `_extract_decision` and `_extract_rewritten_query` reliably parse these tags even when the LLM generates additional explanatory text, creating an explicit contract between the model and the application logic.

### Query Rewriting Guidelines

To ensure rewritten queries maintain fidelity to user intent, [`src/memu/prompts/retrieve/query_rewriter_judger.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/retrieve/query_rewriter_judger.py) provides editorial guidelines for the rewriting step. These constraints prevent semantic drift while allowing the LLM to expand abbreviations, resolve pronouns, and add contextual specificity that improves vector search performance.

## Code Implementation: From Prompt to Parsed Result

### High-Level API Usage

```python
from memu.client.openai_wrapper import OpenAIClient
from memu.app.service import MemuService

service = MemuService(
    llm_client=OpenAIClient(model="gpt-4o-mini"),
    # … other dependencies …

)

# A single user query

queries = [{"role": "user", "content": {"text": "What were my favorite movies last year?"}}]

response = await service.retrieve(queries)

print(response["rewritten_query"])   # -> "What were my favorite movies last year? (include any stored ratings)"

print(response["needs_retrieval"])   # -> True

```

### Internal Decision Flow

```python

# Inside RetrieveMixin._decide_if_retrieval_needed

history_text = self._format_query_context(context_queries)   # formats past messages

content_text = retrieved_content or "No content retrieved yet."

user_prompt = PRE_RETRIEVAL_USER_PROMPT.format(
    query=self._escape_prompt_value(query),
    conversation_history=self._escape_prompt_value(history_text),
    retrieved_content=self._escape_prompt_value(content_text),
)

# Send to LLM

response = await client.chat(user_prompt, system_prompt=PRE_RETRIEVAL_SYSTEM_PROMPT)

# Parse result

decision = self._extract_decision(response)          # "RETRIEVE" or "NO_RETRIEVE"

rewritten = self._extract_rewritten_query(response) # may be None

```

### Expected LLM Output Format

```xml
<decision>RETRIEVE</decision>
<rewritten_query>
What were my favorite movies last year? Please include any stored ratings and watch dates.
</rewritten_query>

```

## Design Rationale: Why This Architecture Works

The memU retrieval system achieves reliability through three architectural principles:

**Explicit Contract** – By forcing the model to output structured XML tags, the system eliminates ambiguity in parsing. The `_extract_decision` method can deterministically locate the decision regardless of surrounding text, preventing the "chatty LLM" problem where models add conversational filler.

**Iterative Rewrites** – The tiered retrieval architecture (category → item → resource) allows the LLM to progressively tighten queries. After each tier, the model receives newly retrieved content and can issue a more specific rewrite for the next search, creating a feedback loop that improves relevance with each iteration.

**Separation of Concerns** – Intent routing, sufficiency checking, and ranking logic remain independent workflow steps. Each step maintains its own LLM client configuration and prompt templates, allowing fine-grained optimization without cross-cutting dependencies. This modularity enables testing and versioning of individual components in [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py) without affecting the broader pipeline.

## Summary

- **Intent prediction** in memU relies on explicit XML-tagged outputs from `PRE_RETRIEVAL_SYSTEM_PROMPT` to distinguish between `RETRIEVE` and `NO_RETRIEVE` states.
- **Query rewriting** occurs in `_decide_if_retrieval_needed` within [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py), using conversation history and retrieved content to reformulate queries for optimal vector search.
- The workflow supports **iterative refinement** through sufficiency checks after each retrieval tier, allowing progressive query optimization.
- **Prompt templates** in [`src/memu/prompts/retrieve/pre_retrieval_decision.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/retrieve/pre_retrieval_decision.py) and [`query_rewriter_judger.py`](https://github.com/NevaMind-AI/memU/blob/main/query_rewriter_judger.py) enforce consistent LLM behavior without hard-coded business logic.

## Frequently Asked Questions

### What triggers a RETRIEVE decision in memU?

The LLM returns `RETRIEVE` when the query references past events, user-specific data, personal preferences, or explicitly asks to recall stored information. According to the system prompt in [`src/memu/prompts/retrieve/pre_retrieval_decision.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/retrieve/pre_retrieval_decision.py), requests involving historical user context or personal memory banks require external retrieval, while generic knowledge questions receive `NO_RETRIEVE`.

### How does memU handle queries that don't need external memory?

When the intent router predicts `NO_RETRIEVE`, the system sets `needs_retrieval` to `False` and bypasses the vector database entirely. The `active_query` remains set to the original (or minimally processed) query, and the pipeline proceeds directly to response generation without executing category, item, or resource retrieval tiers.

### Can the query rewriting process be disabled?

Yes. The `WorkflowState` includes a `skip_rewrite` flag that suppresses the rewriting step when set to `True`, typically during single-turn conversations or when the `route_intention` configuration is disabled. In these cases, `_rag_route_intention` passes the `original_query` through unchanged without invoking `_decide_if_retrieval_needed`.

### Where are the prompt templates defined in the codebase?

The primary prompt templates reside in [`src/memu/prompts/retrieve/pre_retrieval_decision.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/retrieve/pre_retrieval_decision.py), which exports `PRE_RETRIEVAL_SYSTEM_PROMPT` and `PRE_RETRIEVAL_USER_PROMPT`. Editorial guidelines for query rewriting live in [`src/memu/prompts/retrieve/query_rewriter_judger.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/retrieve/query_rewriter_judger.py). These files define the exact wording and XML structure that enable deterministic parsing in [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py).