# How Hermes Agent Implements Session Search Using FTS5 and LLM Summarization

> Discover how Hermes Agent enhances session search with FTS5 for fast retrieval and LLM summarization for concise, relevant results. Learn about our efficient two-stage pipeline.

- Repository: [Nous Research/hermes-agent](https://github.com/NousResearch/hermes-agent)
- Tags: deep-dive
- Published: 2026-03-09

---

**Hermes Agent implements session search through a two-stage pipeline that first queries a SQLite FTS5 virtual table for relevant messages, then uses a cheap auxiliary LLM to generate query-focused summaries of the matching sessions.**

Hermes Agent maintains a persistent memory of past conversations in a SQLite database. When the model needs to recall historic context, it invokes the `session_search` tool, which combines fast full-text indexing with intelligent summarization to deliver relevant memories without bloating the main context window.

## Stage 1: Full-Text Search with SQLite FTS5

### Database Schema and Virtual Tables

The storage layer in [`hermes_state.py`](https://github.com/NousResearch/hermes-agent/blob/main/hermes_state.py) defines three core tables:

* **`sessions`** – Stores conversation metadata (ID, source, model, timestamps).
* **`messages`** – Stores individual messages with roles (user, assistant, tool).
* **`messages_fts`** – An FTS5 virtual table that indexes `messages.content` for high-performance text search.

The FTS5 virtual table enables boolean queries, prefix matching, and ranked results using the built-in `rank` column.

### The search_messages Implementation

The `SessionDB.search_messages` method constructs a dynamic SQL query that joins the FTS5 virtual table with the messages and sessions tables:

```python
def search_messages(self,
                    query: str,
                    source_filter: List[str] = None,
                    role_filter: List[str] = None,
                    limit: int = 20,
                    offset: int = 0) -> List[Dict[str, Any]]:
    where_clauses = ["messages_fts MATCH ?"]
    params = [query]

    # Filter by source (CLI, Telegram, Discord, etc.)

    if source_filter:
        source_placeholders = ",".join("?" for _ in source_filter)
        where_clauses.append(f"s.source IN ({source_placeholders})")
        params.extend(source_filter)

    # Optional role filter (user, assistant, tool)

    if role_filter:
        role_placeholders = ",".join("?" for _ in role_filter)
        where_clauses.append(f"m.role IN ({role_placeholders})")
        params.extend(role_filter)

    where_sql = " AND ".join(where_clauses)
    params.extend([limit, offset])

    sql = f"""
        SELECT
            m.id,
            m.session_id,
            m.role,
            snippet(messages_fts, 0, '>>>', '<<<', '...', 40) AS snippet,
            m.content,
            m.timestamp,
            m.tool_name,
            s.source,
            s.model,
            s.started_at AS session_started
        FROM messages_fts
        JOIN messages m ON m.id = messages_fts.rowid
        JOIN sessions s ON s.id = m.session_id
        WHERE {where_sql}
        ORDER BY rank
        LIMIT ? OFFSET ?
    """
    cursor = self._conn.execute(sql, params)
    return [dict(row) for row in cursor.fetchall()]

```

The query leverages the **FTS5 `MATCH` operator** for boolean logic and the **`snippet()` function** to return highlighted excerpts showing where terms appear. Results are ordered by the FTS5 `rank` column, ensuring the most relevant matches appear first.

## Stage 2: LLM-Powered Session Summarization

### Transcript Preparation and Truncation

Raw message histories can exceed context limits, so [`session_search_tool.py`](https://github.com/NousResearch/hermes-agent/blob/main/session_search_tool.py) implements focused truncation:

1. **Deduplication**: Groups matches by `session_id`, keeping only the top *N* sessions (default 3, max 5).
2. **Context extraction**: Loads the full conversation via `SessionDB.get_messages_as_conversation()`.
3. **Focused truncation**: The `_truncate_around_matches()` function centers the transcript on the first occurrence of any query term, trimming to **100,000 characters** (`MAX_SESSION_CHARS`).

This ensures the auxiliary LLM receives only relevant context, reducing latency and cost.

### Auxiliary LLM Configuration

The summarizer uses a cheap, fast auxiliary model rather than the main conversation LLM. The `get_async_text_auxiliary_client()` function in [`agent/auxiliary_client.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/auxiliary_client.py) resolves to **Google Gemini Flash** by default (specifically `google/gemini-3-flash-preview` via OpenRouter).

Using a lightweight model for summarization keeps the main agent's context window free for high-value reasoning while minimizing API costs.

### The _summarize_session Function

The core summarization logic constructs a focused prompt and streams the result:

```python
async def _summarize_session(conversation_text: str,
                             query: str,
                             session_meta: Dict[str, Any]) -> Optional[str]:
    system_prompt = (
        "You are reviewing a past conversation transcript to help recall what happened. "
        "Summarize the conversation with a focus on the search topic. Include what the user asked, "
        "what actions were taken, key decisions made, and any unresolved items."
    )
    
    user_prompt = (
        f"Search topic: {query}\n"
        f"Session source: {session_meta['source']}\n"
        f"Session date: {session_meta['started']}\n\n"
        f"CONVERSATION TRANSCRIPT:\n{conversation_text}\n\n"
        f"Summarize this conversation with focus on: {query}"
    )
    
    response = await _async_aux_client.chat.completions.create(
        model=_SUMMARIZER_MODEL,
        messages=[
            {"role": "system", "content": system_prompt},
            {"role": "user", "content": user_prompt},
        ],
        temperature=0.1,
        max_tokens=MAX_SUMMARY_TOKENS,
    )
    return response.choices[0].message.content.strip()

```

The **temperature of 0.1** ensures deterministic, factual output, while the structured prompt forces the model to focus on the specific search query rather than generating generic conversation summaries.

### Parallel Execution

All session summaries are generated concurrently using `asyncio.gather()` (lines 86-92 in [`session_search_tool.py`](https://github.com/NousResearch/hermes-agent/blob/main/session_search_tool.py)). The implementation handles both async and sync contexts: if an event loop is already running, it delegates to a thread pool; otherwise, it creates a fresh loop. Robust retry logic (max 3 attempts with exponential backoff) ensures reliability even when the auxiliary API is flaky.

## Integration with the Agent Loop

The `session_search` tool is registered in [`tools/registry.py`](https://github.com/NousResearch/hermes-agent/blob/main/tools/registry.py), making it available to the LLM during tool selection. When the model decides historical context is needed, [`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py) (around line 2570) dispatches the call to `tools.session_search_tool.session_search`.

The tool receives the current `SessionDB` instance and active session ID, allowing it to exclude the ongoing conversation from results. The final output is a JSON payload structured as:

```json
{
  "success": true,
  "query": "docker networking",
  "results": [
    {
      "session_id": "c3b7…",
      "when": "March 01, 2026 at 02:15 PM",
      "source": "cli",
      "model": "anthropic/claude-opus-4.6",
      "summary": "The user wanted to set up Docker networking ... "
    }
  ],
  "count": 2,
  "sessions_searched": 2
}

```

This JSON is injected back into the conversation as a tool result, giving the main LLM concise, relevant historical context without requiring it to process thousands of tokens of raw chat logs.

## Practical Usage Examples

### Direct Python API

You can invoke the search pipeline directly from Python for debugging or custom integrations:

```python
from tools.session_search_tool import session_search
from hermes_state import SessionDB

# Open the persistent database (default: ~/.hermes/state.db)

db = SessionDB()

# Search for sessions mentioning deployment

result_json = session_search(
    query="kubernetes deployment",
    limit=3,
    db=db
)

print(result_json)  # JSON string with summarized sessions

```

### Hermes CLI Interaction

Within an active Hermes session, simply ask about past conversations:

```bash
$ hermes
⚕ ❯ Can you remind me how we configured the PostgreSQL replication last month?

```

The agent’s LLM automatically invokes `session_search` with the appropriate query, retrieves summarized results from the FTS5 index, and synthesizes an answer based on the historical context.

### Gateway Integrations (Telegram, Discord)

The same flow operates across all gateway interfaces. When a user sends a message via Telegram or Discord, the gateway server processes the request through [`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py), which handles the tool dispatch exactly as it does in the CLI. The user receives a contextual reply referencing their prior sessions without any platform-specific differences in the search implementation.

## Summary

Hermes Agent’s session search architecture delivers efficient long-term memory through a hybrid approach:

* **SQLite FTS5 indexing** provides sub-second full-text search across all historical messages, supporting boolean logic, prefix wildcards, and relevance-ranked results via the `messages_fts` virtual table.
* **Intelligent truncation** focuses the context window on query-relevant sections, keeping auxiliary LLM costs low by processing only 100,000 characters per session around match locations.
* **Auxiliary LLM summarization** uses fast, cheap models (Gemini Flash by default) to generate focused recaps, preserving the main model’s context window for high-value reasoning.
* **Async parallel execution** processes multiple session summaries concurrently with robust retry logic, ensuring responsive performance even when searching across years of chat history.

## Frequently Asked Questions

### How does the FTS5 search handle complex query syntax?

The implementation uses the standard FTS5 `MATCH` operator, which supports boolean operators (`AND`, `OR`, `NOT`), phrase queries with double quotes, and prefix wildcards using the asterisk (e.g., `deploy*` matches "deployment"). The `search_messages` method in [`hermes_state.py`](https://github.com/NousResearch/hermes-agent/blob/main/hermes_state.py) passes the query directly to the virtual table, allowing the full expressive power of FTS5 query syntax while returning results ranked by the built-in relevance algorithm.

### Why use a separate auxiliary LLM instead of the main model for summarization?

The auxiliary LLM (configured in [`agent/auxiliary_client.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/auxiliary_client.py)) serves a distinct purpose from the main conversation model. By default, it uses `google/gemini-3-flash-preview` via OpenRouter—a fast, inexpensive model optimized for throughput. This separation prevents the main model's context window from being consumed by lengthy historical transcripts, reduces API costs (since the auxiliary model is cheaper per token), and allows parallel processing of multiple session summaries without blocking the primary inference loop.

### What happens if the search returns too many matching sessions?

The `session_search` tool in [`tools/session_search_tool.py`](https://github.com/NousResearch/hermes-agent/blob/main/tools/session_search_tool.py) implements deduplication and limiting logic to prevent context overflow. After executing the FTS5 query, it groups results by `session_id` and retains only the top *N* distinct sessions (default 3, maximum 5). For each selected session, it truncates the transcript to 100,000 characters centered around the first query match. This ensures that even if the FTS5 query matches hundreds of messages across dozens of sessions, the summarization stage only processes the most relevant few.

### Can I search within specific conversation sources (CLI vs Telegram vs Discord)?

Yes, the `search_messages` method supports source filtering through the `source_filter` parameter. When the tool is invoked from different gateways (CLI, Telegram, Discord, etc.), it can pass specific source identifiers to narrow results to conversations from that platform. The SQL query dynamically constructs an `IN` clause for the `s.source` column, allowing you to restrict searches to specific channels while still leveraging the FTS5 full-text index for fast retrieval.