How Hermes Agent Implements Session Search Using FTS5 and LLM Summarization
Hermes Agent implements session search through a two-stage pipeline that first queries a SQLite FTS5 virtual table for relevant messages, then uses a cheap auxiliary LLM to generate query-focused summaries of the matching sessions.
Hermes Agent maintains a persistent memory of past conversations in a SQLite database. When the model needs to recall historic context, it invokes the session_search tool, which combines fast full-text indexing with intelligent summarization to deliver relevant memories without bloating the main context window.
Stage 1: Full-Text Search with SQLite FTS5
Database Schema and Virtual Tables
The storage layer in hermes_state.py defines three core tables:
sessions– Stores conversation metadata (ID, source, model, timestamps).messages– Stores individual messages with roles (user, assistant, tool).messages_fts– An FTS5 virtual table that indexesmessages.contentfor high-performance text search.
The FTS5 virtual table enables boolean queries, prefix matching, and ranked results using the built-in rank column.
The search_messages Implementation
The SessionDB.search_messages method constructs a dynamic SQL query that joins the FTS5 virtual table with the messages and sessions tables:
def search_messages(self,
query: str,
source_filter: List[str] = None,
role_filter: List[str] = None,
limit: int = 20,
offset: int = 0) -> List[Dict[str, Any]]:
where_clauses = ["messages_fts MATCH ?"]
params = [query]
# Filter by source (CLI, Telegram, Discord, etc.)
if source_filter:
source_placeholders = ",".join("?" for _ in source_filter)
where_clauses.append(f"s.source IN ({source_placeholders})")
params.extend(source_filter)
# Optional role filter (user, assistant, tool)
if role_filter:
role_placeholders = ",".join("?" for _ in role_filter)
where_clauses.append(f"m.role IN ({role_placeholders})")
params.extend(role_filter)
where_sql = " AND ".join(where_clauses)
params.extend([limit, offset])
sql = f"""
SELECT
m.id,
m.session_id,
m.role,
snippet(messages_fts, 0, '>>>', '<<<', '...', 40) AS snippet,
m.content,
m.timestamp,
m.tool_name,
s.source,
s.model,
s.started_at AS session_started
FROM messages_fts
JOIN messages m ON m.id = messages_fts.rowid
JOIN sessions s ON s.id = m.session_id
WHERE {where_sql}
ORDER BY rank
LIMIT ? OFFSET ?
"""
cursor = self._conn.execute(sql, params)
return [dict(row) for row in cursor.fetchall()]
The query leverages the FTS5 MATCH operator for boolean logic and the snippet() function to return highlighted excerpts showing where terms appear. Results are ordered by the FTS5 rank column, ensuring the most relevant matches appear first.
Stage 2: LLM-Powered Session Summarization
Transcript Preparation and Truncation
Raw message histories can exceed context limits, so session_search_tool.py implements focused truncation:
- Deduplication: Groups matches by
session_id, keeping only the top N sessions (default 3, max 5). - Context extraction: Loads the full conversation via
SessionDB.get_messages_as_conversation(). - Focused truncation: The
_truncate_around_matches()function centers the transcript on the first occurrence of any query term, trimming to 100,000 characters (MAX_SESSION_CHARS).
This ensures the auxiliary LLM receives only relevant context, reducing latency and cost.
Auxiliary LLM Configuration
The summarizer uses a cheap, fast auxiliary model rather than the main conversation LLM. The get_async_text_auxiliary_client() function in agent/auxiliary_client.py resolves to Google Gemini Flash by default (specifically google/gemini-3-flash-preview via OpenRouter).
Using a lightweight model for summarization keeps the main agent's context window free for high-value reasoning while minimizing API costs.
The _summarize_session Function
The core summarization logic constructs a focused prompt and streams the result:
async def _summarize_session(conversation_text: str,
query: str,
session_meta: Dict[str, Any]) -> Optional[str]:
system_prompt = (
"You are reviewing a past conversation transcript to help recall what happened. "
"Summarize the conversation with a focus on the search topic. Include what the user asked, "
"what actions were taken, key decisions made, and any unresolved items."
)
user_prompt = (
f"Search topic: {query}\n"
f"Session source: {session_meta['source']}\n"
f"Session date: {session_meta['started']}\n\n"
f"CONVERSATION TRANSCRIPT:\n{conversation_text}\n\n"
f"Summarize this conversation with focus on: {query}"
)
response = await _async_aux_client.chat.completions.create(
model=_SUMMARIZER_MODEL,
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_prompt},
],
temperature=0.1,
max_tokens=MAX_SUMMARY_TOKENS,
)
return response.choices[0].message.content.strip()
The temperature of 0.1 ensures deterministic, factual output, while the structured prompt forces the model to focus on the specific search query rather than generating generic conversation summaries.
Parallel Execution
All session summaries are generated concurrently using asyncio.gather() (lines 86-92 in session_search_tool.py). The implementation handles both async and sync contexts: if an event loop is already running, it delegates to a thread pool; otherwise, it creates a fresh loop. Robust retry logic (max 3 attempts with exponential backoff) ensures reliability even when the auxiliary API is flaky.
Integration with the Agent Loop
The session_search tool is registered in tools/registry.py, making it available to the LLM during tool selection. When the model decides historical context is needed, run_agent.py (around line 2570) dispatches the call to tools.session_search_tool.session_search.
The tool receives the current SessionDB instance and active session ID, allowing it to exclude the ongoing conversation from results. The final output is a JSON payload structured as:
{
"success": true,
"query": "docker networking",
"results": [
{
"session_id": "c3b7…",
"when": "March 01, 2026 at 02:15 PM",
"source": "cli",
"model": "anthropic/claude-opus-4.6",
"summary": "The user wanted to set up Docker networking ... "
}
],
"count": 2,
"sessions_searched": 2
}
This JSON is injected back into the conversation as a tool result, giving the main LLM concise, relevant historical context without requiring it to process thousands of tokens of raw chat logs.
Practical Usage Examples
Direct Python API
You can invoke the search pipeline directly from Python for debugging or custom integrations:
from tools.session_search_tool import session_search
from hermes_state import SessionDB
# Open the persistent database (default: ~/.hermes/state.db)
db = SessionDB()
# Search for sessions mentioning deployment
result_json = session_search(
query="kubernetes deployment",
limit=3,
db=db
)
print(result_json) # JSON string with summarized sessions
Hermes CLI Interaction
Within an active Hermes session, simply ask about past conversations:
$ hermes
⚕ ❯ Can you remind me how we configured the PostgreSQL replication last month?
The agent’s LLM automatically invokes session_search with the appropriate query, retrieves summarized results from the FTS5 index, and synthesizes an answer based on the historical context.
Gateway Integrations (Telegram, Discord)
The same flow operates across all gateway interfaces. When a user sends a message via Telegram or Discord, the gateway server processes the request through run_agent.py, which handles the tool dispatch exactly as it does in the CLI. The user receives a contextual reply referencing their prior sessions without any platform-specific differences in the search implementation.
Summary
Hermes Agent’s session search architecture delivers efficient long-term memory through a hybrid approach:
- SQLite FTS5 indexing provides sub-second full-text search across all historical messages, supporting boolean logic, prefix wildcards, and relevance-ranked results via the
messages_ftsvirtual table. - Intelligent truncation focuses the context window on query-relevant sections, keeping auxiliary LLM costs low by processing only 100,000 characters per session around match locations.
- Auxiliary LLM summarization uses fast, cheap models (Gemini Flash by default) to generate focused recaps, preserving the main model’s context window for high-value reasoning.
- Async parallel execution processes multiple session summaries concurrently with robust retry logic, ensuring responsive performance even when searching across years of chat history.
Frequently Asked Questions
How does the FTS5 search handle complex query syntax?
The implementation uses the standard FTS5 MATCH operator, which supports boolean operators (AND, OR, NOT), phrase queries with double quotes, and prefix wildcards using the asterisk (e.g., deploy* matches "deployment"). The search_messages method in hermes_state.py passes the query directly to the virtual table, allowing the full expressive power of FTS5 query syntax while returning results ranked by the built-in relevance algorithm.
Why use a separate auxiliary LLM instead of the main model for summarization?
The auxiliary LLM (configured in agent/auxiliary_client.py) serves a distinct purpose from the main conversation model. By default, it uses google/gemini-3-flash-preview via OpenRouter—a fast, inexpensive model optimized for throughput. This separation prevents the main model's context window from being consumed by lengthy historical transcripts, reduces API costs (since the auxiliary model is cheaper per token), and allows parallel processing of multiple session summaries without blocking the primary inference loop.
What happens if the search returns too many matching sessions?
The session_search tool in tools/session_search_tool.py implements deduplication and limiting logic to prevent context overflow. After executing the FTS5 query, it groups results by session_id and retains only the top N distinct sessions (default 3, maximum 5). For each selected session, it truncates the transcript to 100,000 characters centered around the first query match. This ensures that even if the FTS5 query matches hundreds of messages across dozens of sessions, the summarization stage only processes the most relevant few.
Can I search within specific conversation sources (CLI vs Telegram vs Discord)?
Yes, the search_messages method supports source filtering through the source_filter parameter. When the tool is invoked from different gateways (CLI, Telegram, Discord, etc.), it can pass specific source identifiers to narrow results to conversations from that platform. The SQL query dynamically constructs an IN clause for the s.source column, allowing you to restrict searches to specific channels while still leveraging the FTS5 full-text index for fast retrieval.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →