How the Context Compression System in Hermes Agent Handles Long Conversations
Hermes Agent's context compression system automatically summarizes middle conversation turns when token counts approach a configurable threshold, preserving critical first and recent messages while maintaining valid tool-call structures.
The NousResearch/hermes-agent repository implements a sophisticated context compression system to prevent token overflow during extended multi-turn dialogues. By intelligently condensing historical conversation data while protecting essential context boundaries, this system ensures reliable long-term agent operation without manual intervention. The implementation centers on agent/context_compressor.py with supporting utilities in agent/model_metadata.py.
Triggering Compression: Token Thresholds and Preflight Checks
The system monitors token usage proactively to intercept overflowing contexts before API calls fail.
Calculating the Compression Threshold
During initialization, the compressor queries get_model_context_length() from agent/model_metadata.py to determine the model's maximum context size. It then calculates a trigger point using threshold_percent (defaulting to 0.85 or 85% of capacity).
self.context_length = get_model_context_length(model, base_url=base_url) # model_metadata.py
self.threshold_tokens = int(self.context_length * threshold_percent)
This conservative default ensures compression activates before hitting hard limits, leaving buffer room for the summarization process itself.
Preflight Detection with Rough Estimation
Before every LLM request, the agent invokes should_compress_preflight(messages), which performs a cheap token estimate via estimate_messages_tokens_rough. When the estimated count exceeds threshold_tokens, the compression pipeline triggers immediately.
Source: lines 70‑74 of context_compressor.py【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L70-L74】
Protecting Critical Conversation Boundaries
The algorithm selectively preserves conversation structure to maintain coherence and tool functionality.
Safeguarding First and Last Turns
The system always protects the first N (protect_first_n, default 3) and last M (protect_last_n, default 4) messages from summarization. These parameters ensure the model retains initial system instructions and recent context while condensing older middle turns.
Source: class docstring and __init__ values【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L21-L35】
Aligning Tool-Call Boundaries
Because Hermes Agent emits tool-call / tool-result pairs, the compressor adjusts indices to prevent slicing these groups. The helper methods _align_boundary_forward and _align_boundary_backward ensure whole tool interactions remain intact:
_align_boundary_forwardskips forward while encounteringtoolroles_align_boundary_backwardmoves the end index backward if the preceding message contains assistanttool_calls
This alignment prevents "no tool call found" errors that would occur if partial tool sequences reached the LLM.
Source: methods _align_boundary_forward and _align_boundary_backward【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L73-L90】
Generating and Inserting Context Summaries
Middle turns undergo transformation into compact summaries that capture essential information without the full token cost.
Flattening and Summarizing Middle Turns
The _generate_summary method (lines 94‑104) first flattens candidate messages into readable text with role prefixes and truncated content. Tool-call names append to the text for reference. The system sends this flattened text to an auxiliary summarization model (defaulting to Gemini Flash) with specific instructions to prefix responses with [CONTEXT SUMMARY]:.
Source: _generate_summary → lines 94‑104【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L94-L104】
Auxiliary Model Fallback Strategy
If the auxiliary client fails or lacks configuration, the compressor transparently falls back to the user's primary model via _get_fallback_client(). The code automatically prepends the [CONTEXT SUMMARY]: prefix if the model response omits it, ensuring consistent formatting regardless of which model generates the summary.
Source: _call_summary_model and fallback logic【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L46-L62】【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L73-L84】
Reconstructing the Message List
The compress method (lines 38‑52) rebuilds the message array through a three-step process:
- Copies all protected turns unchanged
- Inserts the summary as a single user-role message (
{"role":"user","content":summary}) - Appends remaining tail turns
If no summary model is available, the system drops the middle block entirely and emits a warning, allowing the conversation to continue with reduced context rather than failing.
Source: compress method lines 38‑52【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L38-L52】
Maintaining Tool-Call Integrity After Compression
Removing middle turns risks creating orphaned tool results or incomplete call sequences. The private method _sanitize_tool_pairs (lines 13‑71) resolves these inconsistencies through two mechanisms:
- Deletion: Removes any
toolmessages whosetool_call_idlacks a matching assistanttool_callsentry - Stub insertion: Adds placeholder results for any remaining
tool_callsthat lost their corresponding output
This sanitization guarantees the API receives well-formed message sequences even after aggressive compression.
Source: _sanitize_tool_pairs lines 13‑71【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L13-L71】
Monitoring Compression Statistics
Unless operating in quiet_mode, the system outputs statistical feedback after each compression event, reporting:
- Tokens saved compared to the original message count
- Cumulative compressions performed during the current session
This telemetry helps developers tune threshold_percent and summary target sizes based on actual conversation patterns.
Source: final part of compress【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L55-L62】
Implementation Examples
Basic Compression Setup
from agent.context_compressor import ContextCompressor
compressor = ContextCompressor(
model="anthropic/claude-opus-4.6",
threshold_percent=0.85, # start compressing at 85% of context
protect_first_n=3,
protect_last_n=4,
summary_target_tokens=2500, # aim for a ~2.5k-token summary
)
# `messages` is the list of OpenAI-style messages
if compressor.should_compress_preflight(messages):
messages = compressor.compress(messages, current_tokens=compressor.last_prompt_tokens)
Handling Fallback Scenarios
The compressor automatically handles missing auxiliary keys by failing over to the primary model:
# Attempts auxiliary model (e.g., Gemini Flash) first,
# then silently switches to the main model via OPENAI_BASE_URL
summary = compressor._generate_summary(turns_to_summarize)
Summary
- Threshold-based activation: The context compression system monitors token counts via
should_compress_preflight()and triggers at 85% of model capacity by default - Protected boundaries: First 3 and last 4 turns remain intact, with special alignment logic for tool-call pairs
- Intelligent summarization: Middle turns flatten into text and process through an auxiliary model (with primary model fallback), inserting results as user-role messages
- Tool-call safety:
_sanitize_tool_pairsremoves orphaned results and injects stubs for missing outputs, maintaining API-valid sequences - Operational transparency: Real-time statistics track tokens saved and compression frequency for performance tuning
Frequently Asked Questions
How does the context compression system determine when to summarize messages?
The system calculates a threshold at 85% of the model's total context length (configurable via threshold_percent) during initialization in agent/context_compressor.py. Before each API call, should_compress_preflight() performs a rough token estimate; if the count exceeds this threshold, compression activates immediately to prevent context window overflow.
What happens to tool calls that span the compression boundary?
The compressor uses _align_boundary_forward and _align_boundary_backward methods to adjust indices around tool-call pairs, ensuring it never slices a tool interaction in half. Additionally, _sanitize_tool_pairs validates the final message list by deleting orphaned tool results and adding stub responses for any tool calls that lost their outputs during summarization.
Can I customize how many messages are protected from compression?
Yes, the ContextCompressor class accepts protect_first_n (default 3) and protect_last_n (default 4) parameters during instantiation. These values ensure initial system instructions and recent conversation context remain uncompressed while middle turns undergo summarization.
What occurs if the auxiliary summarization model is unavailable?
The system implements transparent fallback logic via _get_fallback_client(). If the auxiliary model (typically Gemini Flash) cannot be reached, the compressor automatically routes summary requests to the user's primary model configured through OPENAI_BASE_URL. If all models fail, the middle turns are dropped with a warning rather than crashing the conversation loop.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →