How the Context Compression System in Hermes Agent Handles Long Conversations

Hermes Agent's context compression system automatically summarizes middle conversation turns when token counts approach a configurable threshold, preserving critical first and recent messages while maintaining valid tool-call structures.

The NousResearch/hermes-agent repository implements a sophisticated context compression system to prevent token overflow during extended multi-turn dialogues. By intelligently condensing historical conversation data while protecting essential context boundaries, this system ensures reliable long-term agent operation without manual intervention. The implementation centers on agent/context_compressor.py with supporting utilities in agent/model_metadata.py.

Triggering Compression: Token Thresholds and Preflight Checks

The system monitors token usage proactively to intercept overflowing contexts before API calls fail.

Calculating the Compression Threshold

During initialization, the compressor queries get_model_context_length() from agent/model_metadata.py to determine the model's maximum context size. It then calculates a trigger point using threshold_percent (defaulting to 0.85 or 85% of capacity).

self.context_length = get_model_context_length(model, base_url=base_url)   # model_metadata.py

self.threshold_tokens = int(self.context_length * threshold_percent)

This conservative default ensures compression activates before hitting hard limits, leaving buffer room for the summarization process itself.

Preflight Detection with Rough Estimation

Before every LLM request, the agent invokes should_compress_preflight(messages), which performs a cheap token estimate via estimate_messages_tokens_rough. When the estimated count exceeds threshold_tokens, the compression pipeline triggers immediately.

Source: lines 70‑74 of context_compressor.py【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L70-L74】

Protecting Critical Conversation Boundaries

The algorithm selectively preserves conversation structure to maintain coherence and tool functionality.

Safeguarding First and Last Turns

The system always protects the first N (protect_first_n, default 3) and last M (protect_last_n, default 4) messages from summarization. These parameters ensure the model retains initial system instructions and recent context while condensing older middle turns.

Source: class docstring and __init__ values【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L21-L35】

Aligning Tool-Call Boundaries

Because Hermes Agent emits tool-call / tool-result pairs, the compressor adjusts indices to prevent slicing these groups. The helper methods _align_boundary_forward and _align_boundary_backward ensure whole tool interactions remain intact:

  • _align_boundary_forward skips forward while encountering tool roles
  • _align_boundary_backward moves the end index backward if the preceding message contains assistant tool_calls

This alignment prevents "no tool call found" errors that would occur if partial tool sequences reached the LLM.

Source: methods _align_boundary_forward and _align_boundary_backward【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L73-L90】

Generating and Inserting Context Summaries

Middle turns undergo transformation into compact summaries that capture essential information without the full token cost.

Flattening and Summarizing Middle Turns

The _generate_summary method (lines 94‑104) first flattens candidate messages into readable text with role prefixes and truncated content. Tool-call names append to the text for reference. The system sends this flattened text to an auxiliary summarization model (defaulting to Gemini Flash) with specific instructions to prefix responses with [CONTEXT SUMMARY]:.

Source: _generate_summary → lines 94‑104【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L94-L104】

Auxiliary Model Fallback Strategy

If the auxiliary client fails or lacks configuration, the compressor transparently falls back to the user's primary model via _get_fallback_client(). The code automatically prepends the [CONTEXT SUMMARY]: prefix if the model response omits it, ensuring consistent formatting regardless of which model generates the summary.

Source: _call_summary_model and fallback logic【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L46-L62】【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L73-L84】

Reconstructing the Message List

The compress method (lines 38‑52) rebuilds the message array through a three-step process:

  1. Copies all protected turns unchanged
  2. Inserts the summary as a single user-role message ({"role":"user","content":summary})
  3. Appends remaining tail turns

If no summary model is available, the system drops the middle block entirely and emits a warning, allowing the conversation to continue with reduced context rather than failing.

Source: compress method lines 38‑52【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L38-L52】

Maintaining Tool-Call Integrity After Compression

Removing middle turns risks creating orphaned tool results or incomplete call sequences. The private method _sanitize_tool_pairs (lines 13‑71) resolves these inconsistencies through two mechanisms:

  • Deletion: Removes any tool messages whose tool_call_id lacks a matching assistant tool_calls entry
  • Stub insertion: Adds placeholder results for any remaining tool_calls that lost their corresponding output

This sanitization guarantees the API receives well-formed message sequences even after aggressive compression.

Source: _sanitize_tool_pairs lines 13‑71【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L13-L71】

Monitoring Compression Statistics

Unless operating in quiet_mode, the system outputs statistical feedback after each compression event, reporting:

  • Tokens saved compared to the original message count
  • Cumulative compressions performed during the current session

This telemetry helps developers tune threshold_percent and summary target sizes based on actual conversation patterns.

Source: final part of compress【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L55-L62】

Implementation Examples

Basic Compression Setup

from agent.context_compressor import ContextCompressor

compressor = ContextCompressor(
    model="anthropic/claude-opus-4.6",
    threshold_percent=0.85,      # start compressing at 85% of context

    protect_first_n=3,
    protect_last_n=4,
    summary_target_tokens=2500,  # aim for a ~2.5k-token summary

)

# `messages` is the list of OpenAI-style messages

if compressor.should_compress_preflight(messages):
    messages = compressor.compress(messages, current_tokens=compressor.last_prompt_tokens)

Handling Fallback Scenarios

The compressor automatically handles missing auxiliary keys by failing over to the primary model:


# Attempts auxiliary model (e.g., Gemini Flash) first,

# then silently switches to the main model via OPENAI_BASE_URL

summary = compressor._generate_summary(turns_to_summarize)

Summary

  • Threshold-based activation: The context compression system monitors token counts via should_compress_preflight() and triggers at 85% of model capacity by default
  • Protected boundaries: First 3 and last 4 turns remain intact, with special alignment logic for tool-call pairs
  • Intelligent summarization: Middle turns flatten into text and process through an auxiliary model (with primary model fallback), inserting results as user-role messages
  • Tool-call safety: _sanitize_tool_pairs removes orphaned results and injects stubs for missing outputs, maintaining API-valid sequences
  • Operational transparency: Real-time statistics track tokens saved and compression frequency for performance tuning

Frequently Asked Questions

How does the context compression system determine when to summarize messages?

The system calculates a threshold at 85% of the model's total context length (configurable via threshold_percent) during initialization in agent/context_compressor.py. Before each API call, should_compress_preflight() performs a rough token estimate; if the count exceeds this threshold, compression activates immediately to prevent context window overflow.

What happens to tool calls that span the compression boundary?

The compressor uses _align_boundary_forward and _align_boundary_backward methods to adjust indices around tool-call pairs, ensuring it never slices a tool interaction in half. Additionally, _sanitize_tool_pairs validates the final message list by deleting orphaned tool results and adding stub responses for any tool calls that lost their outputs during summarization.

Can I customize how many messages are protected from compression?

Yes, the ContextCompressor class accepts protect_first_n (default 3) and protect_last_n (default 4) parameters during instantiation. These values ensure initial system instructions and recent conversation context remain uncompressed while middle turns undergo summarization.

What occurs if the auxiliary summarization model is unavailable?

The system implements transparent fallback logic via _get_fallback_client(). If the auxiliary model (typically Gemini Flash) cannot be reached, the compressor automatically routes summary requests to the user's primary model configured through OPENAI_BASE_URL. If all models fail, the middle turns are dropped with a warning rather than crashing the conversation loop.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →