# How the Context Compression System in Hermes Agent Handles Long Conversations

> Learn how Hermes Agent's context compression system manages long conversations by summarizing middle turns, preserving key messages for optimal AI interaction.

- Repository: [Nous Research/hermes-agent](https://github.com/NousResearch/hermes-agent)
- Tags: deep-dive
- Published: 2026-03-09

---

**Hermes Agent's context compression system automatically summarizes middle conversation turns when token counts approach a configurable threshold, preserving critical first and recent messages while maintaining valid tool-call structures.**

The `NousResearch/hermes-agent` repository implements a sophisticated **context compression system** to prevent token overflow during extended multi-turn dialogues. By intelligently condensing historical conversation data while protecting essential context boundaries, this system ensures reliable long-term agent operation without manual intervention. The implementation centers on [`agent/context_compressor.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/context_compressor.py) with supporting utilities in [`agent/model_metadata.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/model_metadata.py).

## Triggering Compression: Token Thresholds and Preflight Checks

The system monitors token usage proactively to intercept overflowing contexts before API calls fail.

### Calculating the Compression Threshold

During initialization, the compressor queries `get_model_context_length()` from [`agent/model_metadata.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/model_metadata.py) to determine the model's maximum context size. It then calculates a trigger point using `threshold_percent` (defaulting to **0.85** or 85% of capacity).

```python
self.context_length = get_model_context_length(model, base_url=base_url)   # model_metadata.py

self.threshold_tokens = int(self.context_length * threshold_percent)

```

This conservative default ensures compression activates before hitting hard limits, leaving buffer room for the summarization process itself.

### Preflight Detection with Rough Estimation

Before every LLM request, the agent invokes `should_compress_preflight(messages)`, which performs a cheap token estimate via `estimate_messages_tokens_rough`. When the estimated count exceeds `threshold_tokens`, the compression pipeline triggers immediately.

*Source: lines 70‑74 of [`context_compressor.py`](https://github.com/NousResearch/hermes-agent/blob/main/context_compressor.py)*【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L70-L74】

## Protecting Critical Conversation Boundaries

The algorithm selectively preserves conversation structure to maintain coherence and tool functionality.

### Safeguarding First and Last Turns

The system always protects the **first N** (`protect_first_n`, default **3**) and **last M** (`protect_last_n`, default **4**) messages from summarization. These parameters ensure the model retains initial system instructions and recent context while condensing older middle turns.

*Source: class docstring and `__init__` values*【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L21-L35】

### Aligning Tool-Call Boundaries

Because Hermes Agent emits *tool-call / tool-result* pairs, the compressor adjusts indices to prevent slicing these groups. The helper methods `_align_boundary_forward` and `_align_boundary_backward` ensure whole tool interactions remain intact:

- `_align_boundary_forward` skips forward while encountering `tool` roles
- `_align_boundary_backward` moves the end index backward if the preceding message contains assistant `tool_calls`

This alignment prevents "no tool call found" errors that would occur if partial tool sequences reached the LLM.

*Source: methods `_align_boundary_forward` and `_align_boundary_backward`*【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L73-L90】

## Generating and Inserting Context Summaries

Middle turns undergo transformation into compact summaries that capture essential information without the full token cost.

### Flattening and Summarizing Middle Turns

The `_generate_summary` method (lines 94‑104) first flattens candidate messages into readable text with role prefixes and truncated content. Tool-call names append to the text for reference. The system sends this flattened text to an **auxiliary summarization model** (defaulting to Gemini Flash) with specific instructions to prefix responses with `[CONTEXT SUMMARY]:`.

*Source: `_generate_summary` → lines 94‑104*【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L94-L104】

### Auxiliary Model Fallback Strategy

If the auxiliary client fails or lacks configuration, the compressor transparently falls back to the user's primary model via `_get_fallback_client()`. The code automatically prepends the `[CONTEXT SUMMARY]:` prefix if the model response omits it, ensuring consistent formatting regardless of which model generates the summary.

*Source: `_call_summary_model` and fallback logic*【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L46-L62】【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L73-L84】

### Reconstructing the Message List

The `compress` method (lines 38‑52) rebuilds the message array through a three-step process:

1. Copies all protected turns unchanged
2. Inserts the **summary as a single user-role message** (`{"role":"user","content":summary}`)
3. Appends remaining tail turns

If no summary model is available, the system drops the middle block entirely and emits a warning, allowing the conversation to continue with reduced context rather than failing.

*Source: `compress` method lines 38‑52*【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L38-L52】

## Maintaining Tool-Call Integrity After Compression

Removing middle turns risks creating orphaned tool results or incomplete call sequences. The private method `_sanitize_tool_pairs` (lines 13‑71) resolves these inconsistencies through two mechanisms:

- **Deletion**: Removes any `tool` messages whose `tool_call_id` lacks a matching assistant `tool_calls` entry
- **Stub insertion**: Adds placeholder results for any remaining `tool_calls` that lost their corresponding output

This sanitization guarantees the API receives well-formed message sequences even after aggressive compression.

*Source: `_sanitize_tool_pairs` lines 13‑71*【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L13-L71】

## Monitoring Compression Statistics

Unless operating in `quiet_mode`, the system outputs statistical feedback after each compression event, reporting:

- **Tokens saved** compared to the original message count
- **Cumulative compressions** performed during the current session

This telemetry helps developers tune `threshold_percent` and summary target sizes based on actual conversation patterns.

*Source: final part of `compress`*【/cache/repos/github.com/NousResearch/hermes-agent/main/agent/context_compressor.py#L55-L62】

## Implementation Examples

### Basic Compression Setup

```python
from agent.context_compressor import ContextCompressor

compressor = ContextCompressor(
    model="anthropic/claude-opus-4.6",
    threshold_percent=0.85,      # start compressing at 85% of context

    protect_first_n=3,
    protect_last_n=4,
    summary_target_tokens=2500,  # aim for a ~2.5k-token summary

)

# `messages` is the list of OpenAI-style messages

if compressor.should_compress_preflight(messages):
    messages = compressor.compress(messages, current_tokens=compressor.last_prompt_tokens)

```

### Handling Fallback Scenarios

The compressor automatically handles missing auxiliary keys by failing over to the primary model:

```python

# Attempts auxiliary model (e.g., Gemini Flash) first,

# then silently switches to the main model via OPENAI_BASE_URL

summary = compressor._generate_summary(turns_to_summarize)

```

## Summary

- **Threshold-based activation**: The context compression system monitors token counts via `should_compress_preflight()` and triggers at 85% of model capacity by default
- **Protected boundaries**: First 3 and last 4 turns remain intact, with special alignment logic for tool-call pairs
- **Intelligent summarization**: Middle turns flatten into text and process through an auxiliary model (with primary model fallback), inserting results as user-role messages
- **Tool-call safety**: `_sanitize_tool_pairs` removes orphaned results and injects stubs for missing outputs, maintaining API-valid sequences
- **Operational transparency**: Real-time statistics track tokens saved and compression frequency for performance tuning

## Frequently Asked Questions

### How does the context compression system determine when to summarize messages?

The system calculates a threshold at 85% of the model's total context length (configurable via `threshold_percent`) during initialization in [`agent/context_compressor.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/context_compressor.py). Before each API call, `should_compress_preflight()` performs a rough token estimate; if the count exceeds this threshold, compression activates immediately to prevent context window overflow.

### What happens to tool calls that span the compression boundary?

The compressor uses `_align_boundary_forward` and `_align_boundary_backward` methods to adjust indices around tool-call pairs, ensuring it never slices a tool interaction in half. Additionally, `_sanitize_tool_pairs` validates the final message list by deleting orphaned tool results and adding stub responses for any tool calls that lost their outputs during summarization.

### Can I customize how many messages are protected from compression?

Yes, the `ContextCompressor` class accepts `protect_first_n` (default 3) and `protect_last_n` (default 4) parameters during instantiation. These values ensure initial system instructions and recent conversation context remain uncompressed while middle turns undergo summarization.

### What occurs if the auxiliary summarization model is unavailable?

The system implements transparent fallback logic via `_get_fallback_client()`. If the auxiliary model (typically Gemini Flash) cannot be reached, the compressor automatically routes summary requests to the user's primary model configured through `OPENAI_BASE_URL`. If all models fail, the middle turns are dropped with a warning rather than crashing the conversation loop.