# AIAgent Class Architecture and Agent Loop in Hermes Agent: A Deep Dive

> Explore the AIAgent class architecture and agent loop in NousResearch/hermes-agent. Understand LLM interactions, tool execution, and context management within run_agent.py.

- Repository: [Nous Research/hermes-agent](https://github.com/NousResearch/hermes-agent)
- Tags: deep-dive
- Published: 2026-03-09

---

**The `AIAgent` class in [`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py) orchestrates LLM interactions through a robust while-loop that handles tool execution, context compression, and interrupt handling, exposing simple `chat()` and complex `run_conversation()` entry points.**

The `AIAgent` class serves as the core runtime engine for the NousResearch/hermes-agent project, encapsulating all state required for a single agent session. This architecture supports everything from model selection and tool discovery to trajectory logging and graceful interruption, making it a production-ready implementation for autonomous LLM agents.

## Core Architecture of the AIAgent Class

The `AIAgent` class is defined in **[`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py)** and initializes a comprehensive runtime environment during instantiation. The constructor assembles multiple subsystems—including context compression, session persistence, and interrupt handling—while caching the system prompt for efficient prefix-caching on Claude models.

### Initialization and Runtime State

During initialization, the constructor establishes the LLM endpoint configuration through attributes like `model`, `base_url`, `api_key`, and `provider`, defaulting to OpenRouter with Claude Opus. It also sets global safety limits via `max_iterations` and `iteration_budget`, which cap the total number of tool-calling steps across the entire session.

The initialization sequence dynamically discovers available tools through the `tools` and `valid_tool_names` attributes, filtering by enabled or disabled toolsets. It also instantiates the `context_compressor` for automatic token management and configures optional persistence layers including `session_db` for SQLite storage and `session_log_file` for JSON debugging logs.

### Key Attributes and Configuration

The class maintains several specialized fields for advanced functionality:

- **`memory_store`** and **`honcho`**: Support long-term memory and cross-session user modeling
- **`ephemeral_system_prompt`** and **`prefill_messages`**: Allow prompt fragments injected only at API-call time without polluting session logs
- **`interrupt`** handling fields (`_interrupt_requested`, `_interrupt_message`): Enable external threads (CLI or gateway) to safely abort the loop
- **`quiet_mode`**, **`verbose_logging`**, **`tool_progress_callback`**: Control UI feedback and logging granularity

The constructor prints a concise status banner (unless `quiet_mode` is enabled) and caches the final system prompt to optimize performance on Claude models【/cache/repos/github.com/NousResearch/hermes-agent/main/run_agent.py#L142-L170】.

## The Agent Loop Implementation in run_conversation()

The `run_conversation()` method implements the core agent loop—a sophisticated while-loop that iteratively calls the LLM, processes tool calls, manages context windows, and handles interruptions. This method returns the complete message history as a dictionary, unlike the simplified `chat()` method.

### Loop Structure and Guard Conditions

The loop guard ensures safe execution boundaries:

```python
while api_call_count < self.max_iterations and self.iteration_budget.remaining > 0:

```

This condition appears at line 3033-3036 in [`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py)【/cache/repos/github.com/NousResearch/hermes-agent/main/run_agent.py#L3033-L3036】, preventing infinite loops and enforcing budget constraints across the entire conversation.

### Message Preparation and API Calls

Each iteration begins by checking for interrupt requests, then builds the API payload through the `_build_api_kwargs` helper method. This process:

1. Copies the current message list and embeds any `reasoning_content`
2. Strips unsupported fields for API compatibility
3. Prepends the cached system prompt combined with `ephemeral_system_prompt`
4. Applies optional prompt-caching for Claude models

The `_interruptible_api_call` wrapper manages the actual HTTP call with exponential-backoff retries, handling rate-limit errors and payload-too-large conditions by triggering automatic context compression. Invalid responses trigger retry logic with back-off, while non-retryable client errors abort the loop.

### Tool Execution and Context Management

After receiving the LLM response, the loop parses the output to separate normal text, tool calls, and reasoning tags. The `_execute_tool_calls` helper validates each tool call against `valid_tool_names`, rejects malformed JSON arguments (with up to 3 retries), and dispatches execution through the central tool registry.

Special handling includes budget refunds for `execute_code` calls (considered cheap operations) and comprehensive error capture. After tool execution, the loop conditionally triggers `_compress_context` when token usage approaches the model's limit, summarizing middle turns to maintain the conversation window.

The `_persist_session` method incrementally writes the session state to both the JSON log file and SQLite database, ensuring durability for debugging and analysis.

## Entry Points: chat() vs run_conversation()

The `AIAgent` class exposes two distinct interfaces for different use cases:

**`chat(message: str) → str`**: Provides a simplified one-shot query interface that returns only the final assistant response as a string. This method is ideal for simple integrations where the full conversation history and tool execution details are not required.

**`run_conversation(...) → dict`**: Offers the full-featured agent loop that returns the complete message history, tool trajectories, and metadata. This method supports advanced features including interrupt handling, context compression, and session persistence, making it suitable for production deployments requiring observability and control.

Both methods utilize the same underlying loop infrastructure, ensuring consistent behavior across simple and complex usage patterns.

## Summary

- The **`AIAgent`** class in [`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py) encapsulates all runtime state for Hermes Agent sessions, from model configuration to interrupt handling.
- **Initialization** establishes LLM endpoints, tool registries, context compressors, and optional persistence layers (SQLite/JSON).
- The **`run_conversation()`** method implements a robust while-loop guarded by `max_iterations` and `iteration_budget` constraints.
- **Core loop phases** include interrupt detection, API payload construction with prompt caching, exponential-backoff retries, tool validation/execution, and automatic context compression.
- **Two entry points** serve different needs: `chat()` for simple string responses and `run_conversation()` for full trajectory access and production features.

## Frequently Asked Questions

### What is the AIAgent class in Hermes Agent?

The `AIAgent` class is the core runtime engine defined in [`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py) that manages a complete agent session. It handles LLM API calls, tool discovery and execution, context window management, interrupt handling, and session persistence, providing both simple (`chat`) and advanced (`run_conversation`) interfaces for interacting with the agent.

### How does the agent loop handle infinite loops and budget constraints?

The agent loop uses a dual-guard condition: `while api_call_count < self.max_iterations and self.iteration_budget.remaining > 0`. This ensures the loop terminates after a configurable number of API calls or when the iteration budget is exhausted. Additionally, the loop tracks tool execution costs and refunds iterations for cheap operations like `execute_code`, optimizing budget utilization across the conversation.

### What happens when the conversation context exceeds token limits?

When token usage approaches the model's limit, the `_compress_context` method automatically summarizes middle turns of the conversation to reduce the context window size. This compression occurs during the loop's context management phase, ensuring the API payload stays within limits while preserving the most recent and relevant conversation history. The system also handles payload-too-large errors from the API by triggering immediate compression and retry with exponential backoff.

### How does the AIAgent class support interrupt handling during long-running tasks?

The class maintains interrupt handling fields (`_interrupt_requested`, `_interrupt_message`) that external threads (such as CLI interfaces or gateway processes) can set to request safe loop termination. At the start of each iteration, the loop checks for interrupt requests through the `_interruptible_api_call` wrapper, which raises an `InterruptError` if requested, allowing the agent to abort gracefully while preserving the current session state in the persistence layer.