AIAgent Class Architecture and Agent Loop in Hermes Agent: A Deep Dive
The AIAgent class in run_agent.py orchestrates LLM interactions through a robust while-loop that handles tool execution, context compression, and interrupt handling, exposing simple chat() and complex run_conversation() entry points.
The AIAgent class serves as the core runtime engine for the NousResearch/hermes-agent project, encapsulating all state required for a single agent session. This architecture supports everything from model selection and tool discovery to trajectory logging and graceful interruption, making it a production-ready implementation for autonomous LLM agents.
Core Architecture of the AIAgent Class
The AIAgent class is defined in run_agent.py and initializes a comprehensive runtime environment during instantiation. The constructor assembles multiple subsystems—including context compression, session persistence, and interrupt handling—while caching the system prompt for efficient prefix-caching on Claude models.
Initialization and Runtime State
During initialization, the constructor establishes the LLM endpoint configuration through attributes like model, base_url, api_key, and provider, defaulting to OpenRouter with Claude Opus. It also sets global safety limits via max_iterations and iteration_budget, which cap the total number of tool-calling steps across the entire session.
The initialization sequence dynamically discovers available tools through the tools and valid_tool_names attributes, filtering by enabled or disabled toolsets. It also instantiates the context_compressor for automatic token management and configures optional persistence layers including session_db for SQLite storage and session_log_file for JSON debugging logs.
Key Attributes and Configuration
The class maintains several specialized fields for advanced functionality:
memory_storeandhoncho: Support long-term memory and cross-session user modelingephemeral_system_promptandprefill_messages: Allow prompt fragments injected only at API-call time without polluting session logsinterrupthandling fields (_interrupt_requested,_interrupt_message): Enable external threads (CLI or gateway) to safely abort the loopquiet_mode,verbose_logging,tool_progress_callback: Control UI feedback and logging granularity
The constructor prints a concise status banner (unless quiet_mode is enabled) and caches the final system prompt to optimize performance on Claude models【/cache/repos/github.com/NousResearch/hermes-agent/main/run_agent.py#L142-L170】.
The Agent Loop Implementation in run_conversation()
The run_conversation() method implements the core agent loop—a sophisticated while-loop that iteratively calls the LLM, processes tool calls, manages context windows, and handles interruptions. This method returns the complete message history as a dictionary, unlike the simplified chat() method.
Loop Structure and Guard Conditions
The loop guard ensures safe execution boundaries:
while api_call_count < self.max_iterations and self.iteration_budget.remaining > 0:
This condition appears at line 3033-3036 in run_agent.py【/cache/repos/github.com/NousResearch/hermes-agent/main/run_agent.py#L3033-L3036】, preventing infinite loops and enforcing budget constraints across the entire conversation.
Message Preparation and API Calls
Each iteration begins by checking for interrupt requests, then builds the API payload through the _build_api_kwargs helper method. This process:
- Copies the current message list and embeds any
reasoning_content - Strips unsupported fields for API compatibility
- Prepends the cached system prompt combined with
ephemeral_system_prompt - Applies optional prompt-caching for Claude models
The _interruptible_api_call wrapper manages the actual HTTP call with exponential-backoff retries, handling rate-limit errors and payload-too-large conditions by triggering automatic context compression. Invalid responses trigger retry logic with back-off, while non-retryable client errors abort the loop.
Tool Execution and Context Management
After receiving the LLM response, the loop parses the output to separate normal text, tool calls, and reasoning tags. The _execute_tool_calls helper validates each tool call against valid_tool_names, rejects malformed JSON arguments (with up to 3 retries), and dispatches execution through the central tool registry.
Special handling includes budget refunds for execute_code calls (considered cheap operations) and comprehensive error capture. After tool execution, the loop conditionally triggers _compress_context when token usage approaches the model's limit, summarizing middle turns to maintain the conversation window.
The _persist_session method incrementally writes the session state to both the JSON log file and SQLite database, ensuring durability for debugging and analysis.
Entry Points: chat() vs run_conversation()
The AIAgent class exposes two distinct interfaces for different use cases:
chat(message: str) → str: Provides a simplified one-shot query interface that returns only the final assistant response as a string. This method is ideal for simple integrations where the full conversation history and tool execution details are not required.
run_conversation(...) → dict: Offers the full-featured agent loop that returns the complete message history, tool trajectories, and metadata. This method supports advanced features including interrupt handling, context compression, and session persistence, making it suitable for production deployments requiring observability and control.
Both methods utilize the same underlying loop infrastructure, ensuring consistent behavior across simple and complex usage patterns.
Summary
- The
AIAgentclass inrun_agent.pyencapsulates all runtime state for Hermes Agent sessions, from model configuration to interrupt handling. - Initialization establishes LLM endpoints, tool registries, context compressors, and optional persistence layers (SQLite/JSON).
- The
run_conversation()method implements a robust while-loop guarded bymax_iterationsanditeration_budgetconstraints. - Core loop phases include interrupt detection, API payload construction with prompt caching, exponential-backoff retries, tool validation/execution, and automatic context compression.
- Two entry points serve different needs:
chat()for simple string responses andrun_conversation()for full trajectory access and production features.
Frequently Asked Questions
What is the AIAgent class in Hermes Agent?
The AIAgent class is the core runtime engine defined in run_agent.py that manages a complete agent session. It handles LLM API calls, tool discovery and execution, context window management, interrupt handling, and session persistence, providing both simple (chat) and advanced (run_conversation) interfaces for interacting with the agent.
How does the agent loop handle infinite loops and budget constraints?
The agent loop uses a dual-guard condition: while api_call_count < self.max_iterations and self.iteration_budget.remaining > 0. This ensures the loop terminates after a configurable number of API calls or when the iteration budget is exhausted. Additionally, the loop tracks tool execution costs and refunds iterations for cheap operations like execute_code, optimizing budget utilization across the conversation.
What happens when the conversation context exceeds token limits?
When token usage approaches the model's limit, the _compress_context method automatically summarizes middle turns of the conversation to reduce the context window size. This compression occurs during the loop's context management phase, ensuring the API payload stays within limits while preserving the most recent and relevant conversation history. The system also handles payload-too-large errors from the API by triggering immediate compression and retry with exponential backoff.
How does the AIAgent class support interrupt handling during long-running tasks?
The class maintains interrupt handling fields (_interrupt_requested, _interrupt_message) that external threads (such as CLI interfaces or gateway processes) can set to request safe loop termination. At the start of each iteration, the loop checks for interrupt requests through the _interruptible_api_call wrapper, which raises an InterruptError if requested, allowing the agent to abort gracefully while preserving the current session state in the persistence layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →