How Hermes Agent Implements Prompt Caching with Anthropic's cache_control
Hermes Agent reduces API token costs by up to 75% in multi-turn conversations by automatically injecting cache_control breakpoints into Anthropic API requests, caching the system prompt and last three messages for reuse.
The NousResearch/hermes-agent repository implements a sophisticated Hermes Agent prompt caching system that leverages Anthropic's cache_control parameter to minimize redundant token processing. By strategically placing ephemeral cache markers within message payloads, the agent ensures that static conversation context—such as system instructions and recent dialogue history—persists across API calls without incurring repeated token charges.
Understanding Anthropic's cache_control Parameter
Anthropic's API accepts a cache_control object on any message or text block to enable prompt caching. When the type field is set to ephemeral, the model stores that specific content fragment in a short-lived cache for subsequent reuse within the same conversation session.
The optional ttl (time-to-live) field extends cache persistence beyond the default window. Valid values include "5m" (five minutes), "1h" (one hour), or other duration strings accepted by the Anthropic API. Without explicit TTL configuration, cached fragments expire at the end of the immediate request cycle.
The system_and_3 Caching Strategy in Hermes Agent
Hermes Agent implements the system_and_3 strategy defined in agent/prompt_caching.py. This approach places up to four cache breakpoints per API request, balancing cache hit probability against API payload complexity.
Breakpoint Placement Rules
The strategy follows a deterministic priority order for marker placement:
| Breakpoint | Target Message | Position |
|---|---|---|
| 1️⃣ | System prompt | First message (role="system") |
| 2️⃣ | Third most recent non-system message | Last three non-system entries |
| 3️⃣ | Second most recent non-system message | Last three non-system entries |
| 4️⃣ | Most recent non-system message | Last three non-system entries |
If fewer than four eligible messages exist, the system applies markers to all available candidates. Non-system messages include user, assistant, and tool roles.
Message Format Handling
The _apply_cache_marker helper function in agent/prompt_caching.py handles three distinct content formats to ensure Anthropic API compatibility:
-
String content: Plain text strings are converted into a single text block object:
{"type": "text", "text": "...", "cache_control": {"type": "ephemeral"}} -
Block list content: When
contentis already a list of blocks (text, image, etc.), thecache_controlmarker attaches to the last block in the list—the only position Anthropic's API currently recognizes for caching directives. -
Tool messages: Messages with
role="tool"receive thecache_controlobject at the top level of the message dictionary rather than nested within content blocks.
All transformations return deep copies of the message list to prevent mutation of the original conversation history.
Implementation Details in the Codebase
Core Caching Functions
The caching logic resides in agent/prompt_caching.py, which exports two primary functions:
-
_apply_cache_marker(message, ttl): A pure function that returns a new message dictionary with the appropriatecache_controlstructure injected based on content format. It accepts an optionalttlparameter (defaulting to"5m") to customize cache duration. -
apply_anthropic_cache_control(messages, cache_ttl="5m"): The main entry point that implements thesystem_and_3strategy. It iterates through the message list, identifies eligible breakpoints (system + last three non-system), and delegates marker application to_apply_cache_marker.
Runtime Integration
The agent activates caching during the request preparation phase in run_agent.py (approximately lines 3115-3121):
if self._use_prompt_caching:
api_messages = apply_anthropic_cache_control(
api_messages, cache_ttl=self._cache_ttl
)
The _use_prompt_caching boolean automatically enables when:
- The model provider is Anthropic (e.g.,
claude-3-opus-20240229) - The base URL points to OpenRouter (which proxies Anthropic models with caching support)
Users can override the default 5-minute TTL by passing a cache_ttl argument during agent initialization (e.g., "1h" for one-hour persistence).
Practical Usage Examples
Direct API Usage
For custom implementations outside the agent framework, import the caching utility directly:
from agent.prompt_caching import apply_anthropic_cache_control
messages = [
{"role": "system", "content": "You are a helpful coding assistant."},
{"role": "user", "content": "Write a Python function to parse JSON."},
{"role": "assistant", "content": "Here's a solution using the json module..."},
{"role": "user", "content": "Add error handling for malformed JSON."}
]
# Apply system_and_3 strategy with 5-minute cache
cached_messages = apply_anthropic_cache_control(messages, cache_ttl="5m")
The resulting cached_messages contains cache_control markers on the system prompt and the last three non-system messages, ready for transmission to Anthropic's API.
Automatic Caching in Agent Workflows
When using the high-level agent interface, caching requires no manual intervention:
from hermes_agent import AIAgent
agent = AIAgent(
model="anthropic/claude-3-opus-20240229",
max_iterations=10,
cache_ttl="5m" # Optional: customize TTL
)
# The agent automatically caches prompts across turns
response = agent.chat("Explain the architecture of transformers.")
follow_up = agent.chat("How does attention work specifically?")
The agent detects the Claude model and OpenRouter endpoint, enabling _use_prompt_caching automatically and injecting breakpoints before each API call.
Customizing Cache TTL
Extend cache persistence for long-running sessions or background tasks:
# One-hour cache for extended workflows
cached = apply_anthropic_cache_control(messages, cache_ttl="1h")
# Or via agent configuration
agent = AIAgent(
model="anthropic/claude-3-sonnet-20240229",
cache_ttl="1h"
)
Longer TTL values reduce cache misses in conversations with pauses between turns, though they consume cache quota for extended periods.
Performance Impact
According to the implementation in NousResearch/hermes-agent, the system_and_3 strategy yields approximately 75% token cost reduction for long multi-turn conversations. By caching the system prompt (often thousands of tokens in agent contexts) and the recent conversation history, subsequent API calls only process new user input and generate fresh completions, while reusing cached context at a fraction of the standard token price.
Summary
- Hermes Agent prompt caching leverages Anthropic's
cache_controlparameter withtype: ephemeralto store reusable prompt fragments. - The
system_and_3strategy places up to four breakpoints: one on the system prompt and three on the most recent non-system messages. - The
_apply_cache_markerhelper inagent/prompt_caching.pyhandles string-to-block conversion and attaches markers to the last block in multi-block messages. - Caching activates automatically for Claude models on OpenRouter in
run_agent.py, with a default 5-minute TTL configurable viacache_ttl. - Proper implementation reduces token costs by approximately 75% in sustained conversations.
Frequently Asked Questions
How does Hermes Agent decide which messages to cache?
Hermes Agent follows the system_and_3 strategy defined in agent/prompt_caching.py. It always attempts to cache the system prompt first, then applies cache markers to the last three non-system messages (user, assistant, or tool roles). If fewer than four total messages exist, it caches all eligible messages. This approach balances cache hit probability with API payload efficiency.
What is the default cache duration, and how can I change it?
The default TTL (time-to-live) is "5m" (five minutes), as specified in the apply_anthropic_cache_control function signature. You can extend this duration by passing a different string to the cache_ttl parameter, such as "1h" for one hour or "30m" for thirty minutes. This is useful for long-running agent workflows where conversations may pause between turns.
Does prompt caching work with all AI models in Hermes Agent?
No, prompt caching is specific to Anthropic models (such as Claude 3 Opus or Claude 3 Sonnet) accessed through the OpenRouter API endpoint. The agent automatically detects eligible configurations by checking _use_prompt_caching, which enables only when the model provider is Anthropic and the base URL points to OpenRouter. Other model providers (OpenAI, Google, etc.) do not currently support this specific caching mechanism.
How much does prompt caching reduce API costs?
According to the implementation in NousResearch/hermes-agent, the system_and_3 caching strategy reduces token costs by approximately 75% for long, multi-turn conversations. This significant reduction occurs because the system prompt (often containing extensive instructions or context) and the recent conversation history are stored in Anthropic's ephemeral cache and reused across API calls, so you only pay for processing new input tokens and generated output tokens on subsequent requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →