# How Hermes Agent Implements Prompt Caching with Anthropic's cache_control

> Discover how Hermes Agent slashes API costs by up to 75% using Anthropic's cache_control for efficient prompt caching in multi-turn conversations. Learn about prompt caching.

- Repository: [Nous Research/hermes-agent](https://github.com/NousResearch/hermes-agent)
- Tags: how-to-guide
- Published: 2026-03-09

---

**Hermes Agent reduces API token costs by up to 75% in multi-turn conversations by automatically injecting `cache_control` breakpoints into Anthropic API requests, caching the system prompt and last three messages for reuse.**

The `NousResearch/hermes-agent` repository implements a sophisticated **Hermes Agent prompt caching** system that leverages Anthropic's `cache_control` parameter to minimize redundant token processing. By strategically placing ephemeral cache markers within message payloads, the agent ensures that static conversation context—such as system instructions and recent dialogue history—persists across API calls without incurring repeated token charges.

## Understanding Anthropic's cache_control Parameter

Anthropic's API accepts a `cache_control` object on any message or text block to enable **prompt caching**. When the `type` field is set to **`ephemeral`**, the model stores that specific content fragment in a short-lived cache for subsequent reuse within the same conversation session.

The optional `ttl` (time-to-live) field extends cache persistence beyond the default window. Valid values include `"5m"` (five minutes), `"1h"` (one hour), or other duration strings accepted by the Anthropic API. Without explicit TTL configuration, cached fragments expire at the end of the immediate request cycle.

## The system_and_3 Caching Strategy in Hermes Agent

Hermes Agent implements the **`system_and_3`** strategy defined in [`agent/prompt_caching.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/prompt_caching.py). This approach places up to **four cache breakpoints** per API request, balancing cache hit probability against API payload complexity.

### Breakpoint Placement Rules

The strategy follows a deterministic priority order for marker placement:

| Breakpoint | Target Message | Position |
|------------|----------------|----------|
| 1️⃣ | System prompt | First message (`role="system"`) |
| 2️⃣ | Third most recent non-system message | Last three non-system entries |
| 3️⃣ | Second most recent non-system message | Last three non-system entries |
| 4️⃣ | Most recent non-system message | Last three non-system entries |

If fewer than four eligible messages exist, the system applies markers to all available candidates. Non-system messages include `user`, `assistant`, and `tool` roles.

### Message Format Handling

The `_apply_cache_marker` helper function in [`agent/prompt_caching.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/prompt_caching.py) handles three distinct content formats to ensure Anthropic API compatibility:

1. **String content**: Plain text strings are converted into a single text block object: `{"type": "text", "text": "...", "cache_control": {"type": "ephemeral"}}`

2. **Block list content**: When `content` is already a list of blocks (text, image, etc.), the `cache_control` marker attaches to the **last block** in the list—the only position Anthropic's API currently recognizes for caching directives.

3. **Tool messages**: Messages with `role="tool"` receive the `cache_control` object at the top level of the message dictionary rather than nested within content blocks.

All transformations return **deep copies** of the message list to prevent mutation of the original conversation history.

## Implementation Details in the Codebase

### Core Caching Functions

The caching logic resides in [`agent/prompt_caching.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/prompt_caching.py), which exports two primary functions:

- **`_apply_cache_marker(message, ttl)`**: A pure function that returns a new message dictionary with the appropriate `cache_control` structure injected based on content format. It accepts an optional `ttl` parameter (defaulting to `"5m"`) to customize cache duration.

- **`apply_anthropic_cache_control(messages, cache_ttl="5m")`**: The main entry point that implements the `system_and_3` strategy. It iterates through the message list, identifies eligible breakpoints (system + last three non-system), and delegates marker application to `_apply_cache_marker`.

### Runtime Integration

The agent activates caching during the request preparation phase in [`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py) (approximately lines 3115-3121):

```python
if self._use_prompt_caching:
    api_messages = apply_anthropic_cache_control(
        api_messages, cache_ttl=self._cache_ttl
    )

```

The `_use_prompt_caching` boolean automatically enables when:
- The model provider is Anthropic (e.g., `claude-3-opus-20240229`)
- The base URL points to OpenRouter (which proxies Anthropic models with caching support)

Users can override the default **5-minute TTL** by passing a `cache_ttl` argument during agent initialization (e.g., `"1h"` for one-hour persistence).

## Practical Usage Examples

### Direct API Usage

For custom implementations outside the agent framework, import the caching utility directly:

```python
from agent.prompt_caching import apply_anthropic_cache_control

messages = [
    {"role": "system", "content": "You are a helpful coding assistant."},
    {"role": "user", "content": "Write a Python function to parse JSON."},
    {"role": "assistant", "content": "Here's a solution using the json module..."},
    {"role": "user", "content": "Add error handling for malformed JSON."}
]

# Apply system_and_3 strategy with 5-minute cache

cached_messages = apply_anthropic_cache_control(messages, cache_ttl="5m")

```

The resulting `cached_messages` contains `cache_control` markers on the system prompt and the last three non-system messages, ready for transmission to Anthropic's API.

### Automatic Caching in Agent Workflows

When using the high-level agent interface, caching requires no manual intervention:

```python
from hermes_agent import AIAgent

agent = AIAgent(
    model="anthropic/claude-3-opus-20240229",
    max_iterations=10,
    cache_ttl="5m"  # Optional: customize TTL

)

# The agent automatically caches prompts across turns

response = agent.chat("Explain the architecture of transformers.")
follow_up = agent.chat("How does attention work specifically?")

```

The agent detects the Claude model and OpenRouter endpoint, enabling `_use_prompt_caching` automatically and injecting breakpoints before each API call.

### Customizing Cache TTL

Extend cache persistence for long-running sessions or background tasks:

```python

# One-hour cache for extended workflows

cached = apply_anthropic_cache_control(messages, cache_ttl="1h")

# Or via agent configuration

agent = AIAgent(
    model="anthropic/claude-3-sonnet-20240229",
    cache_ttl="1h"
)

```

Longer TTL values reduce cache misses in conversations with pauses between turns, though they consume cache quota for extended periods.

## Performance Impact

According to the implementation in `NousResearch/hermes-agent`, the **system_and_3** strategy yields approximately **75% token cost reduction** for long multi-turn conversations. By caching the system prompt (often thousands of tokens in agent contexts) and the recent conversation history, subsequent API calls only process new user input and generate fresh completions, while reusing cached context at a fraction of the standard token price.

## Summary

- **Hermes Agent prompt caching** leverages Anthropic's `cache_control` parameter with `type: ephemeral` to store reusable prompt fragments.
- The **`system_and_3`** strategy places up to four breakpoints: one on the system prompt and three on the most recent non-system messages.
- The **`_apply_cache_marker`** helper in [`agent/prompt_caching.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/prompt_caching.py) handles string-to-block conversion and attaches markers to the last block in multi-block messages.
- Caching activates automatically for Claude models on OpenRouter in [`run_agent.py`](https://github.com/NousResearch/hermes-agent/blob/main/run_agent.py), with a default **5-minute TTL** configurable via `cache_ttl`.
- Proper implementation reduces token costs by approximately **75%** in sustained conversations.

## Frequently Asked Questions

### How does Hermes Agent decide which messages to cache?

Hermes Agent follows the **`system_and_3`** strategy defined in [`agent/prompt_caching.py`](https://github.com/NousResearch/hermes-agent/blob/main/agent/prompt_caching.py). It always attempts to cache the system prompt first, then applies cache markers to the last three non-system messages (user, assistant, or tool roles). If fewer than four total messages exist, it caches all eligible messages. This approach balances cache hit probability with API payload efficiency.

### What is the default cache duration, and how can I change it?

The default **TTL (time-to-live)** is **`"5m"`** (five minutes), as specified in the `apply_anthropic_cache_control` function signature. You can extend this duration by passing a different string to the `cache_ttl` parameter, such as `"1h"` for one hour or `"30m"` for thirty minutes. This is useful for long-running agent workflows where conversations may pause between turns.

### Does prompt caching work with all AI models in Hermes Agent?

No, **prompt caching is specific to Anthropic models** (such as Claude 3 Opus or Claude 3 Sonnet) accessed through the OpenRouter API endpoint. The agent automatically detects eligible configurations by checking `_use_prompt_caching`, which enables only when the model provider is Anthropic and the base URL points to OpenRouter. Other model providers (OpenAI, Google, etc.) do not currently support this specific caching mechanism.

### How much does prompt caching reduce API costs?

According to the implementation in `NousResearch/hermes-agent`, the **system_and_3** caching strategy reduces token costs by approximately **75%** for long, multi-turn conversations. This significant reduction occurs because the system prompt (often containing extensive instructions or context) and the recent conversation history are stored in Anthropic's ephemeral cache and reused across API calls, so you only pay for processing new input tokens and generated output tokens on subsequent requests.