# DeepAgents Prompt Caching for Anthropic: Reducing LLM Latency with Automatic Middleware

> Reduce Anthropic LLM latency with DeepAgents prompt caching. Automatically cache Claude API responses to eliminate redundant calls and speed up your agents without manual setup.

- Repository: [LangChain/deepagents](https://github.com/langchain-ai/deepagents)
- Tags: how-to-guide
- Published: 2026-03-17

---

**DeepAgents automatically wires `AnthropicPromptCachingMiddleware` into every agent to cache Claude API responses in an LRU cache, eliminating redundant network calls and reducing latency without manual configuration.**

DeepAgents, the LangChain-based framework maintained in the `langchain-ai/deepagents` repository, implements built-in prompt caching strategies specifically optimized for Anthropic's Claude models. The framework automatically intercepts duplicate prompts across all agent tiers—general-purpose sub-agents, user-defined sub-agents, and the main agent—serving cached responses from memory rather than hitting the API repeatedly.

## How DeepAgents Implements Anthropic Prompt Caching

### Automatic Middleware Injection

In [`libs/deepagents/deepagents/graph.py`](https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/deepagents/graph.py), the `create_deep_agent` factory function automatically appends `AnthropicPromptCachingMiddleware` to every agent's middleware stack. This middleware, imported from `langchain_anthropic.middleware`, maintains an in-memory LRU cache of request-response pairs and sits before the tool-calling layer to intercept duplicate `ChatAnthropic` invocations.

The middleware is injected at three specific points within the factory:

**General-purpose sub-agent** (lines 199-203):

```python
gp_middleware = [
    TodoListMiddleware(),
    FilesystemMiddleware(backend=backend),
    create_summarization_middleware(model, backend),
    AnthropicPromptCachingMiddleware(unsupported_model_behavior="ignore"),
    PatchToolCallsMiddleware(),
]

```

**User-provided sub-agents** (lines 28-34):

```python
subagent_middleware = [
    TodoListMiddleware(),
    FilesystemMiddleware(backend=backend),
    create_summarization_middleware(subagent_model, backend),
    AnthropicPromptCachingMiddleware(unsupported_model_behavior="ignore"),
    PatchToolCallsMiddleware(),
]

```

**Main agent stack** (lines 71-73):

```python
deepagent_middleware.extend([
    FilesystemMiddleware(backend=backend),
    SubAgentMiddleware(...),
    create_summarization_middleware(model, backend),
    AnthropicPromptCachingMiddleware(unsupported_model_behavior="ignore"),
    PatchToolCallsMiddleware(),
])

```

### Provider-Agnostic Safety

The `unsupported_model_behavior="ignore"` parameter ensures the middleware acts as a no-op when using non-Anthropic models. This design guarantees that the same agent configuration works seamlessly across OpenAI, Google, and other providers without raising runtime errors or impacting performance.

## Configuring Prompt Caching Behavior

### Default Caching (Zero Configuration)

By default, calling `create_deep_agent()` enables prompt caching automatically for the lifetime of the Python process:

```python
from deepagents import create_deep_agent

agent = create_deep_agent()  # Anthropic caching enabled automatically

```

### Customizing Cache Size and TTL

To adjust the default LRU cache settings, manually instantiate `AnthropicPromptCachingMiddleware` with the `maxsize` (default 128) and `ttl` parameters:

```python
from deepagents import create_deep_agent
from langchain_anthropic.middleware import AnthropicPromptCachingMiddleware

custom_middleware = AnthropicPromptCachingMiddleware(
    maxsize=512,      # Store up to 512 unique prompts

    ttl=3600,         # Cache entries expire after 1 hour

    unsupported_model_behavior="ignore",
)

agent = create_deep_agent(middleware=[custom_middleware])

```

### Persistent Caching Across Sessions

For caching that survives process restarts, implement a LangGraph `BaseCache` and pass it via the `cache` parameter. This propagates to the underlying `create_agent` call for graph-level persistence:

```python
from deepagents import create_deep_agent
from langgraph.cache.base import BaseCache

# Example with a hypothetical SQLite-backed implementation

persistent_cache = SqlCache(db_path="anthropic_cache.sqlite")
agent = create_deep_agent(cache=persistent_cache)

```

### Disabling Prompt Caching

To disable caching entirely, pass an empty middleware list or configure specific sub-agents to exclude the middleware:

```python

# Disable for the main agent

agent = create_deep_agent(middleware=[])

# Disable for specific sub-agents only

subagents = [
    {
        "name": "no_cache_researcher",
        "model": "anthropic:claude-3-5-sonnet",
        "middleware": [],  # Excludes AnthropicPromptCachingMiddleware

    }
]
agent = create_deep_agent(subagents=subagents)

```

## Summary

- **DeepAgents automatically inserts `AnthropicPromptCachingMiddleware`** into every agent and sub-agent via the `create_deep_agent` factory in [`libs/deepagents/deepagents/graph.py`](https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/deepagents/graph.py), covering lines 28-34, 71-73, and 199-203.
- **The LRU cache stores 128 prompts by default** with configurable `maxsize` and `ttl` parameters to balance memory usage against hit rates.
- **Provider-agnostic safety** is enforced via `unsupported_model_behavior="ignore"`, ensuring non-Anthropic models operate without interference.
- **Persistent caching** is supported through LangGraph's `BaseCache` interface via the `cache` parameter for cross-session storage.

## Frequently Asked Questions

### Does DeepAgents cache prompts for providers other than Anthropic?

No. The `AnthropicPromptCachingMiddleware` specifically handles Anthropic Claude API calls. However, the `unsupported_model_behavior="ignore"` setting ensures the middleware silently passes through for other providers like OpenAI or Google, allowing mixed-model agent configurations without errors.

### What is the default cache size and can it be increased?

The default LRU cache stores **128** unique prompt-response pairs. You can increase this by passing a custom `maxsize` parameter when manually instantiating `AnthropicPromptCachingMiddleware`, or decrease it to reduce memory consumption in resource-constrained environments.

### Does the cache persist when I restart my Python application?

No, the default implementation uses an in-memory LRU cache that clears when the process exits. To persist cache across restarts, implement a custom `BaseCache` (such as a SQLite-backed cache) and pass it to `create_deep_agent()` via the `cache` parameter, which propagates to the underlying LangGraph agent.

### How do I completely disable prompt caching for a specific sub-agent?

Pass an empty `middleware` list in the sub-agent configuration dictionary when calling `create_deep_agent()`. This overrides the default middleware stack for that specific sub-agent, excluding `AnthropicPromptCachingMiddleware` while keeping it enabled for other agents in the graph.