DeepAgents Prompt Caching for Anthropic: Reducing LLM Latency with Automatic Middleware

DeepAgents automatically wires AnthropicPromptCachingMiddleware into every agent to cache Claude API responses in an LRU cache, eliminating redundant network calls and reducing latency without manual configuration.

DeepAgents, the LangChain-based framework maintained in the langchain-ai/deepagents repository, implements built-in prompt caching strategies specifically optimized for Anthropic's Claude models. The framework automatically intercepts duplicate prompts across all agent tiers—general-purpose sub-agents, user-defined sub-agents, and the main agent—serving cached responses from memory rather than hitting the API repeatedly.

How DeepAgents Implements Anthropic Prompt Caching

Automatic Middleware Injection

In libs/deepagents/deepagents/graph.py, the create_deep_agent factory function automatically appends AnthropicPromptCachingMiddleware to every agent's middleware stack. This middleware, imported from langchain_anthropic.middleware, maintains an in-memory LRU cache of request-response pairs and sits before the tool-calling layer to intercept duplicate ChatAnthropic invocations.

The middleware is injected at three specific points within the factory:

General-purpose sub-agent (lines 199-203):

gp_middleware = [
    TodoListMiddleware(),
    FilesystemMiddleware(backend=backend),
    create_summarization_middleware(model, backend),
    AnthropicPromptCachingMiddleware(unsupported_model_behavior="ignore"),
    PatchToolCallsMiddleware(),
]

User-provided sub-agents (lines 28-34):

subagent_middleware = [
    TodoListMiddleware(),
    FilesystemMiddleware(backend=backend),
    create_summarization_middleware(subagent_model, backend),
    AnthropicPromptCachingMiddleware(unsupported_model_behavior="ignore"),
    PatchToolCallsMiddleware(),
]

Main agent stack (lines 71-73):

deepagent_middleware.extend([
    FilesystemMiddleware(backend=backend),
    SubAgentMiddleware(...),
    create_summarization_middleware(model, backend),
    AnthropicPromptCachingMiddleware(unsupported_model_behavior="ignore"),
    PatchToolCallsMiddleware(),
])

Provider-Agnostic Safety

The unsupported_model_behavior="ignore" parameter ensures the middleware acts as a no-op when using non-Anthropic models. This design guarantees that the same agent configuration works seamlessly across OpenAI, Google, and other providers without raising runtime errors or impacting performance.

Configuring Prompt Caching Behavior

Default Caching (Zero Configuration)

By default, calling create_deep_agent() enables prompt caching automatically for the lifetime of the Python process:

from deepagents import create_deep_agent

agent = create_deep_agent()  # Anthropic caching enabled automatically

Customizing Cache Size and TTL

To adjust the default LRU cache settings, manually instantiate AnthropicPromptCachingMiddleware with the maxsize (default 128) and ttl parameters:

from deepagents import create_deep_agent
from langchain_anthropic.middleware import AnthropicPromptCachingMiddleware

custom_middleware = AnthropicPromptCachingMiddleware(
    maxsize=512,      # Store up to 512 unique prompts

    ttl=3600,         # Cache entries expire after 1 hour

    unsupported_model_behavior="ignore",
)

agent = create_deep_agent(middleware=[custom_middleware])

Persistent Caching Across Sessions

For caching that survives process restarts, implement a LangGraph BaseCache and pass it via the cache parameter. This propagates to the underlying create_agent call for graph-level persistence:

from deepagents import create_deep_agent
from langgraph.cache.base import BaseCache

# Example with a hypothetical SQLite-backed implementation

persistent_cache = SqlCache(db_path="anthropic_cache.sqlite")
agent = create_deep_agent(cache=persistent_cache)

Disabling Prompt Caching

To disable caching entirely, pass an empty middleware list or configure specific sub-agents to exclude the middleware:


# Disable for the main agent

agent = create_deep_agent(middleware=[])

# Disable for specific sub-agents only

subagents = [
    {
        "name": "no_cache_researcher",
        "model": "anthropic:claude-3-5-sonnet",
        "middleware": [],  # Excludes AnthropicPromptCachingMiddleware

    }
]
agent = create_deep_agent(subagents=subagents)

Summary

  • DeepAgents automatically inserts AnthropicPromptCachingMiddleware into every agent and sub-agent via the create_deep_agent factory in libs/deepagents/deepagents/graph.py, covering lines 28-34, 71-73, and 199-203.
  • The LRU cache stores 128 prompts by default with configurable maxsize and ttl parameters to balance memory usage against hit rates.
  • Provider-agnostic safety is enforced via unsupported_model_behavior="ignore", ensuring non-Anthropic models operate without interference.
  • Persistent caching is supported through LangGraph's BaseCache interface via the cache parameter for cross-session storage.

Frequently Asked Questions

Does DeepAgents cache prompts for providers other than Anthropic?

No. The AnthropicPromptCachingMiddleware specifically handles Anthropic Claude API calls. However, the unsupported_model_behavior="ignore" setting ensures the middleware silently passes through for other providers like OpenAI or Google, allowing mixed-model agent configurations without errors.

What is the default cache size and can it be increased?

The default LRU cache stores 128 unique prompt-response pairs. You can increase this by passing a custom maxsize parameter when manually instantiating AnthropicPromptCachingMiddleware, or decrease it to reduce memory consumption in resource-constrained environments.

Does the cache persist when I restart my Python application?

No, the default implementation uses an in-memory LRU cache that clears when the process exits. To persist cache across restarts, implement a custom BaseCache (such as a SQLite-backed cache) and pass it to create_deep_agent() via the cache parameter, which propagates to the underlying LangGraph agent.

How do I completely disable prompt caching for a specific sub-agent?

Pass an empty middleware list in the sub-agent configuration dictionary when calling create_deep_agent(). This overrides the default middleware stack for that specific sub-agent, excluding AnthropicPromptCachingMiddleware while keeping it enabled for other agents in the graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →