How Context Management and Cache Control Optimize Open Agents Calls
Open Agents uses provider-specific cache-control headers and aggressive message compaction to reduce token costs and prevent context window overflow during multi-turn LLM workflows.
Open Agents, developed by Vercel Labs, orchestrates complex LLM workflows that frequently invoke external tools and maintain lengthy conversational histories. Efficient context management and cache control optimize Open Agents calls by minimizing redundant computation and shrinking payload sizes before they reach the model provider.
Implementing Cache Control for Anthropic Models
The addCacheControl Mechanism
In packages/agent/context-management/cache-control.ts, the addCacheControl function detects Anthropic models via isAnthropicModel and injects cacheControl: { type: "ephemeral" } into the final tool or message. This marks the boundary for Anthropic's incremental cache, allowing the provider to store tool results locally and reuse them across subsequent identical calls.
Token Cost Reduction
By targeting the last tool or message in a sequence, cache control prevents redundant round-trips for repeated tool invocations. This optimization is particularly impactful for Anthropic models that support ephemeral cache break-points, significantly cutting token usage in multi-step agent workflows.
Shrinking Context Windows with Aggressive Compaction
The Compaction Pipeline
The packages/agent/context-management/aggressive-compaction-helpers.ts file implements a four-stage pipeline that reduces conversation history size without losing critical workflow state:
indexToolCalls– Maps every tool call to its position in the message array.findPendingCompactionCandidates– Identifies tool calls that are no longer referenced in recent turns.estimateCompactionSavings– Calculates potential token reduction before performing replacement.compactToolData– Replaces full tool inputs and outputs with a minimal placeholder object.
Preserving Workflow Integrity
Compaction replaces verbose tool data with a compact marker like { compacted: true, message: "[compacted]" }. This preserves the structural integrity of the conversation for downstream processing while ensuring the payload stays within model context limits.
Integration Workflow
During agent execution, these optimizations operate in three phases:
Preparation – Developers wrap tool sets or message arrays with addCacheControl before invoking generateText or custom agents. The function validates the model provider and appends cache directives only for supported backends.
Execution – The LLM receives requests where the final tool or message carries ephemeral cache headers. Anthropic's runtime stores these results, enabling cache hits on subsequent identical tool invocations.
Post-processing – After each turn, the aggressive compaction pipeline runs via the helper functions in aggressive-compaction-helpers.ts. It indexes tool calls, identifies stale data, estimates savings, and rewrites the message history to minimize the next request's payload size.
Summary
- Cache control in
packages/agent/context-management/cache-control.tsinjects provider-specific headers likecacheControl: { type: "ephemeral" }for Anthropic models, enabling incremental caching of tool results. - Aggressive compaction in
packages/agent/context-management/aggressive-compaction-helpers.tsusesindexToolCalls,findPendingCompactionCandidates,estimateCompactionSavings, andcompactToolDatato shrink conversation history. - These mechanisms work sequentially: cache control prevents redundant computation during execution, while compaction reduces payload size between turns.
- Together, they minimize token costs and prevent context window overflow in multi-step Open Agents workflows.
Frequently Asked Questions
What is the difference between cache control and aggressive compaction in Open Agents?
Cache control operates at the provider level by marking specific tools or messages with ephemeral cache headers, allowing the LLM provider to store and reuse results. Aggressive compaction operates at the application level by rewriting message history to remove verbose tool data that is no longer needed, reducing the payload size sent to the provider.
Which model providers support the cache control optimization?
Currently, the addCacheControl function specifically detects Anthropic models using isAnthropicModel and injects cacheControl: { type: "ephemeral" } only for those providers. Other providers may ignore these options or handle caching through different mechanisms.
How does aggressive compaction preserve workflow state while reducing tokens?
The compaction pipeline identifies tool calls that are no longer referenced in recent conversation turns and replaces their full input and output objects with a minimal placeholder like { compacted: true, message: "[compacted]" }. This preserves the message structure for downstream processing while eliminating the token-heavy payload details.
When should developers manually invoke compaction helpers versus relying on automatic processing?
The Open Agents runtime automatically applies aggressive compaction after each turn in the agent workflow. Developers typically only need to manually invoke functions like compactToolData when building custom agent loops or when implementing specialized conversation management that bypasses the standard runtime post-processing pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →