# How Context Management and Cache Control Optimize Open Agents Calls

> Discover how Open Agents optimizes LLM calls with context management and cache control, reducing token costs and preventing overflow for efficient multi-turn workflows.

- Repository: [Vercel Labs/open-agents](https://github.com/vercel-labs/open-agents)
- Tags: performance
- Published: 2026-04-16

---

**Open Agents uses provider-specific cache-control headers and aggressive message compaction to reduce token costs and prevent context window overflow during multi-turn LLM workflows.**

Open Agents, developed by Vercel Labs, orchestrates complex LLM workflows that frequently invoke external tools and maintain lengthy conversational histories. Efficient context management and cache control optimize Open Agents calls by minimizing redundant computation and shrinking payload sizes before they reach the model provider.

## Implementing Cache Control for Anthropic Models

### The addCacheControl Mechanism

In [`packages/agent/context-management/cache-control.ts`](https://github.com/vercel-labs/open-agents/blob/main/packages/agent/context-management/cache-control.ts), the `addCacheControl` function detects Anthropic models via `isAnthropicModel` and injects `cacheControl: { type: "ephemeral" }` into the final tool or message. This marks the boundary for Anthropic's incremental cache, allowing the provider to store tool results locally and reuse them across subsequent identical calls.

### Token Cost Reduction

By targeting the last tool or message in a sequence, cache control prevents redundant round-trips for repeated tool invocations. This optimization is particularly impactful for Anthropic models that support ephemeral cache break-points, significantly cutting token usage in multi-step agent workflows.

## Shrinking Context Windows with Aggressive Compaction

### The Compaction Pipeline

The [`packages/agent/context-management/aggressive-compaction-helpers.ts`](https://github.com/vercel-labs/open-agents/blob/main/packages/agent/context-management/aggressive-compaction-helpers.ts) file implements a four-stage pipeline that reduces conversation history size without losing critical workflow state:

1. **`indexToolCalls`** – Maps every tool call to its position in the message array.
2. **`findPendingCompactionCandidates`** – Identifies tool calls that are no longer referenced in recent turns.
3. **`estimateCompactionSavings`** – Calculates potential token reduction before performing replacement.
4. **`compactToolData`** – Replaces full tool inputs and outputs with a minimal placeholder object.

### Preserving Workflow Integrity

Compaction replaces verbose tool data with a compact marker like `{ compacted: true, message: "[compacted]" }`. This preserves the structural integrity of the conversation for downstream processing while ensuring the payload stays within model context limits.

## Integration Workflow

During agent execution, these optimizations operate in three phases:

**Preparation** – Developers wrap tool sets or message arrays with `addCacheControl` before invoking `generateText` or custom agents. The function validates the model provider and appends cache directives only for supported backends.

**Execution** – The LLM receives requests where the final tool or message carries ephemeral cache headers. Anthropic's runtime stores these results, enabling cache hits on subsequent identical tool invocations.

**Post-processing** – After each turn, the aggressive compaction pipeline runs via the helper functions in [`aggressive-compaction-helpers.ts`](https://github.com/vercel-labs/open-agents/blob/main/aggressive-compaction-helpers.ts). It indexes tool calls, identifies stale data, estimates savings, and rewrites the message history to minimize the next request's payload size.

## Summary

- **Cache control** in [`packages/agent/context-management/cache-control.ts`](https://github.com/vercel-labs/open-agents/blob/main/packages/agent/context-management/cache-control.ts) injects provider-specific headers like `cacheControl: { type: "ephemeral" }` for Anthropic models, enabling incremental caching of tool results.
- **Aggressive compaction** in [`packages/agent/context-management/aggressive-compaction-helpers.ts`](https://github.com/vercel-labs/open-agents/blob/main/packages/agent/context-management/aggressive-compaction-helpers.ts) uses `indexToolCalls`, `findPendingCompactionCandidates`, `estimateCompactionSavings`, and `compactToolData` to shrink conversation history.
- These mechanisms work sequentially: cache control prevents redundant computation during execution, while compaction reduces payload size between turns.
- Together, they minimize token costs and prevent context window overflow in multi-step Open Agents workflows.

## Frequently Asked Questions

### What is the difference between cache control and aggressive compaction in Open Agents?

Cache control operates at the provider level by marking specific tools or messages with ephemeral cache headers, allowing the LLM provider to store and reuse results. Aggressive compaction operates at the application level by rewriting message history to remove verbose tool data that is no longer needed, reducing the payload size sent to the provider.

### Which model providers support the cache control optimization?

Currently, the `addCacheControl` function specifically detects Anthropic models using `isAnthropicModel` and injects `cacheControl: { type: "ephemeral" }` only for those providers. Other providers may ignore these options or handle caching through different mechanisms.

### How does aggressive compaction preserve workflow state while reducing tokens?

The compaction pipeline identifies tool calls that are no longer referenced in recent conversation turns and replaces their full input and output objects with a minimal placeholder like `{ compacted: true, message: "[compacted]" }`. This preserves the message structure for downstream processing while eliminating the token-heavy payload details.

### When should developers manually invoke compaction helpers versus relying on automatic processing?

The Open Agents runtime automatically applies aggressive compaction after each turn in the agent workflow. Developers typically only need to manually invoke functions like `compactToolData` when building custom agent loops or when implementing specialized conversation management that bypasses the standard runtime post-processing pipeline.