How Agent Compaction Manages Long Conversations in Cloudflare OS
The agent compaction system in Cloudflare OS automatically condenses conversation history when token usage reaches 85% of the model's input budget, preserving critical state through intelligent summarization while maintaining full canonical history for replay.
Agent compaction is a core subsystem of the Workshop backend in cloudflare/cloudflare-os. It solves the fundamental problem of unbounded context growth in AI-driven chat sessions while ensuring that tool calls, reverts, and pending operations remain intact. By projecting messages into a weighted representation and computing optimal cut points, the system keeps prompts within model limits without losing semantic continuity.
Detecting When to Compact
The compaction trigger has two entry points: automatic threshold detection and explicit user command.
Automatic Threshold Detection
Before each turn, the system compares accumulated context tokens against a configured budget:
// packages/workshop-backend/src/agent-compaction.ts
export function shouldCompactChat(contextTokens: number, inputBudget: number): boolean {
return contextTokens >= inputBudget * COMPACTION_TRIGGER_RATIO;
}
The COMPACTION_TRIGGER_RATIO constant is set to 0.85 (85%). When context tokens hit this threshold, the next turn is flagged for compaction. This provides a safety margin before hard context window limits are reached.
Explicit /compact Command
Users can force immediate compaction regardless of token usage:
export function isCompactionTurn(messages: AiChatMessage[]): boolean {
const last = messages.at(-1);
return last?.type === "slashCommand" && last.request.id.builtin &&
last.request.id.commandId === "compact";
}
This slashCommand message type is defined in packages/workshop-shared/api.ts and processed uniformly with automatic triggers.
Building the Projection Layer
Messages are transformed into a CompactionProjectionMessage structure that enables precise token accounting and boundary placement.
Weighted Message Representation
Each message receives a weight based on serialized length:
// Approximate lines 37-43
function projectionMessageWeight(msg: CompactionProjectionMessage): number {
// Weight derived from JSON-serialized character count
// with role-specific adjustments for tool calls
}
The projection captures:
- Sequence number: Originating position in canonical chat history
canCutflag: Marks valid boundary points where logical records end- Role and content type: Distinguishes user, assistant, and tool messages
Token Estimation Fallback
When the model provider lacks usage metadata, estimateProjectionTokens (lines 45-48) approximates token count from serialized projection length using a linear heuristic.
Finding the Compaction Boundary
The core algorithm walks backward through the projection to locate the optimal cut point.
Backward Accumulation with Budget Target
export function findCompactionBoundary(
projection, inputBudget, contextTokens,
compactedTo = 0, protectedFromSequence?): number {
const tailBudget = inputBudget * 0.30; // 30% retention target
// Walk from newest message backward
let keptTokens = 0;
for (let i = projection.length - 1; i >= 0; i--) {
const msg = projection[i];
keptTokens += msg.weight;
if (keptTokens >= tailBudget && msg.canCut) {
return msg.sequence; // Valid boundary found
}
}
return compactedTo; // Fallback: no compaction
}
Key constraints enforced during this walk:
- Atomic records: Boundary only placed where
canCut === true - Protected sequences: Never cut past
protectedFromSequence(e.g., pending connection requests fromfindProtectedFromSequence, lines 200-216)
Protecting Revert Dependencies
Before finalizing, protectRetainedReverts (lines 54-66) scans retained messages for revert operations. If found, the boundary expands to include the source range of any reverted change, ensuring that replay can resolve change IDs correctly.
Summarizing the Compacted Prefix
The discarded prefix becomes input for a summarization model call that produces a structured hand-off.
Prompt Construction
buildSummaryPrompt (lines 72-108) performs several normalization steps:
- Merges consecutive messages from the same role
- Flattens tool call results into text descriptions
- Strips binary attachments and non-semantic metadata
System Prompt for Summarization
const COMPACTION_SYSTEM_PROMPT = `You are a context summarizer for an AI coding assistant.
Generate a structured hand-off that captures:
- Current task state and objectives
- Key decisions made and their rationale
- Active files, bindings, and proposed changes
- Pending operations requiring attention
Be concise but complete. The summary will replace detailed conversation history.`;
This constant (lines 43-55) ensures consistent summary structure across compaction events.
Creating Compaction Checkpoints
The final phase folds compacted state into a CompactionCheckpoint for future replay.
State Collection
buildCompactionState (lines 72-124) iterates over compacted messages to aggregate:
- Bindings: Variable and resource references
- Pins: Explicitly locked context items
- Proposed changes: Pending edits not yet applied
- Epoch information: Version counters for optimistic concurrency
Checkpoint Storage
The checkpoint stores minimal replay state, while the summary text is appended to the chat log as a regular user message. This separation allows:
- The UI to display "compacted" indicators with full history accessible via paging
- The backend to reconstruct working state from checkpoint plus retained tail
- Diagnostic tools to correlate summaries with their source ranges
Complete Integration Example
Below is the canonical compaction workflow as implemented in the Workshop backend:
import {
shouldCompactChat,
isCompactionTurn,
findCompactionBoundary,
protectRetainedReverts,
buildSummaryPrompt,
buildCompactionState,
} from "@gadgets/workshop-backend/src/agent-compaction";
async function processTurn(
chatMessages: AiChatMessage[],
currentTokens: number,
modelLimits: ModelLimits,
checkpoint: CompactionCheckpoint | null
): Promise<CompactionResult> {
// 1. Decide if compaction is required
const needsCompaction = shouldCompactChat(currentTokens, modelLimits.inputBudget)
|| isCompactionTurn(chatMessages);
if (!needsCompaction) {
return { continueWith: chatMessages, checkpoint };
}
// 2. Build projection and locate boundary
const projection = projectChatMessages(chatMessages);
const rawBoundary = findCompactionBoundary(
projection,
modelLimits.inputBudget,
currentTokens,
checkpoint?.compactedTo ?? 0,
findProtectedFromSequence(chatMessages)
);
// 3. Expand boundary to protect revert dependencies
const safeBoundary = protectRetainedReverts(
rawBoundary,
chatMessages,
checkpoint?.compactedTo ?? 0
);
// 4. Generate summary of discarded prefix
const summaryPrompt = buildSummaryPrompt(
projection,
safeBoundary ?? 0,
modelDescription
);
const summaryText = await callSummarizer(summaryPrompt);
const summaryMessage = createUserMessage(summaryText);
// 5. Construct new checkpoint from compacted state
const newCheckpoint = buildCompactionState(
chatMessages,
safeBoundary ?? 0,
initialBindings,
checkpoint
);
// 6. Return modified message stream with summary injected
const compactedMessages = [
summaryMessage,
...chatMessages.slice(safeBoundary ?? 0)
];
return {
continueWith: compactedMessages,
checkpoint: newCheckpoint
};
}
The helper projectChatMessages (referenced but not in core file) maps AiChatMessage to CompactionProjectionMessage using startsAgentTurn to set canCut boundaries at agent turn boundaries.
Key Source Files
| File | Purpose |
|---|---|
packages/workshop-backend/src/agent-compaction.ts |
Core compaction logic: boundary detection, checkpoint construction, summarization utilities |
packages/workshop-backend/__tests__/agent-compaction.test.ts |
Unit tests for boundary calculation and state folding correctness |
packages/workshop-frontend/src/ChatInterface.compaction.test.ts |
Frontend validation of /compact command handling and summary rendering |
packages/workshop-shared/api.ts |
AiChatMessage type definitions and RPC contracts |
packages/workshop-backend/src/ai-invoke.ts |
Model invocation helpers including zeroUsage() for synthetic messages |
Summary
- Agent compaction triggers at 85% of input budget via
shouldCompactChat, with manual override via/compactcommand detected byisCompactionTurn - Projection weighting converts messages to
CompactionProjectionMessagewithcanCutflags marking valid boundary points - Backward accumulation in
findCompactionBoundarytargets 30% retention budget while respecting protected sequences - Revert protection expands boundaries to include source ranges, preserving replay capability
- Structured summarization via
buildSummaryPromptandCOMPACTION_SYSTEM_PROMPTgenerates concise hand-offs - Checkpoint persistence in
buildCompactionStatemaintains minimal replay state separate from chat log summaries
Frequently Asked Questions
What is the 85% threshold in Cloudflare OS agent compaction?
The COMPACTION_TRIGGER_RATIO = 0.85 constant in agent-compaction.ts defines when automatic compaction activates. When accumulated context tokens reach 85% of the model's input budget, the next turn is processed through the compaction pipeline. This threshold provides headroom for the final turn's tokens while avoiding emergency truncation.
How does agent compaction preserve tool call integrity?
The canCut flag in CompactionProjectionMessage ensures boundaries only occur at complete logical records. Agent turns (including tool calls and their results) are marked atomic via startsAgentTurn. The backward walk in findCompactionBoundary explicitly checks canCut before accepting a boundary, preventing splits mid-transaction.
Can users manually trigger compaction in Cloudflare OS?
Yes. The /compact slash command forces immediate compaction regardless of token usage. isCompactionTurn detects this by checking the last message for type === "slashCommand" with commandId === "compact". This is processed identically to automatic triggers, producing the same checkpoint and summary flow.
What happens to reverted changes during compaction?
protectRetainedReverts scans the retained message tail for revert operations. If any revert's source range falls outside the proposed boundary, the boundary expands to include that range. This ensures that replay can resolve change IDs referenced by retained revert messages, maintaining causal consistency across compaction events.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →