How Agent Compaction Manages Long Conversations in Cloudflare OS

The agent compaction system in Cloudflare OS automatically condenses conversation history when token usage reaches 85% of the model's input budget, preserving critical state through intelligent summarization while maintaining full canonical history for replay.

Agent compaction is a core subsystem of the Workshop backend in cloudflare/cloudflare-os. It solves the fundamental problem of unbounded context growth in AI-driven chat sessions while ensuring that tool calls, reverts, and pending operations remain intact. By projecting messages into a weighted representation and computing optimal cut points, the system keeps prompts within model limits without losing semantic continuity.


Detecting When to Compact

The compaction trigger has two entry points: automatic threshold detection and explicit user command.

Automatic Threshold Detection

Before each turn, the system compares accumulated context tokens against a configured budget:

// packages/workshop-backend/src/agent-compaction.ts
export function shouldCompactChat(contextTokens: number, inputBudget: number): boolean {
  return contextTokens >= inputBudget * COMPACTION_TRIGGER_RATIO;
}

The COMPACTION_TRIGGER_RATIO constant is set to 0.85 (85%). When context tokens hit this threshold, the next turn is flagged for compaction. This provides a safety margin before hard context window limits are reached.

Explicit /compact Command

Users can force immediate compaction regardless of token usage:

export function isCompactionTurn(messages: AiChatMessage[]): boolean {
  const last = messages.at(-1);
  return last?.type === "slashCommand" && last.request.id.builtin && 
         last.request.id.commandId === "compact";
}

This slashCommand message type is defined in packages/workshop-shared/api.ts and processed uniformly with automatic triggers.


Building the Projection Layer

Messages are transformed into a CompactionProjectionMessage structure that enables precise token accounting and boundary placement.

Weighted Message Representation

Each message receives a weight based on serialized length:

// Approximate lines 37-43
function projectionMessageWeight(msg: CompactionProjectionMessage): number {
  // Weight derived from JSON-serialized character count
  // with role-specific adjustments for tool calls
}

The projection captures:

  • Sequence number: Originating position in canonical chat history
  • canCut flag: Marks valid boundary points where logical records end
  • Role and content type: Distinguishes user, assistant, and tool messages

Token Estimation Fallback

When the model provider lacks usage metadata, estimateProjectionTokens (lines 45-48) approximates token count from serialized projection length using a linear heuristic.


Finding the Compaction Boundary

The core algorithm walks backward through the projection to locate the optimal cut point.

Backward Accumulation with Budget Target

export function findCompactionBoundary(
    projection, inputBudget, contextTokens,
    compactedTo = 0, protectedFromSequence?): number {
  const tailBudget = inputBudget * 0.30; // 30% retention target
  
  // Walk from newest message backward
  let keptTokens = 0;
  for (let i = projection.length - 1; i >= 0; i--) {
    const msg = projection[i];
    keptTokens += msg.weight;
    
    if (keptTokens >= tailBudget && msg.canCut) {
      return msg.sequence; // Valid boundary found
    }
  }
  return compactedTo; // Fallback: no compaction
}

Key constraints enforced during this walk:

  • Atomic records: Boundary only placed where canCut === true
  • Protected sequences: Never cut past protectedFromSequence (e.g., pending connection requests from findProtectedFromSequence, lines 200-216)

Protecting Revert Dependencies

Before finalizing, protectRetainedReverts (lines 54-66) scans retained messages for revert operations. If found, the boundary expands to include the source range of any reverted change, ensuring that replay can resolve change IDs correctly.


Summarizing the Compacted Prefix

The discarded prefix becomes input for a summarization model call that produces a structured hand-off.

Prompt Construction

buildSummaryPrompt (lines 72-108) performs several normalization steps:

  • Merges consecutive messages from the same role
  • Flattens tool call results into text descriptions
  • Strips binary attachments and non-semantic metadata

System Prompt for Summarization

const COMPACTION_SYSTEM_PROMPT = `You are a context summarizer for an AI coding assistant.
Generate a structured hand-off that captures:
- Current task state and objectives
- Key decisions made and their rationale
- Active files, bindings, and proposed changes
- Pending operations requiring attention

Be concise but complete. The summary will replace detailed conversation history.`;

This constant (lines 43-55) ensures consistent summary structure across compaction events.


Creating Compaction Checkpoints

The final phase folds compacted state into a CompactionCheckpoint for future replay.

State Collection

buildCompactionState (lines 72-124) iterates over compacted messages to aggregate:

  • Bindings: Variable and resource references
  • Pins: Explicitly locked context items
  • Proposed changes: Pending edits not yet applied
  • Epoch information: Version counters for optimistic concurrency

Checkpoint Storage

The checkpoint stores minimal replay state, while the summary text is appended to the chat log as a regular user message. This separation allows:

  • The UI to display "compacted" indicators with full history accessible via paging
  • The backend to reconstruct working state from checkpoint plus retained tail
  • Diagnostic tools to correlate summaries with their source ranges

Complete Integration Example

Below is the canonical compaction workflow as implemented in the Workshop backend:

import {
  shouldCompactChat,
  isCompactionTurn,
  findCompactionBoundary,
  protectRetainedReverts,
  buildSummaryPrompt,
  buildCompactionState,
} from "@gadgets/workshop-backend/src/agent-compaction";

async function processTurn(
  chatMessages: AiChatMessage[],
  currentTokens: number,
  modelLimits: ModelLimits,
  checkpoint: CompactionCheckpoint | null
): Promise<CompactionResult> {
  
  // 1. Decide if compaction is required
  const needsCompaction = shouldCompactChat(currentTokens, modelLimits.inputBudget) 
                       || isCompactionTurn(chatMessages);
  
  if (!needsCompaction) {
    return { continueWith: chatMessages, checkpoint };
  }
  
  // 2. Build projection and locate boundary
  const projection = projectChatMessages(chatMessages);
  const rawBoundary = findCompactionBoundary(
    projection,
    modelLimits.inputBudget,
    currentTokens,
    checkpoint?.compactedTo ?? 0,
    findProtectedFromSequence(chatMessages)
  );
  
  // 3. Expand boundary to protect revert dependencies
  const safeBoundary = protectRetainedReverts(
    rawBoundary, 
    chatMessages, 
    checkpoint?.compactedTo ?? 0
  );
  
  // 4. Generate summary of discarded prefix
  const summaryPrompt = buildSummaryPrompt(
    projection,
    safeBoundary ?? 0,
    modelDescription
  );
  const summaryText = await callSummarizer(summaryPrompt);
  const summaryMessage = createUserMessage(summaryText);
  
  // 5. Construct new checkpoint from compacted state
  const newCheckpoint = buildCompactionState(
    chatMessages,
    safeBoundary ?? 0,
    initialBindings,
    checkpoint
  );
  
  // 6. Return modified message stream with summary injected
  const compactedMessages = [
    summaryMessage,
    ...chatMessages.slice(safeBoundary ?? 0)
  ];
  
  return {
    continueWith: compactedMessages,
    checkpoint: newCheckpoint
  };
}

The helper projectChatMessages (referenced but not in core file) maps AiChatMessage to CompactionProjectionMessage using startsAgentTurn to set canCut boundaries at agent turn boundaries.


Key Source Files

File Purpose
packages/workshop-backend/src/agent-compaction.ts Core compaction logic: boundary detection, checkpoint construction, summarization utilities
packages/workshop-backend/__tests__/agent-compaction.test.ts Unit tests for boundary calculation and state folding correctness
packages/workshop-frontend/src/ChatInterface.compaction.test.ts Frontend validation of /compact command handling and summary rendering
packages/workshop-shared/api.ts AiChatMessage type definitions and RPC contracts
packages/workshop-backend/src/ai-invoke.ts Model invocation helpers including zeroUsage() for synthetic messages

Summary

  • Agent compaction triggers at 85% of input budget via shouldCompactChat, with manual override via /compact command detected by isCompactionTurn
  • Projection weighting converts messages to CompactionProjectionMessage with canCut flags marking valid boundary points
  • Backward accumulation in findCompactionBoundary targets 30% retention budget while respecting protected sequences
  • Revert protection expands boundaries to include source ranges, preserving replay capability
  • Structured summarization via buildSummaryPrompt and COMPACTION_SYSTEM_PROMPT generates concise hand-offs
  • Checkpoint persistence in buildCompactionState maintains minimal replay state separate from chat log summaries

Frequently Asked Questions

What is the 85% threshold in Cloudflare OS agent compaction?

The COMPACTION_TRIGGER_RATIO = 0.85 constant in agent-compaction.ts defines when automatic compaction activates. When accumulated context tokens reach 85% of the model's input budget, the next turn is processed through the compaction pipeline. This threshold provides headroom for the final turn's tokens while avoiding emergency truncation.

How does agent compaction preserve tool call integrity?

The canCut flag in CompactionProjectionMessage ensures boundaries only occur at complete logical records. Agent turns (including tool calls and their results) are marked atomic via startsAgentTurn. The backward walk in findCompactionBoundary explicitly checks canCut before accepting a boundary, preventing splits mid-transaction.

Can users manually trigger compaction in Cloudflare OS?

Yes. The /compact slash command forces immediate compaction regardless of token usage. isCompactionTurn detects this by checking the last message for type === "slashCommand" with commandId === "compact". This is processed identically to automatic triggers, producing the same checkpoint and summary flow.

What happens to reverted changes during compaction?

protectRetainedReverts scans the retained message tail for revert operations. If any revert's source range falls outside the proposed boundary, the boundary expands to include that range. This ensures that replay can resolve change IDs referenced by retained revert messages, maintaining causal consistency across compaction events.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →