# How Agent Compaction Manages Long Conversations in Cloudflare OS

> Discover how Cloudflare OS agent compaction manages long conversations. It intelligently summarizes history to stay within token limits while preserving full state for replay.

- Repository: [Cloudflare/cloudflare-os](https://github.com/cloudflare/cloudflare-os)
- Tags: internals
- Published: 2026-09-05

---

**The agent compaction system in Cloudflare OS automatically condenses conversation history when token usage reaches 85% of the model's input budget, preserving critical state through intelligent summarization while maintaining full canonical history for replay.**

Agent compaction is a core subsystem of the Workshop backend in [cloudflare/cloudflare-os](https://github.com/cloudflare/cloudflare-os). It solves the fundamental problem of unbounded context growth in AI-driven chat sessions while ensuring that tool calls, reverts, and pending operations remain intact. By projecting messages into a weighted representation and computing optimal cut points, the system keeps prompts within model limits without losing semantic continuity.

---

## Detecting When to Compact

The compaction trigger has two entry points: automatic threshold detection and explicit user command.

### Automatic Threshold Detection

Before each turn, the system compares accumulated context tokens against a configured budget:

```ts
// packages/workshop-backend/src/agent-compaction.ts
export function shouldCompactChat(contextTokens: number, inputBudget: number): boolean {
  return contextTokens >= inputBudget * COMPACTION_TRIGGER_RATIO;
}

```

The `COMPACTION_TRIGGER_RATIO` constant is set to `0.85` (85%). When context tokens hit this threshold, the next turn is flagged for compaction. This provides a safety margin before hard context window limits are reached.

### Explicit `/compact` Command

Users can force immediate compaction regardless of token usage:

```ts
export function isCompactionTurn(messages: AiChatMessage[]): boolean {
  const last = messages.at(-1);
  return last?.type === "slashCommand" && last.request.id.builtin && 
         last.request.id.commandId === "compact";
}

```

This `slashCommand` message type is defined in [`packages/workshop-shared/api.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/workshop-shared/api.ts) and processed uniformly with automatic triggers.

---

## Building the Projection Layer

Messages are transformed into a **`CompactionProjectionMessage`** structure that enables precise token accounting and boundary placement.

### Weighted Message Representation

Each message receives a weight based on serialized length:

```ts
// Approximate lines 37-43
function projectionMessageWeight(msg: CompactionProjectionMessage): number {
  // Weight derived from JSON-serialized character count
  // with role-specific adjustments for tool calls
}

```

The projection captures:

- **Sequence number**: Originating position in canonical chat history
- **`canCut` flag**: Marks valid boundary points where logical records end
- **Role and content type**: Distinguishes user, assistant, and tool messages

### Token Estimation Fallback

When the model provider lacks usage metadata, `estimateProjectionTokens` (lines 45-48) approximates token count from serialized projection length using a linear heuristic.

---

## Finding the Compaction Boundary

The core algorithm walks backward through the projection to locate the optimal cut point.

### Backward Accumulation with Budget Target

```ts
export function findCompactionBoundary(
    projection, inputBudget, contextTokens,
    compactedTo = 0, protectedFromSequence?): number {
  const tailBudget = inputBudget * 0.30; // 30% retention target
  
  // Walk from newest message backward
  let keptTokens = 0;
  for (let i = projection.length - 1; i >= 0; i--) {
    const msg = projection[i];
    keptTokens += msg.weight;
    
    if (keptTokens >= tailBudget && msg.canCut) {
      return msg.sequence; // Valid boundary found
    }
  }
  return compactedTo; // Fallback: no compaction
}

```

Key constraints enforced during this walk:

- **Atomic records**: Boundary only placed where `canCut === true`
- **Protected sequences**: Never cut past `protectedFromSequence` (e.g., pending connection requests from `findProtectedFromSequence`, lines 200-216)

### Protecting Revert Dependencies

Before finalizing, `protectRetainedReverts` (lines 54-66) scans retained messages for **revert operations**. If found, the boundary expands to include the source range of any reverted change, ensuring that replay can resolve change IDs correctly.

---

## Summarizing the Compacted Prefix

The discarded prefix becomes input for a summarization model call that produces a structured hand-off.

### Prompt Construction

`buildSummaryPrompt` (lines 72-108) performs several normalization steps:

- Merges consecutive messages from the same role
- Flattens tool call results into text descriptions
- Strips binary attachments and non-semantic metadata

### System Prompt for Summarization

```ts
const COMPACTION_SYSTEM_PROMPT = `You are a context summarizer for an AI coding assistant.
Generate a structured hand-off that captures:
- Current task state and objectives
- Key decisions made and their rationale
- Active files, bindings, and proposed changes
- Pending operations requiring attention

Be concise but complete. The summary will replace detailed conversation history.`;

```

This constant (lines 43-55) ensures consistent summary structure across compaction events.

---

## Creating Compaction Checkpoints

The final phase folds compacted state into a **`CompactionCheckpoint`** for future replay.

### State Collection

`buildCompactionState` (lines 72-124) iterates over compacted messages to aggregate:

- **Bindings**: Variable and resource references
- **Pins**: Explicitly locked context items
- **Proposed changes**: Pending edits not yet applied
- **Epoch information**: Version counters for optimistic concurrency

### Checkpoint Storage

The checkpoint stores minimal replay state, while the **summary text** is appended to the chat log as a regular user message. This separation allows:

- The UI to display "compacted" indicators with full history accessible via paging
- The backend to reconstruct working state from checkpoint plus retained tail
- Diagnostic tools to correlate summaries with their source ranges

---

## Complete Integration Example

Below is the canonical compaction workflow as implemented in the Workshop backend:

```ts
import {
  shouldCompactChat,
  isCompactionTurn,
  findCompactionBoundary,
  protectRetainedReverts,
  buildSummaryPrompt,
  buildCompactionState,
} from "@gadgets/workshop-backend/src/agent-compaction";

async function processTurn(
  chatMessages: AiChatMessage[],
  currentTokens: number,
  modelLimits: ModelLimits,
  checkpoint: CompactionCheckpoint | null
): Promise<CompactionResult> {
  
  // 1. Decide if compaction is required
  const needsCompaction = shouldCompactChat(currentTokens, modelLimits.inputBudget) 
                       || isCompactionTurn(chatMessages);
  
  if (!needsCompaction) {
    return { continueWith: chatMessages, checkpoint };
  }
  
  // 2. Build projection and locate boundary
  const projection = projectChatMessages(chatMessages);
  const rawBoundary = findCompactionBoundary(
    projection,
    modelLimits.inputBudget,
    currentTokens,
    checkpoint?.compactedTo ?? 0,
    findProtectedFromSequence(chatMessages)
  );
  
  // 3. Expand boundary to protect revert dependencies
  const safeBoundary = protectRetainedReverts(
    rawBoundary, 
    chatMessages, 
    checkpoint?.compactedTo ?? 0
  );
  
  // 4. Generate summary of discarded prefix
  const summaryPrompt = buildSummaryPrompt(
    projection,
    safeBoundary ?? 0,
    modelDescription
  );
  const summaryText = await callSummarizer(summaryPrompt);
  const summaryMessage = createUserMessage(summaryText);
  
  // 5. Construct new checkpoint from compacted state
  const newCheckpoint = buildCompactionState(
    chatMessages,
    safeBoundary ?? 0,
    initialBindings,
    checkpoint
  );
  
  // 6. Return modified message stream with summary injected
  const compactedMessages = [
    summaryMessage,
    ...chatMessages.slice(safeBoundary ?? 0)
  ];
  
  return {
    continueWith: compactedMessages,
    checkpoint: newCheckpoint
  };
}

```

The helper `projectChatMessages` (referenced but not in core file) maps `AiChatMessage` to `CompactionProjectionMessage` using `startsAgentTurn` to set `canCut` boundaries at agent turn boundaries.

---

## Key Source Files

| File | Purpose |
|------|---------|
| [`packages/workshop-backend/src/agent-compaction.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/workshop-backend/src/agent-compaction.ts) | Core compaction logic: boundary detection, checkpoint construction, summarization utilities |
| [`packages/workshop-backend/__tests__/agent-compaction.test.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/workshop-backend/__tests__/agent-compaction.test.ts) | Unit tests for boundary calculation and state folding correctness |
| [`packages/workshop-frontend/src/ChatInterface.compaction.test.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/workshop-frontend/src/ChatInterface.compaction.test.ts) | Frontend validation of `/compact` command handling and summary rendering |
| [`packages/workshop-shared/api.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/workshop-shared/api.ts) | `AiChatMessage` type definitions and RPC contracts |
| [`packages/workshop-backend/src/ai-invoke.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/packages/workshop-backend/src/ai-invoke.ts) | Model invocation helpers including `zeroUsage()` for synthetic messages |

---

## Summary

- **Agent compaction triggers at 85% of input budget** via `shouldCompactChat`, with manual override via `/compact` command detected by `isCompactionTurn`
- **Projection weighting** converts messages to `CompactionProjectionMessage` with `canCut` flags marking valid boundary points
- **Backward accumulation** in `findCompactionBoundary` targets 30% retention budget while respecting protected sequences
- **Revert protection** expands boundaries to include source ranges, preserving replay capability
- **Structured summarization** via `buildSummaryPrompt` and `COMPACTION_SYSTEM_PROMPT` generates concise hand-offs
- **Checkpoint persistence** in `buildCompactionState` maintains minimal replay state separate from chat log summaries

---

## Frequently Asked Questions

### What is the 85% threshold in Cloudflare OS agent compaction?

The `COMPACTION_TRIGGER_RATIO = 0.85` constant in [`agent-compaction.ts`](https://github.com/cloudflare/cloudflare-os/blob/main/agent-compaction.ts) defines when automatic compaction activates. When accumulated context tokens reach 85% of the model's input budget, the next turn is processed through the compaction pipeline. This threshold provides headroom for the final turn's tokens while avoiding emergency truncation.

### How does agent compaction preserve tool call integrity?

The `canCut` flag in `CompactionProjectionMessage` ensures boundaries only occur at complete logical records. Agent turns (including tool calls and their results) are marked atomic via `startsAgentTurn`. The backward walk in `findCompactionBoundary` explicitly checks `canCut` before accepting a boundary, preventing splits mid-transaction.

### Can users manually trigger compaction in Cloudflare OS?

Yes. The `/compact` slash command forces immediate compaction regardless of token usage. `isCompactionTurn` detects this by checking the last message for `type === "slashCommand"` with `commandId === "compact"`. This is processed identically to automatic triggers, producing the same checkpoint and summary flow.

### What happens to reverted changes during compaction?

`protectRetainedReverts` scans the retained message tail for revert operations. If any revert's source range falls outside the proposed boundary, the boundary expands to include that range. This ensures that replay can resolve change IDs referenced by retained revert messages, maintaining causal consistency across compaction events.