How the Automatic Compaction Algorithm Summarizes Older Turns in Prime Agent
The automatic compaction algorithm keeps the most recent ~20,000 tokens, identifies safe cut points that preserve turn boundaries, summarizes discarded older turns using an LLM, and injects the result as a compactionSummary message.
The automatic compaction feature in PrimeIntellect-ai/prime-agent prevents context window overflow by intelligently compressing conversation history. This article explains the complete workflow—from token estimation to summary injection—based on the actual source code implementation.
How Compaction Gets Triggered
The algorithm activates when a session nears its token limit. In packages/coding-agent/src/core/compaction/compaction.ts, the shouldCompact function performs this check:
// compaction.ts#L14-L18
shouldCompact(contextTokens, contextWindow, settings)
This comparison evaluates current token usage against the model's context window minus a reserve buffer of 16,384 tokens. When exceeded, compaction proceeds automatically after assistant turns or can be triggered manually via session.compact().
// Manual compaction trigger
await session.compact(undefined, { skipAbort: true });
// Automatic check in the run loop
if (shouldCompact(contextTokens, model.contextWindow, settings.compaction)) {
await session.compact();
}
Token Estimation Strategy
Before cutting, the system must estimate how many tokens each message consumes. The estimateTokens function (compaction.ts#L20-L78) applies a fast heuristic:
// Characters ÷ 4 estimation for all message types
estimateTokens(message)
This covers user, assistant, tool calls, images, and the special compactionSummary role itself—ensuring the compaction process accounts for its own output size.
Finding Safe Cut Points
The algorithm never arbitrarily slices conversation history. Two functions collaborate to find valid boundaries:
Identifying Valid Locations
findValidCutPoints (compaction.ts#L84-L121) walks session entries and marks safe indices for:
- User messages
- Assistant messages
- Custom messages
- Bash execution results
- Branch summaries
- Existing compaction summaries
Critical constraint: The function never cuts inside a tool-result chain, preserving operational integrity.
Selecting the Cut Point
findCutPoint (compaction.ts#L75-L104) then works backwards from the newest entry:
- Accumulates estimated token sizes
- Stops when keepRecentTokens (default 20,000) budget is reached
- Returns the first entry to retain; everything before becomes the older portion for summarization
Preserving Turn Boundaries
If the calculated cut lands mid-turn, findTurnStartIndex (compaction.ts#L33-L46) walks left to locate the nearest:
- User-type entry, or
- Branch/custom message
This ensures the automatic compaction algorithm summarizes older turns as complete units rather than fragmented exchanges, maintaining conversational coherence in the summary.
Generating the Summary
The older messages undergo transformation through several stages:
- Extract file operations —
extractFileOperationstracks which files were read or modified during the discarded portion - Build the prompt —
SUMMARIZATION_SYSTEM_PROMPTfromutils.tsstructures the summarization request - LLM invocation —
completeSimplecalls the provider (e.g., OpenAI) to generate a concise narrative
The summarization prompt compiles the serialized conversation into a compact format optimized for compression.
Injecting the Compaction Summary
The LLM output becomes a formal message via createCompactionSummaryMessage in packages/coding-agent/src/core/messages.ts#L194-L209:
// Creates the special compactionSummary role message
createCompactionSummaryMessage(summaryText, usageStats)
This compactionSummary entry is inserted at the start of the retained transcript, with usage statistics from the summarization call stored in a new CompactionEntry. The session manager in session-manager.ts persists this restructured history to disk.
Complete Workflow Summary
| Stage | Function/File | Purpose |
|---|---|---|
| Trigger detection | shouldCompact |
Identify context window pressure |
| Token estimation | estimateTokens |
Fast heuristic for all message types |
| Cut point discovery | findValidCutPoints, findCutPoint |
Locate safe, budget-compliant boundary |
| Turn alignment | findTurnStartIndex |
Prevent mid-turn fragmentation |
| Content extraction | extractFileOperations, SUMMARIZATION_SYSTEM_PROMPT |
Prepare summarization context |
| Summary generation | completeSimple (via AI provider) |
LLM-based compression |
| Message creation | createCompactionSummaryMessage |
Formalize as compactionSummary role |
| Persistence | session-manager.ts |
Save new transcript structure |
Summary
- Automatic compaction activates when tokens approach the context window minus a 16,384-token reserve
- The character-based heuristic (
chars ÷ 4) provides fast token estimation across all message types - Safe cut points exclude tool-result chains and respect turn boundaries via backward-walking algorithms
- Approximately 20,000 recent tokens are preserved while older turns get summarized
- The
compactionSummaryrole formally represents compressed history in the transcript - Usage tracking includes token consumption from the summarization call itself
Frequently Asked Questions
What triggers automatic compaction in Prime Agent?
Automatic compaction triggers when shouldCompact detects that current token usage exceeds the model's context window minus a 16,384-token safety buffer. This check runs automatically after each assistant turn, or developers can invoke session.compact() manually.
How does the algorithm decide which turns to keep versus summarize?
The algorithm walks backwards from the newest message, accumulating token estimates until reaching the keepRecentTokens budget (default 20,000). Everything from that point forward stays; everything before gets summarized. If this boundary falls mid-turn, findTurnStartIndex adjusts to the turn's beginning.
Why is there a special compactionSummary message type?
The compactionSummary role in messages.ts distinguishes compressed history from actual conversation turns. This preserves the narrative structure while allowing the system to track how many compaction cycles have occurred and account for their token costs in usage statistics.
Can compaction cut inside tool execution results?
No. The findValidCutPoints function explicitly excludes locations within tool-result chains. This safety mechanism ensures that partial tool outputs— which could corrupt session state—never appear in the retained or summarized portions of the transcript.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →