How Prime Agent's Automatic Compaction Algorithm Preserves Recent Context
TLDR: Prime Agent's automatic compaction algorithm preserves recent context by retaining a configurable slice of the most recent tokens—defaulting to 20,000—while summarizing only older portions of the conversation.
The Prime Agent coding assistant manages long-running sessions through an intelligent compaction system that balances memory efficiency with context retention. Understanding how this automatic compaction algorithm handles recent context is essential for developers tuning agent behavior in extended coding workflows.
The keepRecentTokens Mechanism
At the core of context preservation lies the keepRecentTokens setting. This parameter controls exactly how much of the conversation tail remains untouched during compaction operations.
The default value is 20,000 tokens, exposed through the settings manager in packages/coding-agent/src/core/settings-manager.ts:
// From settings-manager.ts (lines 890-901)
getCompactionKeepRecentTokens(): number {
return this.getSetting('compaction.keepRecentTokens', 20000);
}
This configuration-driven approach allows fine-grained control without modifying core algorithm code.
The findCutPoint Algorithm
The actual preservation logic resides in packages/coding-agent/src/core/compaction/compaction.ts, specifically within the findCutPoint function (lines 360-394). This function determines where to partition the session between preserved recent context and summarized history.
The algorithm executes three sequential steps:
- Reverse traversal — Scan session entries from newest to oldest
- Token accumulation — Sum token counts until reaching
keepRecentTokens - Boundary establishment — Cut at the point where accumulated tokens meet or exceed the threshold
Entries after the cut point remain verbatim; earlier entries get compressed into summaries.
// Conceptual implementation pattern from compaction.ts
function findCutPoint(
entries: PathEntry[],
boundaryStart: number,
boundaryEnd: number,
keepRecentTokens: number // The preservation threshold
): number {
let accumulated = 0;
// Walk backwards from newest entry
for (let i = entries.length - 1; i >= boundaryStart; i--) {
accumulated += estimateTokens(entries[i]);
if (accumulated >= keepRecentTokens) {
return i; // Cut here: keep entries [i, end]
}
}
return boundaryStart; // Keep everything
}
Configuring Context Preservation
Adjust keepRecentTokens based on your workflow requirements:
// Retain more context for complex multi-file refactoring
session.settingsManager.applyOverrides({
compaction: { keepRecentTokens: 50000 }
});
// Minimize retention for simple, stateless tasks
session.settingsManager.applyOverrides({
compaction: { keepRecentTokens: 5000 }
});
After configuration, invoke compaction explicitly:
await session.compact(); // Respects the overridden threshold
Verification Through Testing
The preservation behavior is validated in packages/coding-agent/test/compaction.test.ts (lines 260-300), which exercises various keepRecentTokens scenarios to confirm:
- Exact token count boundaries are respected
- Recent entries maintain full fidelity
- Summarization only affects pre-cut content
Summary
- Default preservation: 20,000 recent tokens via
getCompactionKeepRecentTokens() - Core implementation:
findCutPoint()incompaction.tshandles the partitioning logic - Configurable threshold: Override through
compaction.keepRecentTokenssetting - Guaranteed fidelity: Post-cut entries remain unmodified; only older context is summarized
Frequently Asked Questions
How does keepRecentTokens interact with the overall token limit?
The keepRecentTokens setting operates independently of model context windows. It determines how much recent conversation survives compaction, not the maximum session size. Prime Agent may trigger compaction multiple times as sessions grow, each time preserving the configured recent token count.
What happens if a single message exceeds keepRecentTokens?
When a single entry exceeds the threshold, findCutPoint includes it in the preserved portion. The algorithm accumulates token counts until meeting or exceeding the limit, ensuring complete entries are never split across the preservation boundary.
Can I disable automatic compaction entirely?
While keepRecentTokens can be set to arbitrarily high values, the compaction system is designed for operational sustainability. Disabling compaction risks performance degradation in long sessions rather than explicit failures. Consult the settings manager implementation for available override strategies.
How are tokens estimated before compaction runs?
The compaction system uses approximate token counting based on character heuristics or lightweight tokenization, matching the estimation logic used elsewhere in the session management pipeline. This ensures consistency between when compaction triggers and how keepRecentTokens is calculated.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →