# How Prime Agent's Automatic Compaction Algorithm Preserves Recent Context

> Discover how Prime Agent's automatic compaction algorithm preserves recent context by saving recent tokens and summarizing older conversations for efficient AI memory management.

- Repository: [Prime Intellect/prime-agent](https://github.com/PrimeIntellect-ai/prime-agent)
- Tags: internals
- Published: 2026-09-06

---

**TLDR:** Prime Agent's automatic compaction algorithm preserves recent context by retaining a configurable slice of the most recent tokens—defaulting to 20,000—while summarizing only older portions of the conversation.

The Prime Agent coding assistant manages long-running sessions through an intelligent compaction system that balances memory efficiency with context retention. Understanding how this **automatic compaction algorithm** handles recent context is essential for developers tuning agent behavior in extended coding workflows.

## The keepRecentTokens Mechanism

At the core of context preservation lies the **`keepRecentTokens`** setting. This parameter controls exactly how much of the conversation tail remains untouched during compaction operations.

The default value is **20,000 tokens**, exposed through the settings manager in [`packages/coding-agent/src/core/settings-manager.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/core/settings-manager.ts):

```typescript
// From settings-manager.ts (lines 890-901)
getCompactionKeepRecentTokens(): number {
  return this.getSetting('compaction.keepRecentTokens', 20000);
}

```

This configuration-driven approach allows fine-grained control without modifying core algorithm code.

## The findCutPoint Algorithm

The actual preservation logic resides in [`packages/coding-agent/src/core/compaction/compaction.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/core/compaction/compaction.ts), specifically within the **`findCutPoint`** function (lines 360-394). This function determines where to partition the session between preserved recent context and summarized history.

The algorithm executes three sequential steps:

1. **Reverse traversal** — Scan session entries from newest to oldest
2. **Token accumulation** — Sum token counts until reaching `keepRecentTokens`
3. **Boundary establishment** — Cut at the point where accumulated tokens meet or exceed the threshold

Entries after the cut point remain verbatim; earlier entries get compressed into summaries.

```typescript
// Conceptual implementation pattern from compaction.ts
function findCutPoint(
  entries: PathEntry[],
  boundaryStart: number,
  boundaryEnd: number,
  keepRecentTokens: number  // The preservation threshold
): number {
  let accumulated = 0;
  // Walk backwards from newest entry
  for (let i = entries.length - 1; i >= boundaryStart; i--) {
    accumulated += estimateTokens(entries[i]);
    if (accumulated >= keepRecentTokens) {
      return i; // Cut here: keep entries [i, end]
    }
  }
  return boundaryStart; // Keep everything
}

```

## Configuring Context Preservation

Adjust `keepRecentTokens` based on your workflow requirements:

```typescript
// Retain more context for complex multi-file refactoring
session.settingsManager.applyOverrides({
  compaction: { keepRecentTokens: 50000 }
});

// Minimize retention for simple, stateless tasks
session.settingsManager.applyOverrides({
  compaction: { keepRecentTokens: 5000 }
});

```

After configuration, invoke compaction explicitly:

```typescript
await session.compact(); // Respects the overridden threshold

```

## Verification Through Testing

The preservation behavior is validated in [`packages/coding-agent/test/compaction.test.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/test/compaction.test.ts) (lines 260-300), which exercises various `keepRecentTokens` scenarios to confirm:

- Exact token count boundaries are respected
- Recent entries maintain full fidelity
- Summarization only affects pre-cut content

## Summary

- **Default preservation**: 20,000 recent tokens via `getCompactionKeepRecentTokens()`
- **Core implementation**: `findCutPoint()` in [`compaction.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/compaction.ts) handles the partitioning logic
- **Configurable threshold**: Override through `compaction.keepRecentTokens` setting
- **Guaranteed fidelity**: Post-cut entries remain unmodified; only older context is summarized

## Frequently Asked Questions

### How does `keepRecentTokens` interact with the overall token limit?

The `keepRecentTokens` setting operates independently of model context windows. It determines how much recent conversation survives compaction, not the maximum session size. Prime Agent may trigger compaction multiple times as sessions grow, each time preserving the configured recent token count.

### What happens if a single message exceeds `keepRecentTokens`?

When a single entry exceeds the threshold, `findCutPoint` includes it in the preserved portion. The algorithm accumulates token counts until meeting or exceeding the limit, ensuring complete entries are never split across the preservation boundary.

### Can I disable automatic compaction entirely?

While `keepRecentTokens` can be set to arbitrarily high values, the compaction system is designed for operational sustainability. Disabling compaction risks performance degradation in long sessions rather than explicit failures. Consult the settings manager implementation for available override strategies.

### How are tokens estimated before compaction runs?

The compaction system uses approximate token counting based on character heuristics or lightweight tokenization, matching the estimation logic used elsewhere in the session management pipeline. This ensures consistency between when compaction triggers and how `keepRecentTokens` is calculated.