# How the antinomyhq/forgecode Context Compaction Algorithm Works: A Complete Guide

> Explore the antinomyhq/forgecode context compaction algorithm. Learn how its three-layer pipeline optimizes token usage without losing vital information. Discover tuning strategies.

- Repository: [Forge Code/forgecode](https://github.com/antinomyhq/forgecode)
- Tags: deep-dive
- Published: 2026-04-08

---

**Forgecode automatically compresses conversation history using a three-layer pipeline—trigger detection, strategy-based range selection, and summarization—to prevent token limit exhaustion while preserving critical context.**

The antinomyhq/forgecode context compaction algorithm prevents LLM context window overflow by intelligently summarizing and removing older messages. This open-source Rust implementation processes conversation history through distinct trigger, strategy, and compaction phases to maintain optimal performance. Understanding these layers and their tunable parameters allows developers to balance context retention with token economy.

## The Three-Layer Compaction Architecture

The algorithm operates through three coordinated layers defined in the `forge_domain` and `forge_app` crates:

- **Trigger Layer**: Decides when compaction occurs based on token counts, message limits, or conversation turns (`Compact::should_compact` in [`crates/forge_domain/src/compact/compact_config.rs`](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_domain/src/compact/compact_config.rs))
- **Strategy Layer**: Calculates the exact message range for removal while preserving recent context and tool-call atomicity (`CompactionStrategy` in [`crates/forge_domain/src/compact/strategy.rs`](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_domain/src/compact/strategy.rs))
- **Compaction Layer**: Generates summaries, aggregates usage statistics, and maintains reasoning chains (`Compactor` in [`crates/forge_app/src/compact.rs`](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_app/src/compact.rs))

### Layer 1: Trigger Detection – When to Compact

The `Compact` struct in [`compact_config.rs`](https://github.com/antinomyhq/forgecode/blob/main/compact_config.rs) defines the trigger conditions through seven configurable fields:

```rust
pub struct Compact {
    pub retention_window: usize,
    pub eviction_window: f64,
    pub token_threshold: Option<usize>,
    pub turn_threshold: Option<usize>,
    pub message_threshold: Option<usize>,
    pub on_turn_end: Option<bool>,
    pub model: Option<ModelId>,
}

```

Compaction fires when `should_compact` returns true via any of these four checks:

1. **`should_compact_due_to_tokens`** – Activates when the token count exceeds `token_threshold`
2. **`should_compact_due_to_turns`** – Activates when user turn count exceeds `turn_threshold`
3. **`should_compact_due_to_messages`** – Activates when total message count exceeds `message_threshold`
4. **`should_compact_on_turn_end`** – Activates when `on_turn_end` is true and the last message is from a user

The `retention_window` parameter guarantees that the most recent `N` messages remain untouched, while `eviction_window` (a float between 0.0 and 1.0) specifies the maximum percentage of context that may be removed during normal operation.

### Layer 2: Strategy Computation – Selecting the Range

The `CompactionStrategy` enum in [`strategy.rs`](https://github.com/antinomyhq/forgecode/blob/main/strategy.rs) determines which messages to remove through four variants:

- **`Evict(percentage)`** – Converts a percentage budget into a concrete "preserve-last-N" count via the `to_fixed` method
- **`Retain(N)`** – Directly keeps the last `N` messages
- **`Min(a, b)`** – Selects the more conservative bound between two strategies
- **`Max(a, b)`** – Selects the more aggressive bound (used in "max" compaction mode)

The core algorithm `find_sequence_preserving_last_n` executes these steps:

1. Locates the first assistant message as the `start` boundary
2. Computes `end = len - max_retention - 1`
3. Adjusts `end` backward to prevent splitting tool-call and tool-result pairs (verified via `has_tool_call` and `has_tool_result`)
4. Returns the valid range `(start, end)` or indicates no compaction needed

This ensures the algorithm always compacts a contiguous block starting at the first assistant message while maintaining tool-call atomicity.

### Layer 3: Context Summarization – Generating the Summary

The `Compactor::compact` method in [`crates/forge_app/src/compact.rs`](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_app/src/compact.rs) executes the final transformation:

1. **Builds the strategy** using `eviction.min(retention)` for conservative mode or `eviction.max(retention)` for maximum mode
2. **Filters droppable messages** (such as file attachments) via the transformer pipeline
3. **Creates a `ContextSummary`** using the `SummaryTransformer` in [`transformers/compaction.rs`](https://github.com/antinomyhq/forgecode/blob/main/transformers/compaction.rs), which deduplicates consecutive user messages, trims file-path prefixes, and removes system messages
4. **Renders the summary** using the Mustache template [`forge-partial-summary-frame.md`](https://github.com/antinomyhq/forgecode/blob/main/forge-partial-summary-frame.md)
5. **Aggregates usage statistics** from all `Usage` objects within the compacted range
6. **Preserves reasoning chains** by injecting the last non-empty `reasoning_details` from the removed range into the first remaining assistant message
7. **Splices the context** by replacing the selected range with the single summary entry

The `CompactionHandler` in [`hooks/compaction.rs`](https://github.com/antinomyhq/forgecode/blob/main/hooks/compaction.rs) automatically invokes this service after each LLM response, ensuring continuous context optimization.

## How to Tune Compaction Parameters

| Parameter | Typical Value | Effect on Behavior |
|-----------|---------------|-------------------|
| `retention_window` | 3-5 | Larger values preserve more recent messages, reducing compaction frequency |
| `eviction_window` | 0.2-0.5 | Higher percentages allow more aggressive context removal (clamped to 1.0) |
| `token_threshold` | 3500-7000 | Forces compaction regardless of percentage budget when this limit is hit |
| `turn_threshold` | 10-20 | Triggers compaction after specific user turn counts for long dialogs |
| `message_threshold` | 50-100 | Simple count-based safety net for high-frequency exchanges |
| `on_turn_end` | true/false | Enables immediate post-user compaction for UI-driven cleanup |
| `model` | gpt-3.5-turbo | Specifies a cheaper model for summarization to reduce API costs |

### Practical Tuning Workflow

1. **Start with defaults** – Set `eviction_window = 0.2` and `retention_window = 0` to allow up to 20% eviction while preserving only the most recent context
2. **Monitor logs** – Check the `info!` output from `CompactionHandler` for lines reading "Created context compaction summary" to verify range sizes
3. **Adjust percentage** – Increase `eviction_window` to 0.4 or 0.5 if the context consistently approaches token limits
4. **Set hard limits** – Configure `token_threshold` to 80% of your model's context window (e.g., 3500 for a 4K model) to guarantee compaction before LLM rejection
5. **Preserve recent history** – Increase `retention_window` to 4 or 6 when maintaining the last few user-assistant exchanges is critical for task continuity
6. **Enable turn-end compacting** – Set `on_turn_end = true` for conversational interfaces requiring tight context control after each user input

## Implementation Examples

### Configuring the Compact Parameters

```rust
use forge_domain::Compact;

// Preserve last 3 messages, allow 30% eviction, force at 2000 tokens
let compact_cfg = Compact::new()
    .retention_window(3)
    .eviction_window(0.30)
    .token_threshold(2000)
    .on_turn_end(true);

```

### Manual Compactor Invocation

```rust
use forge_domain::Context;
use forge_app::compact::Compactor;
use forge_app::Environment;

// Initialize with configuration and environment
let env = Environment::default();
let compactor = Compactor::new(compact_cfg, env);

// false = conservative strategy (min), true = aggressive (max)
let compacted_context = compactor.compact(context, false)?;

```

### Hook Registration Pattern

```rust
use forge_app::hooks::compaction::CompactionHandler;
use forge_domain::{Agent, Environment, EventBus};

let agent = Agent::default();
let env = Environment::default();

let handler = CompactionHandler::new(agent.clone(), env.clone());
EventBus::subscribe(handler);

```

### Reading Compaction Logs

Log output indicates compaction trigger and range:

```text
INFO  forge_app::hooks::compaction: Compaction triggered by hook
    agent_id = "my-agent"
INFO  forge_app::compact: Created context compaction summary
    sequence_start = 4
    sequence_end   = 9
    sequence_length = 6

```

The `sequence_start` and `sequence_end` values map directly to the range calculated by `CompactionStrategy::eviction_range`.

## Summary

- The antinomyhq/forgecode algorithm uses three distinct layers: **trigger detection** (`should_compact`), **range strategy** (`CompactionStrategy`), and **summarization** (`Compactor`)
- **Trigger conditions** include token thresholds, turn counts, message counts, and turn-end flags defined in [`compact_config.rs`](https://github.com/antinomyhq/forgecode/blob/main/compact_config.rs)
- **Strategy selection** preserves the last `N` messages and prevents splitting tool-call pairs through the `find_sequence_preserving_last_n` logic in [`strategy.rs`](https://github.com/antinomyhq/forgecode/blob/main/strategy.rs)
- **Compaction execution** generates summaries via `SummaryTransformer`, aggregates usage statistics, and preserves reasoning chains through [`compaction.rs`](https://github.com/antinomyhq/forgecode/blob/main/compaction.rs)
- **Tuning parameters** balance `retention_window` (preservation) against `eviction_window` (percentage budget) and hard thresholds for specific use cases

## Frequently Asked Questions

### How does Forgecode prevent removing tool calls without their results?

The `find_sequence_preserving_last_n` function in [`strategy.rs`](https://github.com/antinomyhq/forgecode/blob/main/strategy.rs) checks for `has_tool_call` and `has_tool_result` when calculating the `end` index. If a tool-call or tool-result message would be bisected by the compaction range, the algorithm adjusts the boundary backward to keep both messages together in the retained context or remove them as a pair.

### What happens if both `token_threshold` and `eviction_window` are configured?

Hard thresholds always take precedence. If `token_threshold` is set to 3500 and the context reaches 3501 tokens, the compactor triggers immediately regardless of whether the `eviction_window` percentage would normally allow more content to remain. This ensures the context never exceeds the LLM's hard limit even if the percentage budget suggests otherwise.

### Can I use a different model for summarization than the main agent?

Yes. The `model` field in `Compact` accepts an optional `ModelId`. When specified, the `Compactor` uses this model for the `SummaryTransformer` pipeline instead of the agent's default model. This allows cost optimization by using lighter models (e.g., GPT-3.5-Turbo) for compression while reserving larger models for reasoning tasks.

### How does the algorithm handle reasoning or chain-of-thought content?

During compaction in [`crates/forge_app/src/compact.rs`](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_app/src/compact.rs), the `Compactor` extracts the **last non-empty** `reasoning_details` from the range being removed. This reasoning content is then injected into the first remaining assistant message after the compaction boundary, ensuring that extended thinking chains and intermediate steps remain available to the LLM for subsequent turns.