How the antinomyhq/forgecode Context Compaction Algorithm Works: A Complete Guide

Forgecode automatically compresses conversation history using a three-layer pipeline—trigger detection, strategy-based range selection, and summarization—to prevent token limit exhaustion while preserving critical context.

The antinomyhq/forgecode context compaction algorithm prevents LLM context window overflow by intelligently summarizing and removing older messages. This open-source Rust implementation processes conversation history through distinct trigger, strategy, and compaction phases to maintain optimal performance. Understanding these layers and their tunable parameters allows developers to balance context retention with token economy.

The Three-Layer Compaction Architecture

The algorithm operates through three coordinated layers defined in the forge_domain and forge_app crates:

Layer 1: Trigger Detection – When to Compact

The Compact struct in compact_config.rs defines the trigger conditions through seven configurable fields:

pub struct Compact {
    pub retention_window: usize,
    pub eviction_window: f64,
    pub token_threshold: Option<usize>,
    pub turn_threshold: Option<usize>,
    pub message_threshold: Option<usize>,
    pub on_turn_end: Option<bool>,
    pub model: Option<ModelId>,
}

Compaction fires when should_compact returns true via any of these four checks:

  1. should_compact_due_to_tokens – Activates when the token count exceeds token_threshold
  2. should_compact_due_to_turns – Activates when user turn count exceeds turn_threshold
  3. should_compact_due_to_messages – Activates when total message count exceeds message_threshold
  4. should_compact_on_turn_end – Activates when on_turn_end is true and the last message is from a user

The retention_window parameter guarantees that the most recent N messages remain untouched, while eviction_window (a float between 0.0 and 1.0) specifies the maximum percentage of context that may be removed during normal operation.

Layer 2: Strategy Computation – Selecting the Range

The CompactionStrategy enum in strategy.rs determines which messages to remove through four variants:

  • Evict(percentage) – Converts a percentage budget into a concrete "preserve-last-N" count via the to_fixed method
  • Retain(N) – Directly keeps the last N messages
  • Min(a, b) – Selects the more conservative bound between two strategies
  • Max(a, b) – Selects the more aggressive bound (used in "max" compaction mode)

The core algorithm find_sequence_preserving_last_n executes these steps:

  1. Locates the first assistant message as the start boundary
  2. Computes end = len - max_retention - 1
  3. Adjusts end backward to prevent splitting tool-call and tool-result pairs (verified via has_tool_call and has_tool_result)
  4. Returns the valid range (start, end) or indicates no compaction needed

This ensures the algorithm always compacts a contiguous block starting at the first assistant message while maintaining tool-call atomicity.

Layer 3: Context Summarization – Generating the Summary

The Compactor::compact method in crates/forge_app/src/compact.rs executes the final transformation:

  1. Builds the strategy using eviction.min(retention) for conservative mode or eviction.max(retention) for maximum mode
  2. Filters droppable messages (such as file attachments) via the transformer pipeline
  3. Creates a ContextSummary using the SummaryTransformer in transformers/compaction.rs, which deduplicates consecutive user messages, trims file-path prefixes, and removes system messages
  4. Renders the summary using the Mustache template forge-partial-summary-frame.md
  5. Aggregates usage statistics from all Usage objects within the compacted range
  6. Preserves reasoning chains by injecting the last non-empty reasoning_details from the removed range into the first remaining assistant message
  7. Splices the context by replacing the selected range with the single summary entry

The CompactionHandler in hooks/compaction.rs automatically invokes this service after each LLM response, ensuring continuous context optimization.

How to Tune Compaction Parameters

Parameter Typical Value Effect on Behavior
retention_window 3-5 Larger values preserve more recent messages, reducing compaction frequency
eviction_window 0.2-0.5 Higher percentages allow more aggressive context removal (clamped to 1.0)
token_threshold 3500-7000 Forces compaction regardless of percentage budget when this limit is hit
turn_threshold 10-20 Triggers compaction after specific user turn counts for long dialogs
message_threshold 50-100 Simple count-based safety net for high-frequency exchanges
on_turn_end true/false Enables immediate post-user compaction for UI-driven cleanup
model gpt-3.5-turbo Specifies a cheaper model for summarization to reduce API costs

Practical Tuning Workflow

  1. Start with defaults – Set eviction_window = 0.2 and retention_window = 0 to allow up to 20% eviction while preserving only the most recent context
  2. Monitor logs – Check the info! output from CompactionHandler for lines reading "Created context compaction summary" to verify range sizes
  3. Adjust percentage – Increase eviction_window to 0.4 or 0.5 if the context consistently approaches token limits
  4. Set hard limits – Configure token_threshold to 80% of your model's context window (e.g., 3500 for a 4K model) to guarantee compaction before LLM rejection
  5. Preserve recent history – Increase retention_window to 4 or 6 when maintaining the last few user-assistant exchanges is critical for task continuity
  6. Enable turn-end compacting – Set on_turn_end = true for conversational interfaces requiring tight context control after each user input

Implementation Examples

Configuring the Compact Parameters

use forge_domain::Compact;

// Preserve last 3 messages, allow 30% eviction, force at 2000 tokens
let compact_cfg = Compact::new()
    .retention_window(3)
    .eviction_window(0.30)
    .token_threshold(2000)
    .on_turn_end(true);

Manual Compactor Invocation

use forge_domain::Context;
use forge_app::compact::Compactor;
use forge_app::Environment;

// Initialize with configuration and environment
let env = Environment::default();
let compactor = Compactor::new(compact_cfg, env);

// false = conservative strategy (min), true = aggressive (max)
let compacted_context = compactor.compact(context, false)?;

Hook Registration Pattern

use forge_app::hooks::compaction::CompactionHandler;
use forge_domain::{Agent, Environment, EventBus};

let agent = Agent::default();
let env = Environment::default();

let handler = CompactionHandler::new(agent.clone(), env.clone());
EventBus::subscribe(handler);

Reading Compaction Logs

Log output indicates compaction trigger and range:

INFO  forge_app::hooks::compaction: Compaction triggered by hook
    agent_id = "my-agent"
INFO  forge_app::compact: Created context compaction summary
    sequence_start = 4
    sequence_end   = 9
    sequence_length = 6

The sequence_start and sequence_end values map directly to the range calculated by CompactionStrategy::eviction_range.

Summary

  • The antinomyhq/forgecode algorithm uses three distinct layers: trigger detection (should_compact), range strategy (CompactionStrategy), and summarization (Compactor)
  • Trigger conditions include token thresholds, turn counts, message counts, and turn-end flags defined in compact_config.rs
  • Strategy selection preserves the last N messages and prevents splitting tool-call pairs through the find_sequence_preserving_last_n logic in strategy.rs
  • Compaction execution generates summaries via SummaryTransformer, aggregates usage statistics, and preserves reasoning chains through compaction.rs
  • Tuning parameters balance retention_window (preservation) against eviction_window (percentage budget) and hard thresholds for specific use cases

Frequently Asked Questions

How does Forgecode prevent removing tool calls without their results?

The find_sequence_preserving_last_n function in strategy.rs checks for has_tool_call and has_tool_result when calculating the end index. If a tool-call or tool-result message would be bisected by the compaction range, the algorithm adjusts the boundary backward to keep both messages together in the retained context or remove them as a pair.

What happens if both token_threshold and eviction_window are configured?

Hard thresholds always take precedence. If token_threshold is set to 3500 and the context reaches 3501 tokens, the compactor triggers immediately regardless of whether the eviction_window percentage would normally allow more content to remain. This ensures the context never exceeds the LLM's hard limit even if the percentage budget suggests otherwise.

Can I use a different model for summarization than the main agent?

Yes. The model field in Compact accepts an optional ModelId. When specified, the Compactor uses this model for the SummaryTransformer pipeline instead of the agent's default model. This allows cost optimization by using lighter models (e.g., GPT-3.5-Turbo) for compression while reserving larger models for reasoning tasks.

How does the algorithm handle reasoning or chain-of-thought content?

During compaction in crates/forge_app/src/compact.rs, the Compactor extracts the last non-empty reasoning_details from the range being removed. This reasoning content is then injected into the first remaining assistant message after the compaction boundary, ensuring that extended thinking chains and intermediate steps remain available to the LLM for subsequent turns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →