How the antinomyhq/forgecode Context Compaction Algorithm Works: A Complete Guide
Forgecode automatically compresses conversation history using a three-layer pipeline—trigger detection, strategy-based range selection, and summarization—to prevent token limit exhaustion while preserving critical context.
The antinomyhq/forgecode context compaction algorithm prevents LLM context window overflow by intelligently summarizing and removing older messages. This open-source Rust implementation processes conversation history through distinct trigger, strategy, and compaction phases to maintain optimal performance. Understanding these layers and their tunable parameters allows developers to balance context retention with token economy.
The Three-Layer Compaction Architecture
The algorithm operates through three coordinated layers defined in the forge_domain and forge_app crates:
- Trigger Layer: Decides when compaction occurs based on token counts, message limits, or conversation turns (
Compact::should_compactincrates/forge_domain/src/compact/compact_config.rs) - Strategy Layer: Calculates the exact message range for removal while preserving recent context and tool-call atomicity (
CompactionStrategyincrates/forge_domain/src/compact/strategy.rs) - Compaction Layer: Generates summaries, aggregates usage statistics, and maintains reasoning chains (
Compactorincrates/forge_app/src/compact.rs)
Layer 1: Trigger Detection – When to Compact
The Compact struct in compact_config.rs defines the trigger conditions through seven configurable fields:
pub struct Compact {
pub retention_window: usize,
pub eviction_window: f64,
pub token_threshold: Option<usize>,
pub turn_threshold: Option<usize>,
pub message_threshold: Option<usize>,
pub on_turn_end: Option<bool>,
pub model: Option<ModelId>,
}
Compaction fires when should_compact returns true via any of these four checks:
should_compact_due_to_tokens– Activates when the token count exceedstoken_thresholdshould_compact_due_to_turns– Activates when user turn count exceedsturn_thresholdshould_compact_due_to_messages– Activates when total message count exceedsmessage_thresholdshould_compact_on_turn_end– Activates whenon_turn_endis true and the last message is from a user
The retention_window parameter guarantees that the most recent N messages remain untouched, while eviction_window (a float between 0.0 and 1.0) specifies the maximum percentage of context that may be removed during normal operation.
Layer 2: Strategy Computation – Selecting the Range
The CompactionStrategy enum in strategy.rs determines which messages to remove through four variants:
Evict(percentage)– Converts a percentage budget into a concrete "preserve-last-N" count via theto_fixedmethodRetain(N)– Directly keeps the lastNmessagesMin(a, b)– Selects the more conservative bound between two strategiesMax(a, b)– Selects the more aggressive bound (used in "max" compaction mode)
The core algorithm find_sequence_preserving_last_n executes these steps:
- Locates the first assistant message as the
startboundary - Computes
end = len - max_retention - 1 - Adjusts
endbackward to prevent splitting tool-call and tool-result pairs (verified viahas_tool_callandhas_tool_result) - Returns the valid range
(start, end)or indicates no compaction needed
This ensures the algorithm always compacts a contiguous block starting at the first assistant message while maintaining tool-call atomicity.
Layer 3: Context Summarization – Generating the Summary
The Compactor::compact method in crates/forge_app/src/compact.rs executes the final transformation:
- Builds the strategy using
eviction.min(retention)for conservative mode oreviction.max(retention)for maximum mode - Filters droppable messages (such as file attachments) via the transformer pipeline
- Creates a
ContextSummaryusing theSummaryTransformerintransformers/compaction.rs, which deduplicates consecutive user messages, trims file-path prefixes, and removes system messages - Renders the summary using the Mustache template
forge-partial-summary-frame.md - Aggregates usage statistics from all
Usageobjects within the compacted range - Preserves reasoning chains by injecting the last non-empty
reasoning_detailsfrom the removed range into the first remaining assistant message - Splices the context by replacing the selected range with the single summary entry
The CompactionHandler in hooks/compaction.rs automatically invokes this service after each LLM response, ensuring continuous context optimization.
How to Tune Compaction Parameters
| Parameter | Typical Value | Effect on Behavior |
|---|---|---|
retention_window |
3-5 | Larger values preserve more recent messages, reducing compaction frequency |
eviction_window |
0.2-0.5 | Higher percentages allow more aggressive context removal (clamped to 1.0) |
token_threshold |
3500-7000 | Forces compaction regardless of percentage budget when this limit is hit |
turn_threshold |
10-20 | Triggers compaction after specific user turn counts for long dialogs |
message_threshold |
50-100 | Simple count-based safety net for high-frequency exchanges |
on_turn_end |
true/false | Enables immediate post-user compaction for UI-driven cleanup |
model |
gpt-3.5-turbo | Specifies a cheaper model for summarization to reduce API costs |
Practical Tuning Workflow
- Start with defaults – Set
eviction_window = 0.2andretention_window = 0to allow up to 20% eviction while preserving only the most recent context - Monitor logs – Check the
info!output fromCompactionHandlerfor lines reading "Created context compaction summary" to verify range sizes - Adjust percentage – Increase
eviction_windowto 0.4 or 0.5 if the context consistently approaches token limits - Set hard limits – Configure
token_thresholdto 80% of your model's context window (e.g., 3500 for a 4K model) to guarantee compaction before LLM rejection - Preserve recent history – Increase
retention_windowto 4 or 6 when maintaining the last few user-assistant exchanges is critical for task continuity - Enable turn-end compacting – Set
on_turn_end = truefor conversational interfaces requiring tight context control after each user input
Implementation Examples
Configuring the Compact Parameters
use forge_domain::Compact;
// Preserve last 3 messages, allow 30% eviction, force at 2000 tokens
let compact_cfg = Compact::new()
.retention_window(3)
.eviction_window(0.30)
.token_threshold(2000)
.on_turn_end(true);
Manual Compactor Invocation
use forge_domain::Context;
use forge_app::compact::Compactor;
use forge_app::Environment;
// Initialize with configuration and environment
let env = Environment::default();
let compactor = Compactor::new(compact_cfg, env);
// false = conservative strategy (min), true = aggressive (max)
let compacted_context = compactor.compact(context, false)?;
Hook Registration Pattern
use forge_app::hooks::compaction::CompactionHandler;
use forge_domain::{Agent, Environment, EventBus};
let agent = Agent::default();
let env = Environment::default();
let handler = CompactionHandler::new(agent.clone(), env.clone());
EventBus::subscribe(handler);
Reading Compaction Logs
Log output indicates compaction trigger and range:
INFO forge_app::hooks::compaction: Compaction triggered by hook
agent_id = "my-agent"
INFO forge_app::compact: Created context compaction summary
sequence_start = 4
sequence_end = 9
sequence_length = 6
The sequence_start and sequence_end values map directly to the range calculated by CompactionStrategy::eviction_range.
Summary
- The antinomyhq/forgecode algorithm uses three distinct layers: trigger detection (
should_compact), range strategy (CompactionStrategy), and summarization (Compactor) - Trigger conditions include token thresholds, turn counts, message counts, and turn-end flags defined in
compact_config.rs - Strategy selection preserves the last
Nmessages and prevents splitting tool-call pairs through thefind_sequence_preserving_last_nlogic instrategy.rs - Compaction execution generates summaries via
SummaryTransformer, aggregates usage statistics, and preserves reasoning chains throughcompaction.rs - Tuning parameters balance
retention_window(preservation) againsteviction_window(percentage budget) and hard thresholds for specific use cases
Frequently Asked Questions
How does Forgecode prevent removing tool calls without their results?
The find_sequence_preserving_last_n function in strategy.rs checks for has_tool_call and has_tool_result when calculating the end index. If a tool-call or tool-result message would be bisected by the compaction range, the algorithm adjusts the boundary backward to keep both messages together in the retained context or remove them as a pair.
What happens if both token_threshold and eviction_window are configured?
Hard thresholds always take precedence. If token_threshold is set to 3500 and the context reaches 3501 tokens, the compactor triggers immediately regardless of whether the eviction_window percentage would normally allow more content to remain. This ensures the context never exceeds the LLM's hard limit even if the percentage budget suggests otherwise.
Can I use a different model for summarization than the main agent?
Yes. The model field in Compact accepts an optional ModelId. When specified, the Compactor uses this model for the SummaryTransformer pipeline instead of the agent's default model. This allows cost optimization by using lighter models (e.g., GPT-3.5-Turbo) for compression while reserving larger models for reasoning tasks.
How does the algorithm handle reasoning or chain-of-thought content?
During compaction in crates/forge_app/src/compact.rs, the Compactor extracts the last non-empty reasoning_details from the range being removed. This reasoning content is then injected into the first remaining assistant message after the compaction boundary, ensuring that extended thinking chains and intermediate steps remain available to the LLM for subsequent turns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →