Forge Compact Configuration: Managing Context with Token Thresholds and Retention Windows

Forge's compact configuration uses token_threshold to trigger context summarization, retention_window to preserve recent messages, and eviction_window to limit how much history can be replaced by summaries.

The compact feature in antinomyhq/forgecode automatically manages conversation context length by intelligently summarizing older messages when thresholds are exceeded. Understanding how token_threshold, retention_window, and eviction_window interact allows you to fine-tune memory usage and maintain relevant conversation history without hitting model token limits.

Understanding the Three Core Parameters

The Compact struct defined in crates/forge_config/src/compact.rs exposes three critical fields that control when and how conversation compaction occurs.

Token Threshold

The token_threshold field is an Option<usize> that acts as the primary trigger for compaction. When the current token count—including both prompt and completion tokens—exceeds this limit, the should_compact method returns true and initiates the summarization process.

[compact]
token_threshold = 1500

If left as None, the token-based trigger is disabled and compaction relies on other thresholds like turn_threshold or message_threshold.

Retention Window

The retention_window field (type usize) guarantees that the N most recent messages remain untouched regardless of other compaction settings. This ensures critical recent context is never summarized, preserving the immediate conversation flow.

[compact]
retention_window = 5

With this configuration, the last five messages in the conversation context are always kept verbatim, even if the token threshold has been exceeded.

Eviction Window

The eviction_window field accepts a Percentage value (0.0 to 1.0) that defines the maximum proportion of the context that may be replaced by a summary. A value of 0.2 means at most 20% of the conversation can be evicted and summarized, ensuring at least 80% of the original messages remain intact.

[compact]
eviction_window = 0.3

This percentage-based limit prevents overly aggressive summarization that might remove important contextual details from the conversation history.

How the Compaction Decision Flow Works

The compaction logic follows a specific decision pipeline implemented across three key files: compact_config.rs, compact.rs (domain), and compact.rs (app).

Step 1: Threshold Evaluation

In crates/forge_domain/src/compact/compact_config.rs, the should_compact method (lines 66-73) evaluates whether compaction is necessary by checking the configured thresholds:

// Logic from should_compact
if compact.token_threshold.map_or(false, |t| token_count >= t) {
    return true;
}

This method returns true if any threshold—token_threshold, turn_threshold, or message_threshold—is satisfied, or if the on_turn_end flag is set.

Step 2: Strategy Construction

When compaction triggers, the Compactor::compact method in crates/forge_app/src/compact.rs (lines 42-44) builds two competing strategies:

let eviction = CompactionStrategy::evict(self.compact.eviction_window);
let retention = CompactionStrategy::retain(self.compact.retention_window);

The eviction strategy defines how much can be removed based on the percentage limit, while the retention strategy defines what must be kept based on the message count.

Step 3: Strategy Intersection

The final compaction range depends on the max parameter passed to compact():

  • If max == true: The retention strategy dominates, keeping only the most recent retention_window messages and summarizing everything else.
  • If max == false: The system uses eviction.min(retention), taking the intersection of both constraints to determine which message range can be safely replaced by a summary.

The identified range is then summarized into a single user-message entry, with original messages dropped and usage statistics accumulated.

Configuration Examples

TOML Configuration

Create a forge.toml file to configure compaction behavior for your agent:

[compact]

# Trigger compaction when context exceeds 1,500 tokens

token_threshold = 1500

# Always preserve the last 5 messages

retention_window = 5

# Allow up to 30% of context to be evicted

eviction_window = 0.3

Programmatic Configuration

Build a Compact instance programmatically using the builder pattern generated by derive_setters:

use forge_config::Compact;

let compact = Compact::new()
    .token_threshold(1500_usize)
    .retention_window(5)
    .eviction_window(forge_config::Percentage::new(0.3).unwrap());

Runtime Usage

Trigger compaction from an agent by checking thresholds and invoking the compactor:

use forge_app::{Compactor, TemplateEngine};
use forge_domain::{Context, Environment};

let env = Environment::default();
let compactor = Compactor::new(compact, env);

if compact.should_compact(&ctx, token_count) {
    // Use intersection of eviction and retention (max = false)
    let compacted = compactor.compact(ctx, false)?;
    // compacted now contains a summary entry
}

Implementation Details

The compaction system resides in three critical files:

When CompactionStrategy::evict and CompactionStrategy::retain are combined via min(), the system calculates the safe range of messages that can be summarized while respecting both the percentage limit and the absolute retention guarantee.

Summary

  • token_threshold triggers compaction when the total token count exceeds the configured limit, serving as the primary "when to compact" control.
  • retention_window protects the most recent N messages from summarization, ensuring recent conversation context remains available verbatim.
  • eviction_window limits compaction aggression by capping the percentage of conversation history that can be replaced by summaries.
  • The compaction flow evaluates thresholds in should_compact, constructs eviction and retention strategies, and applies them via intersection logic in Compactor::compact.

Frequently Asked Questions

What happens if token_threshold is set to None?

When token_threshold is None, the token-based trigger is disabled and compaction will not occur due to token count alone. However, compaction may still trigger if other thresholds like turn_threshold or message_threshold are configured, or if the on_turn_end flag is enabled.

How do retention_window and eviction_window interact during compaction?

These two parameters work as constraints on the compaction range. The retention_window guarantees the newest messages are preserved, while the eviction_window limits how much of the older context can be removed. When max is false, the system takes the intersection (eviction.min(retention)) to find the safe range for summarization. When max is true, only the retention window is respected for maximum compaction.

Can I trigger compaction manually without waiting for thresholds?

Yes, you can invoke Compactor::compact(ctx, max) directly at any time, bypassing the should_compact check. Set max to true to compact everything outside the retention window, or false to respect the eviction percentage limit.

Where are the default values for these parameters defined?

Default values for token_threshold, retention_window, and eviction_window are defined in crates/forge_config/src/compact.rs within the Default implementation of the Compact struct. The test suite in crates/forge_config/tests/schema.rs validates that TOML configurations correctly parse percentage values for eviction_window within the 0.0 to 1.0 range.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →