Understanding the Context Compaction Algorithm in Reasonix: Soft vs Hard Thresholds

The context compaction algorithm in Reasonix employs a two-phase threshold system that first trims stale tool outputs at a soft limit, then generates conversation summaries when token usage hits the configurable compact ratio (default 80%).

The Reasonix inference engine, part of the esengine/DeepSeek-Reasonix open-source project, implements a sophisticated context compaction algorithm to manage long-running LLM sessions without exceeding model context windows. This system dynamically monitors token consumption and applies selective pruning or summarization based on configurable ratio thresholds, balancing memory efficiency with prompt cache reuse.

How the Context Compaction Algorithm Works

At its core, the algorithm tracks the relationship between consumed tokens and the total context window, triggering maintenance operations at two distinct watermarks defined in internal/agent/compact.go.

Soft Compaction Phase

When token usage reaches the softCompactRatio (typically set slightly below the main threshold), Reasonix initiates soft compaction. During this phase, the engine identifies and removes older tool-result messages while preserving the cache-first prefix and the canonical transcript. This prevents low-value tool outputs from forcing premature truncation of valuable conversation history.

According to the source implementation at line 122 of internal/agent/compact.go, the system logs this preservation strategy when the threshold is crossed:

// Context reached soft threshold; maintaining cache-first prefix
"context reached …% of window; keeping cache‑first prefix until compact threshold …%"

Hard Compaction Phase

When usage approaches the high-water mark—calculated as high = int(float64(a.contextWindow) * a.compactRatio) at line 89 of internal/agent/compact.go—the algorithm triggers hard compaction. At this threshold, Reasonix discards the oldest user-assistant exchanges and replaces them with a compact-summary generated by the model.

The compaction event records the current threshold percentage as implemented at line 171:

compactRatio: %.0f%% (window: %d)

This summary becomes a system-like message that preserves semantic context while freeing substantial token capacity for new exchanges.

Ratio Thresholds and Cache Behavior

The compact ratio determines how aggressively Reasonix reclaims context space. The system supports a configurable range between 65% and 85%, with distinct behavioral characteristics at each level:

  • 65% (Low): Triggers compaction at approximately 65% of the context window. This aggressive setting removes stale content quickly but may reduce prompt-cache reuse since fewer historical turns remain available for caching.
  • 70% (Earlier): Triggers slightly before the default threshold, balancing early cleanup with moderate cache retention. This corresponds to the "earlier" preset in the Desktop UI.
  • 80% (Default): The standard threshold documented in site/src/pages/docs.astro (lines 211-219). This setting maximizes cache utilization while preventing overflow, keeping the full conversation intact until 80% of the window is consumed.
  • 85% (High/Cache-First): Delays compaction until the context window is nearly full. This favors maximum cache reuse but risks truncation if token spikes occur suddenly.

Lower ratios compact earlier and trim tool results more aggressively, while higher ratios favor the cache-first strategy, retaining historic context longer.

Configuring Compaction in Reasonix

Users can customize the compaction behavior through both the command-line interface and the graphical settings panel.

Command-Line Interface

The reasonix config compact-ratio command manages threshold settings globally or per-project. The implementation resides in internal/cli/cli.go (lines 2366-2704):


# Display current effective ratio and its source

reasonix config compact-ratio

# Set global default (valid range: 65-85)

reasonix config compact-ratio 75

# Override for current project only

reasonix config compact-ratio --local 70

These settings persist in the configuration store and apply to subsequent agent initializations defined in the Agent struct at lines 251-337 of internal/agent/agent.go.

Desktop Application Settings

The Desktop application exposes the compaction controls in desktop/frontend/src/components/SettingsPanel.tsx (lines 4432-4450). The interface provides preset buttons for 70%, 80%, and 85% thresholds, alongside a custom input field. The selected value displays as "Current threshold: {value}%", with the calculated token count shown for the active model's context window (e.g., "≈ 96k tokens for a 128k window").

Source Code Implementation Details

The context compaction algorithm spans several key components:

File Purpose
internal/agent/compact.go Contains the high-water mark calculation at line 89, soft threshold logging at line 122, and compaction metrics recording at line 171
internal/agent/agent.go Defines the Agent struct storing compactRatio and contextWindow state at lines 251-337
internal/cli/cli.go Implements config compact-ratio command handling at lines 2366-2704
desktop/frontend/src/components/SettingsPanel.tsx Renders the ratio selection UI with preset buttons at lines 4432-4450
site/src/pages/docs.astro Documents the default 80% threshold and valid configuration range at lines 211-219

Summary

  • The context compaction algorithm in Reasonix uses a two-phase approach: soft compaction cleans tool results at an early threshold, while hard compaction summarizes conversation history at the main ratio.
  • The high-water mark calculation high = int(float64(a.contextWindow) * a.compactRatio) at line 89 of internal/agent/compact.go determines when hard compaction triggers, defaulting to 80% of the context window.
  • Ratio thresholds range from 65% (aggressive early compaction) to 85% (cache-first delayed compaction), directly impacting prompt-cache reuse efficiency.
  • Configuration is available via reasonix config compact-ratio in the CLI or through preset buttons in the Desktop Settings Panel.
  • Source implementation resides primarily in internal/agent/compact.go, with state management in internal/agent/agent.go.

Frequently Asked Questions

What is the default compact ratio in Reasonix?

The default compact ratio is 80% of the model's context window. According to the documentation in site/src/pages/docs.astro at lines 211-219, this default balances cache reuse with overflow protection, though users may adjust this value between 65% and 85% based on their specific workload requirements.

How does soft compaction differ from hard compaction?

Soft compaction triggers at the softCompactRatio threshold and removes old tool-result messages while keeping the conversation transcript intact. Hard compaction triggers at the main compactRatio threshold (calculated at line 89 of internal/agent/compact.go) and generates a summary of the oldest exchanges before truncating the canonical history. This two-tier system prevents unnecessary summarization while removing low-value outputs early.

What is the impact of setting the compact ratio to 85%?

Setting the ratio to 85% enables a cache-first strategy that delays compaction until the context window is nearly full. While this maximizes prompt-cache reuse by retaining more historical turns, it increases the risk of sudden context overflow if token consumption spikes unexpectedly. This setting is ideal for short, high-cache-reuse sessions rather than long-running tool-heavy conversations.

Where is the compaction threshold calculated in the source code?

The threshold calculation occurs in internal/agent/compact.go at line 89, where the engine computes high = int(float64(a.contextWindow) * a.compactRatio). This integer value represents the token limit that triggers hard compaction. The Agent struct maintains these values in internal/agent/agent.go at lines 251-337, storing both the contextWindow size and the configured compactRatio.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →