Context Compaction System in Kimi CLI: How and When It Summarizes History
The context compaction system automatically summarizes older conversation messages into a compressed system-style summary whenever token usage exceeds configurable thresholds of the model's context window.
The MoonshotAI/kimi-cli repository implements an intelligent context management mechanism that prevents conversations from exceeding LLM context limits. This system monitors token consumption in real-time and triggers history summarization to condense older exchanges while preserving recent interactions. Understanding when and how this compaction initiates is essential for developers building custom agents or tuning conversation behavior.
What Is the Context Compaction System?
The context compaction system is a memory management layer that keeps the LLM prompt size within model constraints by compressing historical conversation data. When activated, it replaces aged messages with a concise summary generated by the model itself, maintaining semantic context without the token overhead of full message history.
The architecture centers on three core components in src/kimi_cli/soul/compaction.py:
SimpleCompaction– Implements the actual summarization protocol by preparing compact messages, invoking the LLM via the Kompos chat provider, and reconstructing the message list as["system-prefix", <summary>, <preserved_messages>].CompactionResult– A data structure holding the new message sequence and token-usage statistics returned by the summarization call.should_auto_compact– A predicate function evaluated on every agent loop iteration to determine if compaction conditions are met.
When Does History Summarization Initiate?
History summarization triggers automatically inside KimiSoul._agent_loop (located in src/kimi_cli/soul/kimisoul.py) during step 2c – Context Compaction. The system checks compaction conditions on every iteration using the should_auto_compact function:
if should_auto_compact(
self._context.token_count_with_pending,
self._runtime.llm.max_context_size,
trigger_ratio=self._loop_control.compaction_trigger_ratio,
reserved_context_size=self._loop_control.reserved_context_size,
):
logger.info("Context too long, compacting...")
await self.compact_context()
The should_auto_compact function returns True when either of two configurable thresholds is crossed:
Ratio-Based Trigger
Compaction initiates when the current token count reaches a specified fraction of the model's maximum context size:
token_count >= max_context_size * trigger_ratio
This allows proactive compression before hitting hard limits, typically configured between 0.75 and 0.90 depending on conversation volatility.
Reserved-Size Trigger
Alternatively, compaction fires when the token count plus a safety margin would exceed the maximum context size:
token_count + reserved_context_size >= max_context_size
This reserved buffer ensures that pending tool calls or assistant responses have sufficient token budget to complete without truncation.
How the Compaction Workflow Works
Once triggered, the system executes a structured workflow to compress conversation history while maintaining coherence.
The SimpleCompaction Implementation
The SimpleCompaction class handles the transformation logic. It identifies messages to preserve (typically the most recent exchanges), sends the remaining history to the LLM for summarization, and reconstructs the context:
from kimi_cli.soul.compaction import SimpleCompaction, CompactionResult
from kimi_cli.llm import LLM
from kosong.message import Message
# Assume `messages` is a list of Message objects and `llm` is an LLM instance
compactor = SimpleCompaction(max_preserved_messages=2)
result: CompactionResult = await compactor.compact(messages, llm)
# result.messages now contains ["system-prefix", <summary>, <preserved_messages>]
# ready to be fed back to the LLM
The implementation preserves the max_preserved_messages most recent interactions untouched while compressing everything older into a single system-style summary message.
Integration in the Agent Loop
Inside KimiSoul._agent_loop, the compaction check runs before each LLM inference:
# Inside KimiSoul._agent_loop (step 2c)
if should_auto_compact(
self._context.token_count_with_pending,
self._runtime.llm.max_context_size,
trigger_ratio=self._loop_control.compaction_trigger_ratio,
reserved_context_size=self._loop_control.reserved_context_size,
):
await self.compact_context() # delegates to SimpleCompaction
This integration ensures that history summarization occurs transparently during conversation flow, without requiring explicit user intervention.
Configuring the Trigger Thresholds
Both compaction triggers are configurable through the LoopControl configuration object. Developers can tune these parameters via TOML configuration files:
# Example snippet from a Kimi CLI TOML config file
[loop_control]
compaction_trigger_ratio = 0.85 # compact when 85% of context is used
reserved_context_size = 5000 # maintain 5k token safety buffer
With these settings, the context compaction system activates when token usage exceeds 85% of the model's capacity or when fewer than 5,000 tokens remain available. Adjusting compaction_trigger_ratio higher delays summarization (keeping more raw history) but increases the risk of hitting context limits, while lowering it triggers more frequent compaction.
Unit tests in tests/core/test_simple_compaction.py and tests/core/test_context_pending_tokens.py verify these behaviors across various token-count scenarios and edge cases.
Summary
- Context compaction in Kimi CLI prevents context window overflow by summarizing older messages into compressed representations.
- History summarization triggers automatically when token usage exceeds
compaction_trigger_ratio(default ~0.85) of the max context size or whenreserved_context_sizetokens would be exceeded. - The
SimpleCompactionclass insrc/kimi_cli/soul/compaction.pyimplements the summarization logic, invoked fromKimiSoul._agent_loopinsrc/kimi_cli/soul/kimisoul.py. - Both thresholds are configurable via
LoopControlsettings, allowing fine-tuning of the trade-off between historical detail and context safety.
Frequently Asked Questions
How do I manually trigger context compaction in Kimi CLI?
Developers can invoke compaction programmatically by instantiating SimpleCompaction and calling its compact() method with the current message list and LLM instance. This returns a CompactionResult containing the compressed message sequence ready for prompting.
What happens to tool calls and function results during compaction?
The SimpleCompaction implementation preserves the most recent max_preserved_messages (typically 2-4) in their original form, ensuring that pending tool calls and their results remain intact and addressable. Only older exchanges beyond this preservation window get summarized.
Can I disable automatic history summarization?
While you cannot fully disable compaction without risking context overflow errors, you can effectively prevent it by setting compaction_trigger_ratio to 1.0 and reserved_context_size to 0. However, this is not recommended for production use as it will cause failures once the conversation exceeds the model's token limit.
Where does the summarization text come from?
The summary text is generated by the LLM itself through the Kompos chat provider. The SimpleCompaction class sends the messages marked for compression to the model with a system prompt requesting a concise summary, then inserts the returned text as a new system message in the reconstructed context.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →