# Context Compaction System in Kimi CLI: How and When It Summarizes History

> Understand Kimi CLI's context compaction system. Learn how and when it automatically summarizes older messages to optimize token usage for efficient conversation history management.

- Repository: [Moonshot AI/kimi-cli](https://github.com/MoonshotAI/kimi-cli)
- Tags: internals
- Published: 2026-07-22

---

**The context compaction system automatically summarizes older conversation messages into a compressed system-style summary whenever token usage exceeds configurable thresholds of the model's context window.**

The **MoonshotAI/kimi-cli** repository implements an intelligent context management mechanism that prevents conversations from exceeding LLM context limits. This system monitors token consumption in real-time and triggers **history summarization** to condense older exchanges while preserving recent interactions. Understanding when and how this compaction initiates is essential for developers building custom agents or tuning conversation behavior.

## What Is the Context Compaction System?

The **context compaction system** is a memory management layer that keeps the LLM prompt size within model constraints by compressing historical conversation data. When activated, it replaces aged messages with a concise summary generated by the model itself, maintaining semantic context without the token overhead of full message history.

The architecture centers on three core components in [`src/kimi_cli/soul/compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/compaction.py):

- **`SimpleCompaction`** – Implements the actual summarization protocol by preparing compact messages, invoking the LLM via the Kompos chat provider, and reconstructing the message list as `["system-prefix", <summary>, <preserved_messages>]`.
- **`CompactionResult`** – A data structure holding the new message sequence and token-usage statistics returned by the summarization call.
- **`should_auto_compact`** – A predicate function evaluated on every agent loop iteration to determine if compaction conditions are met.

## When Does History Summarization Initiate?

History summarization triggers automatically inside `KimiSoul._agent_loop` (located in [`src/kimi_cli/soul/kimisoul.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/kimisoul.py)) during step **2c – Context Compaction**. The system checks compaction conditions on every iteration using the `should_auto_compact` function:

```python
if should_auto_compact(
        self._context.token_count_with_pending,
        self._runtime.llm.max_context_size,
        trigger_ratio=self._loop_control.compaction_trigger_ratio,
        reserved_context_size=self._loop_control.reserved_context_size,
    ):
    logger.info("Context too long, compacting...")
    await self.compact_context()

```

The `should_auto_compact` function returns **True** when either of two configurable thresholds is crossed:

### Ratio-Based Trigger

Compaction initiates when the current token count reaches a specified fraction of the model's maximum context size:

```python
token_count >= max_context_size * trigger_ratio

```

This allows proactive compression before hitting hard limits, typically configured between 0.75 and 0.90 depending on conversation volatility.

### Reserved-Size Trigger

Alternatively, compaction fires when the token count plus a safety margin would exceed the maximum context size:

```python
token_count + reserved_context_size >= max_context_size

```

This **reserved buffer** ensures that pending tool calls or assistant responses have sufficient token budget to complete without truncation.

## How the Compaction Workflow Works

Once triggered, the system executes a structured workflow to compress conversation history while maintaining coherence.

### The SimpleCompaction Implementation

The `SimpleCompaction` class handles the transformation logic. It identifies messages to preserve (typically the most recent exchanges), sends the remaining history to the LLM for summarization, and reconstructs the context:

```python
from kimi_cli.soul.compaction import SimpleCompaction, CompactionResult
from kimi_cli.llm import LLM
from kosong.message import Message

# Assume `messages` is a list of Message objects and `llm` is an LLM instance

compactor = SimpleCompaction(max_preserved_messages=2)
result: CompactionResult = await compactor.compact(messages, llm)

# result.messages now contains ["system-prefix", <summary>, <preserved_messages>]

# ready to be fed back to the LLM

```

The implementation preserves the `max_preserved_messages` most recent interactions untouched while compressing everything older into a single system-style summary message.

### Integration in the Agent Loop

Inside `KimiSoul._agent_loop`, the compaction check runs before each LLM inference:

```python

# Inside KimiSoul._agent_loop (step 2c)

if should_auto_compact(
        self._context.token_count_with_pending,
        self._runtime.llm.max_context_size,
        trigger_ratio=self._loop_control.compaction_trigger_ratio,
        reserved_context_size=self._loop_control.reserved_context_size,
    ):
    await self.compact_context()  # delegates to SimpleCompaction

```

This integration ensures that **history summarization** occurs transparently during conversation flow, without requiring explicit user intervention.

## Configuring the Trigger Thresholds

Both compaction triggers are configurable through the `LoopControl` configuration object. Developers can tune these parameters via TOML configuration files:

```toml

# Example snippet from a Kimi CLI TOML config file

[loop_control]
compaction_trigger_ratio = 0.85   # compact when 85% of context is used

reserved_context_size = 5000      # maintain 5k token safety buffer

```

With these settings, the **context compaction system** activates when token usage exceeds **85%** of the model's capacity **or** when fewer than **5,000 tokens** remain available. Adjusting `compaction_trigger_ratio` higher delays summarization (keeping more raw history) but increases the risk of hitting context limits, while lowering it triggers more frequent compaction.

Unit tests in [`tests/core/test_simple_compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/tests/core/test_simple_compaction.py) and [`tests/core/test_context_pending_tokens.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/tests/core/test_context_pending_tokens.py) verify these behaviors across various token-count scenarios and edge cases.

## Summary

- **Context compaction** in Kimi CLI prevents context window overflow by summarizing older messages into compressed representations.
- **History summarization** triggers automatically when token usage exceeds `compaction_trigger_ratio` (default ~0.85) of the max context size or when `reserved_context_size` tokens would be exceeded.
- The **`SimpleCompaction`** class in [`src/kimi_cli/soul/compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/compaction.py) implements the summarization logic, invoked from `KimiSoul._agent_loop` in [`src/kimi_cli/soul/kimisoul.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/kimisoul.py).
- Both thresholds are configurable via `LoopControl` settings, allowing fine-tuning of the trade-off between historical detail and context safety.

## Frequently Asked Questions

### How do I manually trigger context compaction in Kimi CLI?

Developers can invoke compaction programmatically by instantiating `SimpleCompaction` and calling its `compact()` method with the current message list and LLM instance. This returns a `CompactionResult` containing the compressed message sequence ready for prompting.

### What happens to tool calls and function results during compaction?

The `SimpleCompaction` implementation preserves the most recent `max_preserved_messages` (typically 2-4) in their original form, ensuring that pending tool calls and their results remain intact and addressable. Only older exchanges beyond this preservation window get summarized.

### Can I disable automatic history summarization?

While you cannot fully disable compaction without risking context overflow errors, you can effectively prevent it by setting `compaction_trigger_ratio` to 1.0 and `reserved_context_size` to 0. However, this is not recommended for production use as it will cause failures once the conversation exceeds the model's token limit.

### Where does the summarization text come from?

The summary text is generated by the LLM itself through the Kompos chat provider. The `SimpleCompaction` class sends the messages marked for compression to the model with a system prompt requesting a concise summary, then inserts the returned text as a new system message in the reconstructed context.