# Kimi CLI Context Compaction Mechanism: Summarizing History in `src/kimi_cli/soul/compaction.py`

> Discover the Kimi CLI context compaction mechanism in src/kimi_cli/soul/compaction.py. Learn how it summarizes history to respect model limits while retaining key exchanges.

- Repository: [Moonshot AI/kimi-cli](https://github.com/MoonshotAI/kimi-cli)
- Tags: internals
- Published: 2026-07-21

---

**The Kimi CLI context compaction mechanism monitors conversation token usage and automatically summarizes older messages into a concise summary when the history approaches the model's context limit, preserving the most recent exchanges in the new shortened history.**

The MoonshotAI/kimi-cli repository relies on an automated context compaction system to keep long-running agent conversations within strict LLM token budgets. The Kimi CLI context compaction mechanism is implemented primarily in [`src/kimi_cli/soul/compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/compaction.py), where a combination of heuristic estimation, configurable triggers, and LLM-driven summarization replaces aging dialogue with a compact summary.

## Core Steps of the Kimi CLI Context Compaction Mechanism

Before any summarization occurs, the runtime approximates token consumption with functions defined in [`src/kimi_cli/soul/compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/compaction.py). According to the MoonshotAI/kimi-cli source code, the system uses a cheap heuristic for sizing and a dual-condition guard for triggering.

### Estimating usage with `estimate_text_tokens`

The `estimate_text_tokens` function iterates over every `Message` in the conversation, inspects its `TextPart` elements, and divides the total character count by four to derive a fast token estimate. This avoids expensive tokenizer calls during the decision phase.

```python
token_count = estimate_text_tokens(context.messages)

```

### Deciding when to compact via `should_auto_compact`

The `should_auto_compact` function returns `True` when either:

- `token_count >= max_context * trigger_ratio` (default ratio `0.9`), **or**
- `token_count + reserved_context_size >= max_context` (default reserved buffer `200`).

These two guards ensure the model never exceeds its window while leaving a safety margin for the next reply.

```python
if should_auto_compact(token_count, max_context, trigger_ratio=0.9,
                       reserved_context_size=200):
    # perform compaction

```

## The `SimpleCompaction` Implementation in [`compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/compaction.py)

The default implementation of the `Compaction` protocol is `SimpleCompaction`, located in [`src/kimi_cli/soul/compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/compaction.py). It handles message selection, prompt construction, LLM invocation, and result packaging.

### Selecting messages with `SimpleCompaction.prepare`

`SimpleCompaction.prepare` walks the history backwards, counting user and assistant messages until it reaches `max_preserved_messages` (default `2`). Everything older than that becomes the `to_compact` slice; the remainder is `to_preserve`. For each message in the compaction slice, a `TextPart` is added that records the role and content. The method then appends the static `COMPACT` template from `kimi_cli.prompts` and any custom instruction supplied by the caller.

### Summarizing with the LLM and building the result

The `compact` method sends the prepared compaction prompt to the LLM via `kosong.step` with:

- system prompt = `"You are a helpful assistant that compacts conversation context."`
- empty toolset (no tool calls allowed during summarization)
- a single-message history consisting of the generated compact-message

The LLM response is filtered to drop any `ThinkPart` artifacts. A new user message is built with the fixed prefix `COMPACTION_OUTPUT_PREFIX` followed by the summary text. The preserved messages are appended unchanged, and the method returns a `CompactionResult` that stores the new message list, exact token usage reported by the provider, and an optional trace ID for debugging.

```python
CompactionResult(messages=compacted_messages, usage=result.usage, ...)

```

## Practical Usage Example

Developers can invoke the Kimi CLI context compaction mechanism directly when testing custom policies or building alternative runtimes. The following pattern estimates tokens, evaluates the trigger, and runs `SimpleCompaction`:

```python
from kimi_cli.soul.compaction import SimpleCompaction, should_auto_compact
from kimi_cli.llm import LLM
from kimi_cli.soul.message import Message

async def maybe_compact(messages: list[Message], llm: LLM, max_ctx: int) -> list[Message]:
    # Quick token estimate

    token_count = sum(len(p.text) for m in messages for p in m.content if isinstance(p, TextPart))
    if should_auto_compact(token_count, max_ctx, trigger_ratio=0.85, reserved_context_size=250):
        compactor = SimpleCompaction(max_preserved_messages=2)
        result = await compactor.compact(messages, llm)
        return list(result.messages)
    return messages

```

In the main runtime located at [`src/kimi_cli/soul/kimisoul.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/kimisoul.py), this step is invoked automatically when the token budget is exceeded.

## Summary

- [`src/kimi_cli/soul/compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/compaction.py) hosts the Kimi CLI context compaction mechanism, combining estimation, trigger logic, and LLM summarization.
- `estimate_text_tokens` provides a fast heuristic by dividing character counts by four.
- `should_auto_compact` uses a ratio-based threshold and a reserved buffer to decide when compaction is necessary.
- `SimpleCompaction.prepare` preserves the most recent two user-assistant exchanges by default and bundles older messages for summarization.
- The compaction prompt is drawn from the `COMPACT` template in `kimi_cli.prompts`, and the resulting summary becomes the opening message of the shortened history.

## Frequently Asked Questions

### What triggers the Kimi CLI context compaction mechanism?

The mechanism triggers when `should_auto_compact` detects that the estimated token count has reached at least `0.9` of the model's maximum context window, or when the current count plus a 200-token reserved buffer would exceed the limit. These thresholds are configurable via the `trigger_ratio` and `reserved_context_size` parameters.

### How does `SimpleCompaction` decide which messages to keep?

`SimpleCompaction.prepare` traverses the conversation history from newest to oldest and retains up to `max_preserved_messages` user and assistant exchanges, which defaults to two. Every older message is moved into the compaction slice and replaced by a single LLM-generated summary.

### Where is the compaction system prompt defined?

The compaction call uses the static system prompt `"You are a helpful assistant that compacts conversation context."` inside [`src/kimi_cli/soul/compaction.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/compaction.py). The instruction template appended to the compaction payload is imported as `COMPACT` from `kimi_cli.prompts`, which guides how the LLM should format the summary.

### Can developers customize the compaction behavior?

Yes. Developers can instantiate `SimpleCompaction` with a different `max_preserved_messages` value or call `should_auto_compact` with custom `trigger_ratio` and `reserved_context_size` arguments. For deeper changes, the `Compaction` protocol can be implemented with an alternative class that replaces the default summarization strategy.