# KimiSoul Agent Loop Architecture: Inside MoonshotAI's Kimi CLI Agent Runtime

> Explore the KimiSoul agent loop architecture in MoonshotAI's Kimi CLI. Understand its async step-based generator for LLM inference, tool execution, and context management.

- Repository: [Moonshot AI/kimi-cli](https://github.com/MoonshotAI/kimi-cli)
- Tags: internals
- Published: 2026-07-22

---

**The KimiSoul agent loop is an async step-based generator that orchestrates LLM inference, tool execution, dynamic context injection, and automatic token compaction within the Kimi CLI runtime.**

The `KimiSoul` class in [`src/kimi_cli/soul/kimisoul.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/kimisoul.py) serves as the central orchestrator for the MoonshotAI/kimi-cli agent runtime. This Python module implements a sophisticated turn-based architecture where each user interaction triggers a **step loop** that manages context windows, executes tools, and handles dynamic prompt injections until a final response is produced.

## Core Components and Class Structure

### The KimiSoul Class Initialization

According to the MoonshotAI/kimi-cli source code, the `KimiSoul` class (lines 28-38) acts as the runtime container holding references to the `Agent`, `Runtime`, and `Context` objects. The `__init__` method (lines 31-84) performs the heavy lifting of wiring together the execution environment:

- Instantiates the `Agent` and `Runtime` interfaces
- Registers **dynamic injection providers** (e.g., `AfkModeInjectionProvider`, `PlanModeInjectionProvider`)
- Initializes the `HookEngine` for extensibility
- Builds the slash-command registry via `_build_slash_commands` (lines 54-78)
- Sets up plan-mode helpers and telemetry trackers

### Public API and Turn Entry Points

The primary entry point for user interactions is the `run` method (lines 60-85). This async method coordinates the pre-turn lifecycle:

1. Creates a temporary approval source for cancellable tool approvals
2. Refreshes OAuth tokens via `self._runtime.oauth.ensure_fresh`
3. Fires the `UserPromptSubmit` hook (unless `skip_user_prompt_hook=True`)
4. Emits the `TurnBegin` event and logs telemetry via `turn_started`
5. Parses slash commands; if detected, executes the matching `SlashCommand` (lines 25-34)
6. Otherwise delegates to `_turn` for standard LLM interaction or `FlowRunner.ralph_loop` for flow-based execution

## The Agent Loop Lifecycle (`_agent_loop`)

The `_agent_loop` method (lines 37-99) is the core async generator implementing the **step loop**. It follows a strict lifecycle documented in its docstring (lines 38-51).

### Turn Initialization Phase

Before entering the iteration cycle, the loop performs setup tasks (lines 58-94):

- Clears stale steer messages from previous turns
- Loads deferred MCP (Model Context Protocol) tools and reports their status
- Prepares the context for the new turn

### The Step Iteration Cycle

Each iteration represents one complete think-act-observe cycle:

- **Guard Check**: Aborts if `max_steps_per_turn` is exceeded (lines 104-106)
- **Step Begin**: Emits the `StepBegin` event (line 110)
- **Context Compaction**: Automatically triggers `should_auto_compact` when token ratios exceed the configured threshold (lines 115-124), invoking `SimpleCompaction` to trim history
- **Checkpoint Creation**: Persists current state via `_checkpoint` (lines 135-137) to enable rollback
- **Step Execution**: Calls `_step()` (line 140) to perform the actual LLM interaction
- **Error Recovery**: Catches `BackToTheFuture` exceptions to revert context and retry (lines 142-176), or raises fatal exceptions

### Outcome Resolution and Turn Termination

After each step, the loop evaluates the `StepOutcome` (lines 179-197). If the `stop_reason` indicates completion (e.g., `"no_tool_calls"`), the generator yields a `TurnOutcome` containing the final assistant message and step count. Otherwise, it consumes pending steers and continues the loop.

## Step Execution Deep Dive (`_step`)

The `_step` method (lines 111-229) implements the eight sub-stages of a single agent step.

### Notification Delivery and Dynamic Injection

For the root session, pending notifications are rendered as messages and passed through a `Notification` hook (lines 133-165). Subsequently, `_collect_injections` (lines 167-178) queries all registered `DynamicInjectionProvider`s to append contextual reminders (like AFK status or plan-mode instructions) as ephemeral user messages.

### LLM Invocation and Retry Logic

The actual inference is handled by `_run_step_once`, wrapped with Tenacity retry logic (lines 192-216). This layer manages transient network failures, timeouts, and exponential backoff before giving up.

### Tool Execution and Context Growth

Following a successful LLM call:

1. `KimiToolset.begin_step` resets per-step tool state
2. Tool calls are dispatched and results awaited (lines 232-260)
3. Assistant messages and tool results are appended to the context (lines 262-274)
4. Usage statistics trigger `StatusUpdate` events for the UI (lines 218-230)

If a `BackToTheFuture` exception is raised during this process, the context reverts to the previous checkpoint and the step retries.

## Supporting Systems

### Dynamic Injection Architecture

Injection providers are registered during initialization (lines 70-80) and invoked during `_collect_injections`. These providers receive lifecycle callbacks such as `_notify_injection_providers_compacted` when context trimming occurs, allowing them to adjust injected content based on the new context window.

### Context Compaction Strategy

When `should_auto_compact` detects excessive token usage (lines 115-122), the loop invokes `self.compact_context()` utilizing the `SimpleCompaction` strategy. Post-compaction, providers are notified to ensure injected reminders remain relevant despite truncated history.

### Slash Command and Hook Integration

The `_build_slash_commands` method aggregates built-in commands from `soul_slash_registry` alongside skill-derived commands (prefixed with `skill:` and `flow:`). The `HookEngine` provides extension points at `UserPromptSubmit`, `Stop`, `StopFailure`, and `Notification` events, allowing external code to intercept or modify the agent loop's behavior.

## Code Examples

Below are practical interactions with the KimiSoul agent loop:

```python

# Execute a complete turn from the UI layer

await kimi_soul.run(user_input="list files", skip_user_prompt_hook=False)

```

```python

# Toggle plan-mode manually (typically invoked via slash command)

await kimi_soul.toggle_plan_mode_from_manual()

```

```python

# Inject a steering message to guide the next step

kimi_soul.steer("Please clarify the last request")

```

Each call flows through the architectural components described above, with `steer` messages being consumed at the outcome resolution phase of `_agent_loop`.

## Summary

- The **KimiSoul agent loop** in [`src/kimi_cli/soul/kimisoul.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/kimisoul.py) uses an async generator pattern (`_agent_loop`) to process turns through discrete steps.
- Each **step** (`_step`) handles dynamic injection, LLM invocation with Tenacity retries, tool execution, and context growth.
- **Context compaction** triggers automatically based on token ratios, using `SimpleCompaction` and notifying injection providers of history changes.
- The **checkpoint system** enables rollback via `BackToTheFuture` exceptions when steps fail or need reversion.
- **Slash commands** and the **HookEngine** provide extensibility points at the turn and step boundaries.

## Frequently Asked Questions

### What triggers the KimiSoul agent loop to terminate a turn?

The loop terminates when `_step` returns a `StepOutcome` with a `stop_reason` indicating completion, such as `"no_tool_calls"`. The `_agent_loop` generator then yields a `TurnOutcome` containing the final assistant message and the total step count executed during that turn.

### How does KimiSoul handle context window limitations?

The loop monitors token usage via `should_auto_compact` (lines 115-124). When the ratio exceeds the configured trigger, it automatically invokes `SimpleCompaction` to trim the context history. Injection providers are notified via `_notify_injection_providers_compacted` so they can adjust their reminders to fit the reduced window.

### What is the purpose of the BackToTheFuture exception?

`BackToTheFuture` is an internal exception type used for **context rollback**. When raised during `_step`, the agent loop catches it (lines 142-176), reverts the context to the last `_checkpoint`, and retries the step. This mechanism handles transient failures or D-Mail style reversion requests without corrupting the conversation state.

### How are dynamic prompts injected during the agent loop?

Dynamic prompts are injected via the `_collect_injections` method (lines 167-178), which queries registered `DynamicInjectionProvider` instances. These providers (such as `AfkModeInjectionProvider` or `PlanModeInjectionProvider`) return reminder text that is appended as ephemeral user messages before the LLM call, ensuring the model receives current operational context like AFK status or active plans.