KimiSoul Agent Loop Architecture: Inside MoonshotAI's Kimi CLI Agent Runtime

The KimiSoul agent loop is an async step-based generator that orchestrates LLM inference, tool execution, dynamic context injection, and automatic token compaction within the Kimi CLI runtime.

The KimiSoul class in src/kimi_cli/soul/kimisoul.py serves as the central orchestrator for the MoonshotAI/kimi-cli agent runtime. This Python module implements a sophisticated turn-based architecture where each user interaction triggers a step loop that manages context windows, executes tools, and handles dynamic prompt injections until a final response is produced.

Core Components and Class Structure

The KimiSoul Class Initialization

According to the MoonshotAI/kimi-cli source code, the KimiSoul class (lines 28-38) acts as the runtime container holding references to the Agent, Runtime, and Context objects. The __init__ method (lines 31-84) performs the heavy lifting of wiring together the execution environment:

  • Instantiates the Agent and Runtime interfaces
  • Registers dynamic injection providers (e.g., AfkModeInjectionProvider, PlanModeInjectionProvider)
  • Initializes the HookEngine for extensibility
  • Builds the slash-command registry via _build_slash_commands (lines 54-78)
  • Sets up plan-mode helpers and telemetry trackers

Public API and Turn Entry Points

The primary entry point for user interactions is the run method (lines 60-85). This async method coordinates the pre-turn lifecycle:

  1. Creates a temporary approval source for cancellable tool approvals
  2. Refreshes OAuth tokens via self._runtime.oauth.ensure_fresh
  3. Fires the UserPromptSubmit hook (unless skip_user_prompt_hook=True)
  4. Emits the TurnBegin event and logs telemetry via turn_started
  5. Parses slash commands; if detected, executes the matching SlashCommand (lines 25-34)
  6. Otherwise delegates to _turn for standard LLM interaction or FlowRunner.ralph_loop for flow-based execution

The Agent Loop Lifecycle (_agent_loop)

The _agent_loop method (lines 37-99) is the core async generator implementing the step loop. It follows a strict lifecycle documented in its docstring (lines 38-51).

Turn Initialization Phase

Before entering the iteration cycle, the loop performs setup tasks (lines 58-94):

  • Clears stale steer messages from previous turns
  • Loads deferred MCP (Model Context Protocol) tools and reports their status
  • Prepares the context for the new turn

The Step Iteration Cycle

Each iteration represents one complete think-act-observe cycle:

  • Guard Check: Aborts if max_steps_per_turn is exceeded (lines 104-106)
  • Step Begin: Emits the StepBegin event (line 110)
  • Context Compaction: Automatically triggers should_auto_compact when token ratios exceed the configured threshold (lines 115-124), invoking SimpleCompaction to trim history
  • Checkpoint Creation: Persists current state via _checkpoint (lines 135-137) to enable rollback
  • Step Execution: Calls _step() (line 140) to perform the actual LLM interaction
  • Error Recovery: Catches BackToTheFuture exceptions to revert context and retry (lines 142-176), or raises fatal exceptions

Outcome Resolution and Turn Termination

After each step, the loop evaluates the StepOutcome (lines 179-197). If the stop_reason indicates completion (e.g., "no_tool_calls"), the generator yields a TurnOutcome containing the final assistant message and step count. Otherwise, it consumes pending steers and continues the loop.

Step Execution Deep Dive (_step)

The _step method (lines 111-229) implements the eight sub-stages of a single agent step.

Notification Delivery and Dynamic Injection

For the root session, pending notifications are rendered as messages and passed through a Notification hook (lines 133-165). Subsequently, _collect_injections (lines 167-178) queries all registered DynamicInjectionProviders to append contextual reminders (like AFK status or plan-mode instructions) as ephemeral user messages.

LLM Invocation and Retry Logic

The actual inference is handled by _run_step_once, wrapped with Tenacity retry logic (lines 192-216). This layer manages transient network failures, timeouts, and exponential backoff before giving up.

Tool Execution and Context Growth

Following a successful LLM call:

  1. KimiToolset.begin_step resets per-step tool state
  2. Tool calls are dispatched and results awaited (lines 232-260)
  3. Assistant messages and tool results are appended to the context (lines 262-274)
  4. Usage statistics trigger StatusUpdate events for the UI (lines 218-230)

If a BackToTheFuture exception is raised during this process, the context reverts to the previous checkpoint and the step retries.

Supporting Systems

Dynamic Injection Architecture

Injection providers are registered during initialization (lines 70-80) and invoked during _collect_injections. These providers receive lifecycle callbacks such as _notify_injection_providers_compacted when context trimming occurs, allowing them to adjust injected content based on the new context window.

Context Compaction Strategy

When should_auto_compact detects excessive token usage (lines 115-122), the loop invokes self.compact_context() utilizing the SimpleCompaction strategy. Post-compaction, providers are notified to ensure injected reminders remain relevant despite truncated history.

Slash Command and Hook Integration

The _build_slash_commands method aggregates built-in commands from soul_slash_registry alongside skill-derived commands (prefixed with skill: and flow:). The HookEngine provides extension points at UserPromptSubmit, Stop, StopFailure, and Notification events, allowing external code to intercept or modify the agent loop's behavior.

Code Examples

Below are practical interactions with the KimiSoul agent loop:


# Execute a complete turn from the UI layer

await kimi_soul.run(user_input="list files", skip_user_prompt_hook=False)

# Toggle plan-mode manually (typically invoked via slash command)

await kimi_soul.toggle_plan_mode_from_manual()

# Inject a steering message to guide the next step

kimi_soul.steer("Please clarify the last request")

Each call flows through the architectural components described above, with steer messages being consumed at the outcome resolution phase of _agent_loop.

Summary

  • The KimiSoul agent loop in src/kimi_cli/soul/kimisoul.py uses an async generator pattern (_agent_loop) to process turns through discrete steps.
  • Each step (_step) handles dynamic injection, LLM invocation with Tenacity retries, tool execution, and context growth.
  • Context compaction triggers automatically based on token ratios, using SimpleCompaction and notifying injection providers of history changes.
  • The checkpoint system enables rollback via BackToTheFuture exceptions when steps fail or need reversion.
  • Slash commands and the HookEngine provide extensibility points at the turn and step boundaries.

Frequently Asked Questions

What triggers the KimiSoul agent loop to terminate a turn?

The loop terminates when _step returns a StepOutcome with a stop_reason indicating completion, such as "no_tool_calls". The _agent_loop generator then yields a TurnOutcome containing the final assistant message and the total step count executed during that turn.

How does KimiSoul handle context window limitations?

The loop monitors token usage via should_auto_compact (lines 115-124). When the ratio exceeds the configured trigger, it automatically invokes SimpleCompaction to trim the context history. Injection providers are notified via _notify_injection_providers_compacted so they can adjust their reminders to fit the reduced window.

What is the purpose of the BackToTheFuture exception?

BackToTheFuture is an internal exception type used for context rollback. When raised during _step, the agent loop catches it (lines 142-176), reverts the context to the last _checkpoint, and retries the step. This mechanism handles transient failures or D-Mail style reversion requests without corrupting the conversation state.

How are dynamic prompts injected during the agent loop?

Dynamic prompts are injected via the _collect_injections method (lines 167-178), which queries registered DynamicInjectionProvider instances. These providers (such as AfkModeInjectionProvider or PlanModeInjectionProvider) return reminder text that is appended as ephemeral user messages before the LLM call, ensuring the model receives current operational context like AFK status or active plans.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →