Kimi CLI Print Mode Architecture: How Non-Interactive Output Works

Kimi CLI’s print mode streams structured events from the core Wire object to specialized printer classes that format output as either plain text or JSON, enabling deterministic, pipe-friendly automation.

The MoonshotAI/kimi-cli repository implements a dual-interface architecture that separates interactive shell experiences from scriptable, non-interactive workflows. When running in print mode, the CLI bypasses the rich terminal UI and instead processes WireMessage events through configurable formatters located in src/kimi_cli/ui/print/visualize.py, producing line-oriented output suitable for CI pipelines and downstream tooling.

Core Components of the Print Mode Pipeline

The architecture centers on four printer implementations and a central visualize coroutine that orchestrates the flow of messages from the runtime to stdout.

Printer Classes and Output Formats

The system selects a printer implementation based on the --output-format (text or stream-json) and --final-only CLI flags parsed in src/kimi_cli/cli/__init__.py:

  • TextPrinter – Forwards every WireMessage directly to rich.print for immediate plain-text rendering.
  • JsonPrinter – Buffers content parts, tool calls, and notifications, emitting well-formed JSON objects only at safe boundaries (e.g., after a tool result completes).
  • FinalOnlyTextPrinter and FinalOnlyJsonPrinter – Discard intermediate reasoning steps and tool executions, preserving only the final assistant message for concise output.

The selection logic is implemented in the visualize coroutine at lines 71–82 of src/kimi_cli/ui/print/visualize.py.

The Wire Object and Message Stream

At the heart of the system is the Wire object defined in src/kimi_cli/soul/wire.py, which acts as the shared channel between the AI runtime and the UI layer. In print mode, the visualizer consumes from wire.ui_side(merge=True), creating a merged UI stream that collapses duplicate events such as incremental content updates into coherent deltas.

How the Print Mode Pipeline Processes Events

The workflow follows a four-stage pipeline that transforms raw wire events into formatted output.

1. Stream Initialization and Printer Selection

When visualize is invoked, it first instantiates the appropriate printer class based on the OutputFormat enum and final_only boolean. This occurs in the initialization block of the coroutine (lines 71–82).

2. Asynchronous Message Consumption

The coroutine enters an asynchronous loop consuming from the merged UI stream (lines 84–92). Each iteration retrieves a WireMessage subclass instance defined in src/kimi_cli/wire/types.py and routes it to the printer’s feed method.

3. Message Type Handling and Buffering

The printer’s feed method pattern-matches on message types to manage internal state:

  • ContentPart – Text deltas are appended to _content_buffer using the _merge_content helper (lines 29–32) to combine incremental updates into coherent blocks.
  • ToolCall and ToolCallPart – Stored in _tool_call_buffer and merged into the most recent active tool invocation.
  • ToolResult and PlanDisplay – Flushed immediately as discrete JSON lines (lines 77–84), ensuring external tools receive complete structural data without waiting for the full conversation to finish.
  • Notification – Buffered until the surrounding assistant message completes (lines 58–66), preventing interleaved output that would break JSON validity.
  • StepBegin, StepInterrupted, and StepRetry – Trigger boundary actions; StepBegin flushes pending assistant messages, while interruptions and retries discard partially-built state (lines 54–58 and 89–94).

4. Final Flush and Shutdown

When the runtime signals completion via QueueShutDown (defined in src/kimi_cli/utils/aioqueue.py), the visualizer calls the printer’s flush method (lines 87–89). This emits any remaining buffered assistant messages and pending notifications, ensuring no data is lost before the process exits.

Usage Examples

Command-Line Interface

Select print mode through CLI flags to control verbosity and format:


# Plain-text, full stream (default print mode)

kimi mytask --output-format=text

# Structured JSON stream (each line is a parseable JSON object)

kimi mytask --output-format=stream-json

# Only the final assistant reply, as plain text

kimi mytask --final-only --output-format=text

# Final reply as JSON for downstream parsing

kimi mytask --final-only --output-format=stream-json

Programmatic Integration

Import the visualization pipeline directly for embedding in Python automation:

from kimi_cli.cli import OutputFormat
from kimi_cli.soul.wire import Wire
from kimi_cli.ui.print.visualize import visualize

# wire is the shared channel that the runtime writes WireMessage objects to

# output_format can be "text" or "stream-json"

# final_only mirrors the CLI flag --final-only

await visualize(
    output_format=OutputFormat("stream-json"),
    final_only=True,
    wire=wire
)

Summary

  • Printer Strategy Pattern: The architecture uses four specialized printer classes (TextPrinter, JsonPrinter, and their FinalOnly variants) selected via CLI flags in src/kimi_cli/ui/print/visualize.py.
  • Event-Driven Streaming: The visualize coroutine consumes merged WireMessage events from wire.ui_side(merge=True), handling content, tool calls, and lifecycle signals asynchronously.
  • Smart Buffering: Content parts merge incrementally, notifications buffer until message completion, and tool results flush immediately to maintain valid JSON boundaries.
  • Clean Shutdown: The QueueShutDown sentinel triggers a final flush, ensuring all buffered output reaches stdout before process termination.

Frequently Asked Questions

What output formats does Kimi CLI print mode support?

Kimi CLI supports two primary output formats controlled by the --output-format flag: text for human-readable plain text rendered via rich.print, and stream-json for machine-parseable JSON lines. Additionally, the --final-only flag restricts output to the final assistant message, available in both text and JSON variants.

How does Kimi CLI handle streaming JSON output?

The JsonPrinter class buffers content parts and tool interactions internally, emitting JSON objects only at safe boundaries such as completed tool results or finished assistant messages. This prevents emitting malformed partial JSON during streaming, ensuring every line written to stdout is a valid, parseable object.

What is the difference between regular and final-only print mode?

Regular print mode streams all events including intermediate reasoning steps, tool calls, and notifications. Final-only mode (activated with --final-only) uses FinalOnlyTextPrinter or FinalOnlyJsonPrinter to discard intermediate steps, preserving only the ultimate assistant response. This is optimal for piping results to other commands or extracting answers in automation scripts.

How are tool calls represented in non-interactive output?

Tool calls are captured as ToolCall and ToolCallPart messages, stored in _tool_call_buffer, and merged into the most recent active invocation. When using JSON output, completed tool calls and their results are emitted as structured JSON objects immediately upon completion (lines 77–84 in visualize.py), allowing external tools to react to function execution in real time.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →