How to Stream Responses in Kimi CLI: A Complete Guide to Real-Time Output

Kimi CLI streams LLM responses by forwarding wire messages from the JSON-RPC server to the Rich Live UI in real-time, controlled via the stream parameter and configurable with the --no-stream flag or show_thinking_stream option.

The MoonshotAI/kimi-cli repository provides a terminal-based client for interacting with Kimi AI models. By default, the tool implements streaming output to display responses character-by-character rather than waiting for the entire generation to complete. This behavior is implemented through an asynchronous pipeline that bridges the core runtime with the terminal interface.

How Streaming Works Under the Hood

The streaming architecture in Kimi CLI consists of two primary layers: the server-side message dispatcher and the client-side rendering engine.

The JSON-RPC Server Layer

Streaming originates in src/kimi_cli/wire/server.py, where the WireServer class manages the connection between the LLM runtime and the user interface. When processing a prompt, the server invokes KimiSoul.run() followed by _stream_wire_messages() (see the while True loop around line 21).

The server determines whether to stream based on the self._is_streaming flag, which evaluates to True when self._cancel_event is not None. This flag is set automatically when the client sends a PromptMessage containing "stream": true. Inside the streaming loop, each incoming Wire event—such as ToolCallRequest, ApprovalRequest, QuestionRequest, or generic UI messages—is immediately forwarded to the client via:

await self._send_msg(JSONRPCEventMessage(...))

Because this runs in an async loop, messages arrive at the UI as soon as they are produced, eliminating wait times for full response generation.

The UI Rendering Layer

On the display side, src/kimi_cli/ui/shell/visualize/_live_view.py implements the live terminal updates using Rich Live. This component maintains an internal buffer (self._streaming_text) that accumulates new content as it arrives. The append method (approximately line 89) adds each chunk to the buffer and triggers a re-render of the markdown or plain text.

Supporting this is src/kimi_cli/ui/shell/visualize/_blocks.py, which defines the renderable blocks composing the output. Together, these files ensure that streamed content appears in-place in your terminal, creating the appearance of real-time typing.

Configuring Streaming Behavior

While streaming is the default mode, Kimi CLI provides multiple mechanisms to control or disable it.

Disabling Streaming via CLI Flag

To turn off streaming for a specific session, use the --no-stream flag. This option is defined in src/kimi_cli/cli/__init__.py and passed to the runtime configuration:

kimi --no-stream

When disabled, the CLI waits for the complete LLM response before displaying anything, simulating a traditional request-response pattern.

Configuring the Thinking Spinner

You can control the visibility of the "Thinking…" spinner that appears during streaming by modifying the show_thinking_stream option in src/kimi_cli/config.py. When set to true, the spinner provides visual feedback while the model generates tokens; when false, the spinner hides but content still streams.

Set this option in your configuration file at ~/.kimi/config.toml:

[runtime]
show_thinking_stream = false

Practical Usage Examples

Example 1: Interactive Shell (Default Streaming)

Launch the interactive shell to see streaming in its default state:

kimi

Type your prompt. You will see the [Thinking…] spinner followed by the response appearing word-by-word as the model generates it.

Example 2: Disable Streaming for a Single Session

Run a specific command without streaming to receive the complete output at once:

kimi --no-stream

Example 3: Programmatic Streaming via JSON-RPC

For custom clients connecting to the Kimi CLI server, enable streaming by setting "stream": true in your JSON-RPC payload:

import aiohttp
import json
import asyncio

async def stream_prompt():
    async with aiohttp.ClientSession() as session:
        ws = await session.ws_connect("http://localhost:8000/session/123/stream")
        await ws.send_json({
            "jsonrpc": "2.0",
            "method": "prompt",
            "params": {
                "prompt": "Explain quantum entanglement.",
                "stream": True  # Enable streaming

            },
            "id": 1
        })
        
        async for msg in ws:
            data = json.loads(msg.data)
            # Handle incremental events: ToolCallRequest, ApprovalRequest, etc.

            print(data["result"]["content"], end="", flush=True)

asyncio.run(stream_prompt())

Example 4: Permanent Configuration Change

To permanently disable the thinking spinner across all sessions:


# ~/.kimi/config.toml

[runtime]
show_thinking_stream = false

Summary

  • Streaming is default in Kimi CLI, implemented through src/kimi_cli/wire/server.py using _stream_wire_messages() to forward Wire events as they occur.
  • Disable per-session with the --no-stream flag defined in src/kimi_cli/cli/__init__.py.
  • UI updates are handled by Rich Live in src/kimi_cli/ui/shell/visualize/_live_view.py, which appends chunks via the append method around line 89.
  • Control visual feedback using the show_thinking_stream boolean in src/kimi_cli/config.py, configurable via ~/.kimi/config.toml.
  • Programmatic access requires setting "stream": true in JSON-RPC PromptMessage payloads to maintain the self._is_streaming flag.

Frequently Asked Questions

Does Kimi CLI stream responses by default?

Yes, streaming is enabled by default in interactive mode. The WireServer class in src/kimi_cli/wire/server.py automatically sets the self._is_streaming flag when processing prompts, causing the _stream_wire_messages() method to forward each token to the UI as it is generated.

How do I disable streaming in Kimi CLI?

Pass the --no-stream flag when launching the CLI, as implemented in src/kimi_cli/cli/__init__.py. This prevents the server from entering streaming mode, causing the application to wait for the complete LLM response before displaying any output.

What is the show_thinking_stream option?

The show_thinking_stream option in src/kimi_cli/config.py controls visibility of the "Thinking…" spinner during generation. When set to true (default), you see a visual indicator while the model streams; when false, the spinner hides but text continues to stream in real-time. Configure this in ~/.kimi/config.toml under the [runtime] section.

Can I use streaming with the JSON-RPC API?

Yes, include "stream": true in the params object of your JSON-RPC PromptMessage. This triggers the streaming logic in src/kimi_cli/wire/server.py, causing the server to send JSONRPCEventMessage objects incrementally via WebSocket rather than returning a single complete response.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →