# How to Stream Responses in Kimi CLI: A Complete Guide to Real-Time Output

> Learn how to stream responses in Kimi CLI for real-time LLM output. This guide explains the stream parameter and configuration options for seamless interaction with the MoonshotAI kimi-cli.

- Repository: [Moonshot AI/kimi-cli](https://github.com/MoonshotAI/kimi-cli)
- Tags: how-to-guide
- Published: 2026-07-25

---

**Kimi CLI streams LLM responses by forwarding wire messages from the JSON-RPC server to the Rich Live UI in real-time, controlled via the `stream` parameter and configurable with the `--no-stream` flag or `show_thinking_stream` option.**

The `MoonshotAI/kimi-cli` repository provides a terminal-based client for interacting with Kimi AI models. By default, the tool implements **streaming output** to display responses character-by-character rather than waiting for the entire generation to complete. This behavior is implemented through an asynchronous pipeline that bridges the core runtime with the terminal interface.

## How Streaming Works Under the Hood

The streaming architecture in Kimi CLI consists of two primary layers: the server-side message dispatcher and the client-side rendering engine.

### The JSON-RPC Server Layer

Streaming originates in [`src/kimi_cli/wire/server.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/wire/server.py), where the `WireServer` class manages the connection between the LLM runtime and the user interface. When processing a prompt, the server invokes `KimiSoul.run()` followed by `_stream_wire_messages()` (see the `while True` loop around line 21).

The server determines whether to stream based on the `self._is_streaming` flag, which evaluates to `True` when `self._cancel_event is not None`. This flag is set automatically when the client sends a `PromptMessage` containing `"stream": true`. Inside the streaming loop, each incoming `Wire` event—such as `ToolCallRequest`, `ApprovalRequest`, `QuestionRequest`, or generic UI messages—is immediately forwarded to the client via:

```python
await self._send_msg(JSONRPCEventMessage(...))

```

Because this runs in an async loop, messages arrive at the UI as soon as they are produced, eliminating wait times for full response generation.

### The UI Rendering Layer

On the display side, [`src/kimi_cli/ui/shell/visualize/_live_view.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/ui/shell/visualize/_live_view.py) implements the live terminal updates using **Rich Live**. This component maintains an internal buffer (`self._streaming_text`) that accumulates new content as it arrives. The `append` method (approximately line 89) adds each chunk to the buffer and triggers a re-render of the markdown or plain text.

Supporting this is [`src/kimi_cli/ui/shell/visualize/_blocks.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/ui/shell/visualize/_blocks.py), which defines the renderable blocks composing the output. Together, these files ensure that streamed content appears in-place in your terminal, creating the appearance of real-time typing.

## Configuring Streaming Behavior

While streaming is the default mode, Kimi CLI provides multiple mechanisms to control or disable it.

### Disabling Streaming via CLI Flag

To turn off streaming for a specific session, use the `--no-stream` flag. This option is defined in [`src/kimi_cli/cli/__init__.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/cli/__init__.py) and passed to the runtime configuration:

```bash
kimi --no-stream

```

When disabled, the CLI waits for the complete LLM response before displaying anything, simulating a traditional request-response pattern.

### Configuring the Thinking Spinner

You can control the visibility of the "Thinking…" spinner that appears during streaming by modifying the `show_thinking_stream` option in [`src/kimi_cli/config.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/config.py). When set to `true`, the spinner provides visual feedback while the model generates tokens; when `false`, the spinner hides but content still streams.

Set this option in your configuration file at `~/.kimi/config.toml`:

```toml
[runtime]
show_thinking_stream = false

```

## Practical Usage Examples

### Example 1: Interactive Shell (Default Streaming)

Launch the interactive shell to see streaming in its default state:

```bash
kimi

```

Type your prompt. You will see the `[Thinking…]` spinner followed by the response appearing word-by-word as the model generates it.

### Example 2: Disable Streaming for a Single Session

Run a specific command without streaming to receive the complete output at once:

```bash
kimi --no-stream

```

### Example 3: Programmatic Streaming via JSON-RPC

For custom clients connecting to the Kimi CLI server, enable streaming by setting `"stream": true` in your JSON-RPC payload:

```python
import aiohttp
import json
import asyncio

async def stream_prompt():
    async with aiohttp.ClientSession() as session:
        ws = await session.ws_connect("http://localhost:8000/session/123/stream")
        await ws.send_json({
            "jsonrpc": "2.0",
            "method": "prompt",
            "params": {
                "prompt": "Explain quantum entanglement.",
                "stream": True  # Enable streaming

            },
            "id": 1
        })
        
        async for msg in ws:
            data = json.loads(msg.data)
            # Handle incremental events: ToolCallRequest, ApprovalRequest, etc.

            print(data["result"]["content"], end="", flush=True)

asyncio.run(stream_prompt())

```

### Example 4: Permanent Configuration Change

To permanently disable the thinking spinner across all sessions:

```toml

# ~/.kimi/config.toml

[runtime]
show_thinking_stream = false

```

## Summary

- **Streaming is default** in Kimi CLI, implemented through [`src/kimi_cli/wire/server.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/wire/server.py) using `_stream_wire_messages()` to forward `Wire` events as they occur.
- **Disable per-session** with the `--no-stream` flag defined in [`src/kimi_cli/cli/__init__.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/cli/__init__.py).
- **UI updates** are handled by Rich Live in [`src/kimi_cli/ui/shell/visualize/_live_view.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/ui/shell/visualize/_live_view.py), which appends chunks via the `append` method around line 89.
- **Control visual feedback** using the `show_thinking_stream` boolean in [`src/kimi_cli/config.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/config.py), configurable via `~/.kimi/config.toml`.
- **Programmatic access** requires setting `"stream": true` in JSON-RPC `PromptMessage` payloads to maintain the `self._is_streaming` flag.

## Frequently Asked Questions

### Does Kimi CLI stream responses by default?

Yes, streaming is enabled by default in interactive mode. The `WireServer` class in [`src/kimi_cli/wire/server.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/wire/server.py) automatically sets the `self._is_streaming` flag when processing prompts, causing the `_stream_wire_messages()` method to forward each token to the UI as it is generated.

### How do I disable streaming in Kimi CLI?

Pass the `--no-stream` flag when launching the CLI, as implemented in [`src/kimi_cli/cli/__init__.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/cli/__init__.py). This prevents the server from entering streaming mode, causing the application to wait for the complete LLM response before displaying any output.

### What is the `show_thinking_stream` option?

The `show_thinking_stream` option in [`src/kimi_cli/config.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/config.py) controls visibility of the "Thinking…" spinner during generation. When set to `true` (default), you see a visual indicator while the model streams; when `false`, the spinner hides but text continues to stream in real-time. Configure this in `~/.kimi/config.toml` under the `[runtime]` section.

### Can I use streaming with the JSON-RPC API?

Yes, include `"stream": true` in the `params` object of your JSON-RPC `PromptMessage`. This triggers the streaming logic in [`src/kimi_cli/wire/server.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/wire/server.py), causing the server to send `JSONRPCEventMessage` objects incrementally via WebSocket rather than returning a single complete response.