How to Use the Chat Command in Kimi CLI: Interactive and Print Modes Explained

Kimi CLI does not expose a standalone chat subcommand; instead, you launch an interactive session by running kimi alone or execute a one-off query using the --print flag combined with --prompt.

The MoonshotAI/kimi-cli repository provides a unified interface for interacting with large language models through your terminal. While many CLI tools implement a dedicated chat subcommand, Kimi CLI takes a different approach by offering two distinct entry points—an interactive shell and a print mode—both backed by the same agent architecture in src/kimi_cli/soul/kimisoul.py.

Chat Interface Modes

Kimi CLI offers two primary ways to engage in conversation with the AI, depending on whether you need an ongoing dialogue or a single response.

Interactive Shell Mode (Default)

Running the CLI without arguments launches the full-screen TUI where you can type messages back-and-forth with the agent. This mode maintains conversation context and supports multi-turn interactions.

kimi

Once executed, the terminal transforms into an interactive interface where you can type queries, receive streaming responses, and continue the conversation thread.

One-Off Print Mode

For scripting or quick queries that require only the final answer without UI overhead, use the print mode. This mode sends a single prompt to the LLM, prints the assistant's response, and immediately exits.

kimi --print --prompt "Explain the difference between async and threading in Python."

You can also use the short form flags:

kimi -p "Summarize the latest Kimi CLI release notes."

Internal Implementation of Chat Functionality

Understanding how Kimi CLI processes chat requests requires examining the codebase path from argument parsing to response generation.

Argument Parsing and Mode Selection

The Typer callback in src/kimi_cli/cli/__init__.py handles the --print (aliased as --output-format) and --prompt options. Lines 70-78 define how the CLI distinguishes between interactive and print modes based on these flags. When --print is detected, the ui variable is set to "print", bypassing the TUI initialization.

Executing Single Requests

In src/kimi_cli/app.py, the KimiCLI class implements the run_print method (lines 62-69). This method executes a single turn with the supplied prompt and returns only the final message content, avoiding the overhead of the interactive shell. The implementation ensures that once the LLM returns a response, the process terminates cleanly.

Chat Provider Abstraction

The actual communication with language model APIs occurs through the abstraction layer defined in src/kimi_cli/llm.py (lines 10-20). The ChatProvider class standardizes interactions across different backends—including OpenAI, Anthropic, and Kimi—allowing the CLI to route requests to the configured provider without changing the interface logic.

Core Agent Loop

The underlying agent logic resides in src/kimi_cli/soul/kimisoul.py. In print mode, this module bypasses the interactive input loop and directly processes the single-turn request, returning the assistant's reply to the run_print caller.

Practical Usage Examples

The following patterns cover common scenarios for interacting with Kimi CLI's chat functionality.

1. Start an interactive session

kimi

2. Execute a one-off query with full flags

kimi --print --prompt "What is the capital of France?"

3. Quick queries using short flags

kimi -p "Refactor this Python function to use list comprehensions" < function.py

4. Pipe JSON input for advanced scripting

echo '{"messages":[{"role":"user","content":"What is a monad?"}]}' | kimi --print --input-format stream-json

This approach allows you to construct complex message histories programmatically and pipe them directly into the CLI.

Key Source Files for Chat Operations

File Role in Chat Workflow
src/kimi_cli/cli/__init__.py Defines the Typer CLI interface, parses --print/--prompt flags, and selects the UI mode based on arguments.
src/kimi_cli/app.py Implements the run_print method that processes single chat requests and returns the final message.
src/kimi_cli/llm.py Contains the ChatProvider abstraction that standardizes communication with various LLM services.
src/kimi_cli/soul/kimisoul.py Houses the core agent loop; in print mode, bypasses interactive shells to return direct responses.
src/kimi_cli/config.py Loads configuration parameters including model selection and API keys required by the chat provider.

Summary

  • Kimi CLI has no dedicated chat subcommand—functionality is accessed through the default interactive shell or print mode flags.
  • Use kimi alone to start the full-screen TUI for multi-turn conversations.
  • Use kimi --print --prompt "..." (or -p) for single queries that return only the final answer.
  • Implementation spans multiple modules: argument parsing in cli/__init__.py, execution logic in app.py, and provider abstraction in llm.py.
  • JSON piping is supported via --input-format stream-json for programmatic integration.

Frequently Asked Questions

Is there a dedicated chat subcommand in Kimi CLI?

No. According to the source code in src/kimi_cli/cli/__init__.py, Kimi CLI does not implement a separate chat subcommand. The functionality is integrated into the main entry point, with behavior determined by the presence of flags like --print and --prompt.

How do I run a single query without launching the interactive UI?

Append the --print (or -p) flag combined with --prompt followed by your query. This invokes the run_print method in src/kimi_cli/app.py, which processes the request once and exits without initializing the TUI.

What file handles the chat provider abstraction?

The ChatProvider class defined in src/kimi_cli/llm.py (lines 10-20) handles the abstraction. This class standardizes interactions across different LLM backends, allowing the CLI to support multiple providers through a unified interface.

Can I pipe input to kimi-cli for automated scripts?

Yes. You can pipe JSON-formatted message arrays to the CLI using the --input-format stream-json flag alongside --print. This enables integration with shell scripts and other automation tools that need to process LLM responses programmatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →