How to Use kimi-cli for Text Summarization: A Complete Guide

You can summarize text in kimi-cli by invoking the /compact slash command during an interactive session or by piping text into the CLI with the --print flag, which triggers the Compaction engine in src/kimi_cli/soul/compaction.py to generate concise LLM-powered summaries.

The MoonshotAI/kimi-cli repository provides a Python-based interactive agent that performs text summarization through a context-compaction mechanism. This architecture condenses conversation history or piped input into concise summaries by sending specialized system prompts to the underlying LLM provider.

Understanding the Summarization Architecture

The text summarization capability in kimi-cli is implemented as a context-compaction step that asks the LLM to produce a concise summary and replaces the original messages with that summary. According to the source code in src/kimi_cli/soul/compaction.py, the flow involves several key components:

Interactive Text Summarization with /compact

The most common way to use kimi-cli for text summarization is through the interactive TUI (Terminal User Interface). After starting a conversation, you can condense the entire discussion history using the /compact command.


# Start an interactive session

$ kimi
> Tell me about the history of artificial intelligence.
> (…conversation continues…)

# Generate a summary of the discussion

> /compact
Compacting…   # spinner appears

[summary] The discussion covered early AI milestones, the rise of neural networks, and recent breakthroughs in large language models.

When you issue /compact, the following occurs in the source code:

  1. The input is caught by the dispatcher in src/kimi_cli/soul/slash.py.
  2. The compact() function calls KimiSoul.compact_context().
  3. The SimpleCompaction class in src/kimi_cli/soul/compaction.py sends the system prompt to the LLM.
  4. The UI layer in src/kimi_cli/ui/shell/visualize/_live_view.py renders the spinner and result.

One-Shot Summarization from the Command Line

For non-interactive use, you can pipe text directly into kimi-cli and request a summary in a single command. This approach uses the --print flag to output the result without launching the full TUI.


# Echo a long paragraph and summarize it immediately

$ echo "Long text about quantum computing principles..." | kimi --model gpt-4o --print "Summarize the above text."
Compacting…
[summary] The text describes fundamental quantum computing concepts including superposition, entanglement, and quantum gates.

Internally, the CLI builds a temporary session, injects your prompt, and automatically triggers the compaction step before exiting. This leverages the same Compaction engine found in src/kimi_cli/soul/compaction.py but without the interactive overhead.

Custom Summarization with Specific Instructions

You can guide the summarization process by providing custom instructions after the /compact command. This allows you to focus on specific topics or exclude irrelevant information.

$ kimi
> /compact keep the discussion about "neural networks" and drop the rest
Compacting…
[summary] The conversation focused on the evolution of neural networks, their applications, and recent research trends.

The custom text following /compact is passed as the custom_instruction parameter to KimiSoul.compact_context() in src/kimi_cli/soul/kimisoul.py. The compaction prompt incorporates this instruction, instructing the LLM to prioritize the requested topics during summarization.

Structured Summarization via the Plan Tool

For complex summarization tasks, you can use the built-in planning tool to create a structured workflow that includes a summarization step.

$ kimi
> /plan Summarize a research article

# The plan tool proposes steps:

1️⃣ Read the article
2️⃣ Summarize key findings
3️⃣ Output the summary

# Accept the plan → the agent runs the steps

The plan tool is implemented in src/kimi_cli/tools/plan/__init__.py, where each step's description field (see line 44) can request concise summaries. This method is useful when you need to combine summarization with other operations like file reading or data extraction.

Summary

  • Primary Method: Use the /compact slash command in interactive mode to summarize conversation history.
  • Non-Interactive Mode: Pipe text into kimi --print for one-shot summarization without launching the TUI.
  • Customization: Append specific instructions to /compact to focus on particular topics or themes.
  • Architecture: The summarization relies on the Compaction class in src/kimi_cli/soul/compaction.py and the COMPACTION_SYSTEM_PROMPT to generate condensed content.
  • Implementation: Key functions include compact() in src/kimi_cli/soul/slash.py and compact_context() in src/kimi_cli/soul/kimisoul.py at line 1412.

Frequently Asked Questions

How does the /compact command differ from standard summarization tools?

The /compact command is deeply integrated into kimi-cli's context management system. Unlike external summarization tools, it replaces the current conversation history with the generated summary in src/kimi_cli/soul/kimisoul.py, effectively resetting the token count while preserving semantic meaning. This allows for extended conversations without hitting context limits.

Can I summarize external files without copying text manually?

Yes. You can pipe file contents directly into kimi-cli using standard Unix pipes: cat article.txt | kimi --print "Summarize this". The CLI treats the piped content as user input and processes it through the same compaction pipeline defined in src/kimi_cli/soul/compaction.py.

Which model performs the actual summarization?

The summarization is handled by the LLM provider configured in src/kimi_cli/llm.py, which interfaces with the kosong package. You can specify the model using the --model flag (e.g., --model gpt-4o), and this selection propagates through to the compaction engine when generating summaries.

What happens to the original conversation after compaction?

According to the implementation in src/kimi_cli/soul/compaction.py, the original messages are replaced by the summary output. The SimpleCompaction class constructs a new context containing only the system prompt and the generated summary, which becomes the new conversation history for subsequent interactions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →