How to Control Max Tokens in Kimi CLI: Environment Variables, Context-Aware Logic, and Per-Request Overrides

Kimi CLI determines the maximum completion token budget through a three-tier priority system: environment variable hard caps, automatic context-window calculations, and per-request generation_overrides.

The MoonshotAI/kimi-cli repository implements a sophisticated token management system that prevents context window overflows while giving developers granular control over generation limits. Learning how to control max tokens in kimi-cli is essential for optimizing both short conversational turns and long-form content generation without triggering provider-side errors.

Understanding the Token Budget Hierarchy

Kimi CLI applies token limits in three distinct stages, each overriding the previous if configured.

Model-Level Hard Cap via Environment Variables

The highest priority control uses the KIMI_MODEL_MAX_COMPLETION_TOKENS environment variable (with a compatibility alias KIMI_MODEL_MAX_TOKENS). When set, this value becomes an explicit upper bound for every request.

In src/kimi_cli/llm.py (lines 68–84), the CLI checks for these variables during initialization. Setting the value to 0 or any negative number disables this clamping entirely, allowing the provider to determine the budget. This global setting ensures consistent behavior across all CLI sessions without modifying code.

Context-Aware Automatic Fallback

When no environment variable is set, Kimi CLI automatically calculates a safe token budget to prevent context window violations. The system estimates consumed tokens via estimate_request_tokens, then computes the remaining capacity.

According to the source code in src/kimi_cli/llm.py (lines 81–96), the compute_max_completion_tokens function subtracts the input token count from the model’s max_context_size. The resulting value is applied to the request in src/kimi_cli/soul/kimisoul.py (lines 1368–1387). This default behavior ensures the model never receives a budget larger than the available context window.

Per-Request Generation Overrides

For temporary adjustments, callers can pass a generation_overrides mapping containing "max_completion_tokens" to modify the budget for a single request. The helper function with_kimi_generation_overrides injects this configuration into the provider instance.

As implemented in src/kimi_cli/llm.py (lines 64–71), this approach allows specific agent skills or interactive commands to use different limits without affecting the global configuration.

Practical Implementation Examples

Configure a Global Token Limit

Set a hard cap for all CLI sessions by exporting the environment variable before running commands:


# Limit all responses to 4096 tokens

export KIMI_MODEL_MAX_COMPLETION_TOKENS=4096

# Or disable clamping entirely (allow provider default)

export KIMI_MODEL_MAX_COMPLETION_TOKENS=0

Override Tokens for a Single Request

Use the Python API to temporarily increase the budget for a specific operation:

from kimi_cli.llm import with_kimi_generation_overrides

# Apply 8000-token limit to this provider instance only

overrides = {"max_completion_tokens": 8000}
chat = with_kimi_generation_overrides(chat, overrides)

# Subsequent requests use the 8K budget without affecting global settings

Inspect the Computed Budget

Debug the automatic calculation to understand how much context remains:

from kimi_cli.llm import compute_max_completion_tokens

max_ctx = 8192          # Model's max_context_size

input_toks = 2000       # Tokens consumed by system prompt and history

budget = compute_max_completion_tokens(
    max_context_size=max_ctx,
    input_tokens=input_toks,
    response_budget=None,  # None triggers automatic calculation

)

print(f"Computed max_completion_tokens = {budget}")

# Output: 6192 (8192 - 2000)

Key Source Files and Functions

  • src/kimi_cli/llm.py – Contains compute_max_completion_tokens, with_kimi_generation_overrides, and environment variable parsing logic for token budget management.
  • src/kimi_cli/soul/kimisoul.py – Applies the computed max_completion_tokens to each generation request at lines 1368–1387.
  • docs/en/configuration/env-vars.md – Documents the KIMI_MODEL_MAX_COMPLETION_TOKENS and KIMI_MODEL_MAX_TOKENS variables.
  • src/kimi_cli/agents/spec.yaml – Demonstrates how agent specifications pass generation_overrides to influence token limits for specific skills.

Summary

  • Environment variables provide the highest priority hard cap via KIMI_MODEL_MAX_COMPLETION_TOKENS, where 0 disables clamping.
  • Automatic calculation prevents context overflow by subtracting input tokens from max_context_size when no override is set.
  • Per-request overrides via generation_overrides allow temporary budget adjustments without changing global configuration.
  • The compute_max_completion_tokens function in src/kimi_cli/llm.py serves as the central authority for budget mathematics.

Frequently Asked Questions

What is the difference between KIMI_MODEL_MAX_COMPLETION_TOKENS and KIMI_MODEL_MAX_TOKENS?

KIMI_MODEL_MAX_TOKENS functions as a compatibility alias for KIMI_MODEL_MAX_COMPLETION_TOKENS. Both variables are checked in src/kimi_cli/llm.py, but KIMI_MODEL_MAX_COMPLETION_TOKENS takes precedence if both are defined.

How does kimi-cli prevent exceeding the model's context window?

The CLI calls estimate_request_tokens to measure input size, then compute_max_completion_tokens calculates max_context_size - input_tokens as the response budget. This logic, implemented in src/kimi_cli/llm.py and applied in src/kimi_cli/soul/kimisoul.py, ensures the total prompt plus completion never exceeds the model's limit.

Can I disable the token limit entirely?

Yes. Set KIMI_MODEL_MAX_COMPLETION_TOKENS=0 (or any negative integer) in your environment, or pass {"max_completion_tokens": 0} in the generation_overrides mapping. This disables client-side clamping and delegates the decision to the API provider.

How do I set token limits for specific agent skills?

Agent configurations in src/kimi_cli/agents/spec.yaml can specify generation_overrides including max_completion_tokens. This allows individual skills to use higher or lower limits than the global default without affecting other CLI functionality.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →