How to Control Max Tokens in Kimi CLI: Environment Variables, Context-Aware Logic, and Per-Request Overrides
Kimi CLI determines the maximum completion token budget through a three-tier priority system: environment variable hard caps, automatic context-window calculations, and per-request generation_overrides.
The MoonshotAI/kimi-cli repository implements a sophisticated token management system that prevents context window overflows while giving developers granular control over generation limits. Learning how to control max tokens in kimi-cli is essential for optimizing both short conversational turns and long-form content generation without triggering provider-side errors.
Understanding the Token Budget Hierarchy
Kimi CLI applies token limits in three distinct stages, each overriding the previous if configured.
Model-Level Hard Cap via Environment Variables
The highest priority control uses the KIMI_MODEL_MAX_COMPLETION_TOKENS environment variable (with a compatibility alias KIMI_MODEL_MAX_TOKENS). When set, this value becomes an explicit upper bound for every request.
In src/kimi_cli/llm.py (lines 68–84), the CLI checks for these variables during initialization. Setting the value to 0 or any negative number disables this clamping entirely, allowing the provider to determine the budget. This global setting ensures consistent behavior across all CLI sessions without modifying code.
Context-Aware Automatic Fallback
When no environment variable is set, Kimi CLI automatically calculates a safe token budget to prevent context window violations. The system estimates consumed tokens via estimate_request_tokens, then computes the remaining capacity.
According to the source code in src/kimi_cli/llm.py (lines 81–96), the compute_max_completion_tokens function subtracts the input token count from the model’s max_context_size. The resulting value is applied to the request in src/kimi_cli/soul/kimisoul.py (lines 1368–1387). This default behavior ensures the model never receives a budget larger than the available context window.
Per-Request Generation Overrides
For temporary adjustments, callers can pass a generation_overrides mapping containing "max_completion_tokens" to modify the budget for a single request. The helper function with_kimi_generation_overrides injects this configuration into the provider instance.
As implemented in src/kimi_cli/llm.py (lines 64–71), this approach allows specific agent skills or interactive commands to use different limits without affecting the global configuration.
Practical Implementation Examples
Configure a Global Token Limit
Set a hard cap for all CLI sessions by exporting the environment variable before running commands:
# Limit all responses to 4096 tokens
export KIMI_MODEL_MAX_COMPLETION_TOKENS=4096
# Or disable clamping entirely (allow provider default)
export KIMI_MODEL_MAX_COMPLETION_TOKENS=0
Override Tokens for a Single Request
Use the Python API to temporarily increase the budget for a specific operation:
from kimi_cli.llm import with_kimi_generation_overrides
# Apply 8000-token limit to this provider instance only
overrides = {"max_completion_tokens": 8000}
chat = with_kimi_generation_overrides(chat, overrides)
# Subsequent requests use the 8K budget without affecting global settings
Inspect the Computed Budget
Debug the automatic calculation to understand how much context remains:
from kimi_cli.llm import compute_max_completion_tokens
max_ctx = 8192 # Model's max_context_size
input_toks = 2000 # Tokens consumed by system prompt and history
budget = compute_max_completion_tokens(
max_context_size=max_ctx,
input_tokens=input_toks,
response_budget=None, # None triggers automatic calculation
)
print(f"Computed max_completion_tokens = {budget}")
# Output: 6192 (8192 - 2000)
Key Source Files and Functions
src/kimi_cli/llm.py– Containscompute_max_completion_tokens,with_kimi_generation_overrides, and environment variable parsing logic for token budget management.src/kimi_cli/soul/kimisoul.py– Applies the computedmax_completion_tokensto each generation request at lines 1368–1387.docs/en/configuration/env-vars.md– Documents theKIMI_MODEL_MAX_COMPLETION_TOKENSandKIMI_MODEL_MAX_TOKENSvariables.src/kimi_cli/agents/spec.yaml– Demonstrates how agent specifications passgeneration_overridesto influence token limits for specific skills.
Summary
- Environment variables provide the highest priority hard cap via
KIMI_MODEL_MAX_COMPLETION_TOKENS, where0disables clamping. - Automatic calculation prevents context overflow by subtracting input tokens from
max_context_sizewhen no override is set. - Per-request overrides via
generation_overridesallow temporary budget adjustments without changing global configuration. - The
compute_max_completion_tokensfunction insrc/kimi_cli/llm.pyserves as the central authority for budget mathematics.
Frequently Asked Questions
What is the difference between KIMI_MODEL_MAX_COMPLETION_TOKENS and KIMI_MODEL_MAX_TOKENS?
KIMI_MODEL_MAX_TOKENS functions as a compatibility alias for KIMI_MODEL_MAX_COMPLETION_TOKENS. Both variables are checked in src/kimi_cli/llm.py, but KIMI_MODEL_MAX_COMPLETION_TOKENS takes precedence if both are defined.
How does kimi-cli prevent exceeding the model's context window?
The CLI calls estimate_request_tokens to measure input size, then compute_max_completion_tokens calculates max_context_size - input_tokens as the response budget. This logic, implemented in src/kimi_cli/llm.py and applied in src/kimi_cli/soul/kimisoul.py, ensures the total prompt plus completion never exceeds the model's limit.
Can I disable the token limit entirely?
Yes. Set KIMI_MODEL_MAX_COMPLETION_TOKENS=0 (or any negative integer) in your environment, or pass {"max_completion_tokens": 0} in the generation_overrides mapping. This disables client-side clamping and delegates the decision to the API provider.
How do I set token limits for specific agent skills?
Agent configurations in src/kimi_cli/agents/spec.yaml can specify generation_overrides including max_completion_tokens. This allows individual skills to use higher or lower limits than the global default without affecting other CLI functionality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →