# How to Control Max Tokens in Kimi CLI: Environment Variables, Context-Aware Logic, and Per-Request Overrides

> Control max tokens in Kimi CLI with environment variables, context-aware logic, and per-request overrides. Learn how to manage your token budget effectively for optimal performance.

- Repository: [Moonshot AI/kimi-cli](https://github.com/MoonshotAI/kimi-cli)
- Tags: how-to-guide
- Published: 2026-07-27

---

**Kimi CLI determines the maximum completion token budget through a three-tier priority system: environment variable hard caps, automatic context-window calculations, and per-request `generation_overrides`.**

The MoonshotAI/kimi-cli repository implements a sophisticated token management system that prevents context window overflows while giving developers granular control over generation limits. Learning how to control max tokens in kimi-cli is essential for optimizing both short conversational turns and long-form content generation without triggering provider-side errors.

## Understanding the Token Budget Hierarchy

Kimi CLI applies token limits in three distinct stages, each overriding the previous if configured.

### Model-Level Hard Cap via Environment Variables

The highest priority control uses the `KIMI_MODEL_MAX_COMPLETION_TOKENS` environment variable (with a compatibility alias `KIMI_MODEL_MAX_TOKENS`). When set, this value becomes an explicit upper bound for every request.

In [`src/kimi_cli/llm.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/llm.py) (lines 68–84), the CLI checks for these variables during initialization. Setting the value to `0` or any negative number disables this clamping entirely, allowing the provider to determine the budget. This global setting ensures consistent behavior across all CLI sessions without modifying code.

### Context-Aware Automatic Fallback

When no environment variable is set, Kimi CLI automatically calculates a safe token budget to prevent context window violations. The system estimates consumed tokens via `estimate_request_tokens`, then computes the remaining capacity.

According to the source code in [`src/kimi_cli/llm.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/llm.py) (lines 81–96), the `compute_max_completion_tokens` function subtracts the input token count from the model’s `max_context_size`. The resulting value is applied to the request in [`src/kimi_cli/soul/kimisoul.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/kimisoul.py) (lines 1368–1387). This default behavior ensures the model never receives a budget larger than the available context window.

### Per-Request Generation Overrides

For temporary adjustments, callers can pass a `generation_overrides` mapping containing `"max_completion_tokens"` to modify the budget for a single request. The helper function `with_kimi_generation_overrides` injects this configuration into the provider instance.

As implemented in [`src/kimi_cli/llm.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/llm.py) (lines 64–71), this approach allows specific agent skills or interactive commands to use different limits without affecting the global configuration.

## Practical Implementation Examples

### Configure a Global Token Limit

Set a hard cap for all CLI sessions by exporting the environment variable before running commands:

```bash

# Limit all responses to 4096 tokens

export KIMI_MODEL_MAX_COMPLETION_TOKENS=4096

# Or disable clamping entirely (allow provider default)

export KIMI_MODEL_MAX_COMPLETION_TOKENS=0

```

### Override Tokens for a Single Request

Use the Python API to temporarily increase the budget for a specific operation:

```python
from kimi_cli.llm import with_kimi_generation_overrides

# Apply 8000-token limit to this provider instance only

overrides = {"max_completion_tokens": 8000}
chat = with_kimi_generation_overrides(chat, overrides)

# Subsequent requests use the 8K budget without affecting global settings

```

### Inspect the Computed Budget

Debug the automatic calculation to understand how much context remains:

```python
from kimi_cli.llm import compute_max_completion_tokens

max_ctx = 8192          # Model's max_context_size

input_toks = 2000       # Tokens consumed by system prompt and history

budget = compute_max_completion_tokens(
    max_context_size=max_ctx,
    input_tokens=input_toks,
    response_budget=None,  # None triggers automatic calculation

)

print(f"Computed max_completion_tokens = {budget}")

# Output: 6192 (8192 - 2000)

```

## Key Source Files and Functions

- **[`src/kimi_cli/llm.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/llm.py)** – Contains `compute_max_completion_tokens`, `with_kimi_generation_overrides`, and environment variable parsing logic for token budget management.
- **[`src/kimi_cli/soul/kimisoul.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/kimisoul.py)** – Applies the computed `max_completion_tokens` to each generation request at lines 1368–1387.
- **[`docs/en/configuration/env-vars.md`](https://github.com/MoonshotAI/kimi-cli/blob/main/docs/en/configuration/env-vars.md)** – Documents the `KIMI_MODEL_MAX_COMPLETION_TOKENS` and `KIMI_MODEL_MAX_TOKENS` variables.
- **[`src/kimi_cli/agents/spec.yaml`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/agents/spec.yaml)** – Demonstrates how agent specifications pass `generation_overrides` to influence token limits for specific skills.

## Summary

- **Environment variables** provide the highest priority hard cap via `KIMI_MODEL_MAX_COMPLETION_TOKENS`, where `0` disables clamping.
- **Automatic calculation** prevents context overflow by subtracting input tokens from `max_context_size` when no override is set.
- **Per-request overrides** via `generation_overrides` allow temporary budget adjustments without changing global configuration.
- The `compute_max_completion_tokens` function in [`src/kimi_cli/llm.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/llm.py) serves as the central authority for budget mathematics.

## Frequently Asked Questions

### What is the difference between KIMI_MODEL_MAX_COMPLETION_TOKENS and KIMI_MODEL_MAX_TOKENS?

`KIMI_MODEL_MAX_TOKENS` functions as a compatibility alias for `KIMI_MODEL_MAX_COMPLETION_TOKENS`. Both variables are checked in [`src/kimi_cli/llm.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/llm.py), but `KIMI_MODEL_MAX_COMPLETION_TOKENS` takes precedence if both are defined.

### How does kimi-cli prevent exceeding the model's context window?

The CLI calls `estimate_request_tokens` to measure input size, then `compute_max_completion_tokens` calculates `max_context_size - input_tokens` as the response budget. This logic, implemented in [`src/kimi_cli/llm.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/llm.py) and applied in [`src/kimi_cli/soul/kimisoul.py`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/soul/kimisoul.py), ensures the total prompt plus completion never exceeds the model's limit.

### Can I disable the token limit entirely?

Yes. Set `KIMI_MODEL_MAX_COMPLETION_TOKENS=0` (or any negative integer) in your environment, or pass `{"max_completion_tokens": 0}` in the `generation_overrides` mapping. This disables client-side clamping and delegates the decision to the API provider.

### How do I set token limits for specific agent skills?

Agent configurations in [`src/kimi_cli/agents/spec.yaml`](https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cli/agents/spec.yaml) can specify `generation_overrides` including `max_completion_tokens`. This allows individual skills to use higher or lower limits than the global default without affecting other CLI functionality.