Which AI Models Are Compatible with Kimi‑CLI? A Complete Provider Guide

Kimi‑CLI supports any large‑language model (LLM) exposed through one of its built‑in provider types, including Kimi, OpenAI, Anthropic, Google Gemini, and Vertex AI models.

The MoonshotAI/kimi-cli repository is architected as a provider‑agnostic command‑line interface. Rather than hardcoding model support, the tool delegates to pluggable provider implementations that dynamically expose available models. This design means compatibility extends to any model string recognized by the underlying provider SDK.

Supported Provider Types

The CLI enumerates available backends through the ProviderType literal defined in src/kimi_cli/llm.py (lines 32‑40). Each provider type maps to a specific implementation in the kosong package and supports distinct model families.

Kimi Models

The kimi provider connects to Moonshot AI’s own model family. Compatible strings include kimi‑1.5‑base, kimi‑2‑lite, and subsequent releases. The implementation resides in kosong.chat_provider.kimi.Kimi.

OpenAI Models

Two provider identifiers cover OpenAI’s API surface:

  • openai_legacy – Compatible with GPT‑4, GPT‑4‑turbo, GPT‑3.5‑turbo, and legacy completion models.
  • openai_responses – Targets newer response‑format endpoints such as gpt‑4o and gpt‑4o‑mini.

Both wrap kosong.chat_provider.openai and accept any model name listed in OpenAI’s model catalog.

Anthropic Claude Models

The anthropic provider drives Claude‑2, Claude‑2.1, Claude‑3‑sonnet, Claude‑3‑haiku, and Claude‑3‑opus via kosong.chat_provider.anthropic. You must use the exact model snapshot string (e.g., claude‑3‑haiku‑20240307) when configuring the CLI.

Google Gemini Models

Identifier google_genai (aliased as gemini) supports Gemini‑1.0‑pro, Gemini‑1.5‑flash, Gemini‑1.5‑pro, and future releases. The provider class is kosong.chat_provider.google. Both names map to the same implementation, so provider = "gemini" and provider = "google_genai" are equivalent in ~/.kimi/config.toml.

Vertex AI Models

The vertexai provider accepts any model deployed on Google Cloud Vertex AI, including text‑bison‑32k and code‑bison‑32k. Implementation lives in kosong.chat_provider.vertexai.

Test and Debug Providers

Internal providers _echo, _scripted_echo, and _chaos exist for debugging and chaos testing. These are not intended for production use but are selectable via the same configuration mechanism.

How Model Compatibility Works

Compatibility is determined at runtime by two Pydantic schemas defined in src/kimi_cli/config.py:

  1. LLMProvider (lines 60‑68) – Stores the type (provider identifier), base_url, and api_key.
  2. LLMModel (lines 69‑71) – Stores the model string, provider reference, max_context_size, and capabilities.

When the CLI initializes, it constructs an LLM wrapper (defined in src/kimi_cli/llm.py) that bridges the configuration to the provider SDK. The LLMModel.capabilities field accepts a set of literals from the ModelCapability type ("image_in", "video_in", "thinking", "always_thinking") defined in src/kimi_cli/llm.py (lines 45‑46). These flags inform the CLI whether the chosen model supports multimodal inputs or extended reasoning modes.

Because the provider SDKs surface their own model catalogs, the set of compatible models is dynamic. You only need the correct provider identifier and the exact model string advertised by the provider’s API.

Configuring Compatible Models

You declare available models in ~/.kimi/config.toml or programmatically via the Python API.

TOML Configuration

[default]
model = "gpt-4o-mini"
provider = "openai_legacy"

[[models]]
provider = "openai_legacy"
model = "gpt-4o"
max_context_size = 128_000
capabilities = ["thinking", "always_thinking"]

[[models]]
provider = "anthropic"
model = "claude-3-haiku-20240307"
max_context_size = 100_000
capabilities = ["image_in", "thinking"]

[[models]]
provider = "gemini"
model = "gemini-1.5-flash"
max_context_size = 1_048_576

Programmatic Configuration

from kimi_cli.config import Config, LLMProvider, LLMModel

cfg = Config(
    default_model="gemini-1.5-flash",
    default_thinking=True,
    providers=[
        LLMProvider(
            type="gemini",
            base_url="https://generativelanguage.googleapis.com/v1beta",
            api_key="YOUR_GEMINI_API_KEY",
        )
    ],
    models=[
        LLMModel(
            provider="gemini",
            model="gemini-1.5-flash",
            max_context_size=1_048_576,
            capabilities={"image_in", "thinking"},
        )
    ],
)

The app.py entry point consumes this configuration to instantiate the LLM class and inject it into the agent loop defined in agentspec.py.

Summary

  • Kimi‑CLI supports six production provider types (kimi, openai_legacy, openai_responses, anthropic, google_genai/gemini, vertexai) plus internal test providers.
  • Model compatibility is dynamic; any model string recognized by the provider SDK is valid.
  • Configuration occurs through LLMModel objects in src/kimi_cli/config.py, linking a provider type to a specific model name and capability set.
  • Capabilities (image_in, video_in, thinking, always_thinking) are declared per‑model to enable multimodal and reasoning features.

Frequently Asked Questions

How do I add a new custom model to Kimi‑CLI?

Add a new [[models]] entry in ~/.kimi/config.toml specifying the provider type and exact model string. Ensure the provider credentials are defined in the providers list. The CLI will validate the configuration against the LLMModel schema in src/kimi_cli/config.py on startup.

Can I use local or self‑hosted models with Kimi‑CLI?

Yes, provided they expose an OpenAI‑compatible API. Use the openai_legacy provider type and set the base_url in LLMProvider to your local endpoint (e.g., http://localhost:8000/v1). The CLI treats any OpenAI‑compatible server as a valid backend.

What is the difference between openai_legacy and openai_responses?

openai_legacy uses the traditional chat completions endpoint, while openai_responses targets OpenAI’s newer responses API format. According to the source in src/kimi_cli/llm.py, both are implemented in kosong.chat_provider.openai but handle request/response serialization differently.

How does Kimi‑CLI handle model capabilities like image input?

Capabilities are declared explicitly in the LLMModel.capabilities set using literals defined in src/kimi_cli/llm.py ("image_in", "video_in", "thinking", "always_thinking"). The CLI checks these flags before sending multimodal content or enabling extended reasoning modes, ensuring the selected model supports the requested operation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →