Kimi CLI OAuth Authentication Flow: Device-Code Login, Keyring Migration, and Token Refresh
Kimi CLI authenticates to the Kimi Code platform through an OAuth 2.0 device-code grant, stores tokens in atomic local JSON files, and automatically migrates legacy keyring entries while refreshing access tokens in the background.
The Kimi CLI project (MoonshotAI/kimi-cli) implements its entire authentication lifecycle in src/kimi_cli/auth/oauth.py. Understanding the Kimi CLI OAuth authentication flow is essential for anyone extending the client or debugging login issues, because the module handles everything from device-code polling to cross-process token refresh locks.
How the Kimi CLI Device-Code OAuth Flow Works
The implementation follows the standard OAuth 2.0 device authorization grant. The OAuthManager orchestrates the process at runtime, splitting work between three stages: device-code acquisition, user verification, and token retrieval.
Device-Code Acquisition
The request_device_authorization coroutine initiates the flow by posting to the Kimi OAuth server:
async def request_device_authorization() -> DeviceAuthorization:
async with (
new_client_session() as session,
session.post(
f"{_oauth_host().rstrip('/')}/api/oauth/device_authorization",
data={"client_id": KIMI_CODE_CLIENT_ID},
headers=_common_headers(),
) as response,
):
data = await response.json(content_type=None)
status = response.status
if status != 200:
raise OAuthError(f"Device authorization failed: {data}")
return DeviceAuthorization(
user_code=str(data["user_code"]),
device_code=str(data["device_code"]),
verification_uri=str(data.get("verification_uri") or ""),
verification_uri_complete=str(data["verification_uri_complete"]),
expires_in=int(data.get("expires_in") or 0) or None,
interval=int(data.get("interval") or 5),
)
This function returns a DeviceAuthorization dataclass containing the user_code, device_code, and verification URLs required for the next step. In src/kimi_cli/auth/oauth.py, the CLI builds the POST request to api/oauth/device_authorization and validates the HTTP status before parsing the response.
User Verification and Token Polling
The high-level login_kimi_code coroutine drives user interaction. It first calls request_device_authorization, yields a verification URL event for the UI layer, and then enters a polling loop:
auth = await request_device_authorization()
# ... yield verification URL event ...
while True:
status, data = await _request_device_token(auth)
if status == 200 and "access_token" in data:
token = OAuthToken.from_response(data)
break
# handle errors, wait `auth.interval` seconds, then retry
If open_browser is enabled, the CLI opens the verification URI automatically via webbrowser.open. The _request_device_token helper repeatedly polls the /api/oauth/token endpoint using the device_code until the user authorizes the request or the code expires.
Token Persistence: File Storage and Keyring Migration
Once the server returns an access token, the CLI must persist it securely. The codebase in src/kimi_cli/auth/oauth.py now prefers file-based storage but retains keyring support for backward compatibility.
The OAuthToken Model
The token itself is defined as a slots-based dataclass:
@dataclass(slots=True)
class OAuthToken:
access_token: str
refresh_token: str
expires_at: float
scope: str
token_type: str
expires_in: float = 0.0
OAuthToken.from_response reconstructs the object from the server's JSON payload, and instances can be serialized to plain dictionaries for storage.
File-Based Token Storage
File helpers such as _credentials_dir, _credentials_path, and _credentials_lock_path resolve a per-key JSON file under the shared user directory (typically ~/.kimi/credentials). The _save_to_file function writes atomically using a temporary file followed by os.replace, preventing corruption if the process crashes. Conversely, _load_from_file reads the JSON and instantiates an OAuthToken.
Keyring Fallback and Automatic Migration
For legacy installations, _load_from_keyring reads a password entry from the OS keyring:
def _load_from_keyring(key: str) -> OAuthToken | None:
raw = keyring.get_password(KEYRING_SERVICE, key)
# ... JSON-decode and return OAuthToken ...
When OAuthRef.storage == "keyring", the load_tokens function first checks the file cache, then falls back to keyring. On a successful keyring load, the token is immediately migrated to a file via _save_to_file and the keyring entry is deleted. This ensures old credentials move to the modern file-based backend without user intervention.
Public API for Saving Tokens
The save_tokens function enforces the migration at write time:
def save_tokens(ref: OAuthRef, token: OAuthToken) -> OAuthRef:
if ref.storage == "keyring":
logger.warning("Keyring storage is deprecated; saving OAuth tokens to file.")
ref = OAuthRef(storage="file", key=ref.key)
_save_to_file(ref.key, token)
return ref
Any attempt to save to keyring is intercepted and redirected to the file backend, returning an updated OAuthRef so that subsequent reads use the correct path.
Runtime Token Management with OAuthManager
OAuthManager is instantiated once per CLI process or web server and centralizes all token operations. According to the source code in src/kimi_cli/auth/oauth.py, it handles loading credentials, migrating legacy storage, caching access tokens, and refreshing them proactively.
Loading and Migrating Legacy Credentials
On initialization, the manager iterates over OAuth references in the global Config using _iter_oauth_refs. It then runs _migrate_oauth_storage to convert any remaining keyring entries to files before the rest of the application needs them.
Background Token Refresh
The refreshing method returns an async context manager that spawns a background task. This task sleeps for REFRESH_INTERVAL_SECONDS (60 seconds) and calls ensure_fresh each cycle. If the process wakes after a long pause (elapsed time greater than twice the interval), the manager forces an immediate refresh. Successful refreshes update the in-memory cache and the live LLM client via _apply_access_token.
Cross-Process Lock for Refresh Safety
Token refresh is guarded by _CrossProcessLock. The implementation uses fcntl.flock on Unix and msvcrt.locking on Windows, ensuring that only a single process performs a refresh even when multiple CLI instances or workers run concurrently.
The Refresh Algorithm
Inside _refresh_tokens, the manager performs the following steps:
- Load the persisted token using
load_tokens. - Check if the token is close to expiry against
_refresh_threshold. - Call
refresh_tokento exchange the refresh token for a new pair. - On HTTP 401 or 403 (
OAuthUnauthorized), mark the refresh token as rejected via_mark_refresh_token_rejected. - Write the new token back with
save_tokensand update the in-memory cache.
This algorithm guarantees that stale tokens are detected and renewed before they expire.
Practical Kimi CLI OAuth Usage Examples
Logging in via the CLI
The simplest entry point is the built-in login command:
kimi login
This command is exposed in src/kimi_cli/cli/info.py and ultimately invokes login_kimi_code. After the user visits the verification URL and authorizes the device, the new token is saved to ~/.kimi/credentials/kimi-code.json.
Resolving API Keys in Code
Applications integrating Kimi CLI can query the current access token programmatically:
from kimi_cli.auth.oauth import OAuthManager, OAuthRef
from kimi_cli.config import load_config
cfg = load_config()
manager = OAuthManager(cfg)
# Resolve the API key for a provider that uses OAuth
api_key = manager.resolve_api_key(
cfg.providers["kimi"].api_key,
cfg.providers["kimi"].oauth,
)
print("Current access token:", api_key)
resolve_api_key returns the cached access token when available; otherwise, it falls back to a static API key. This makes the OAuth flow transparent to downstream API consumers.
Background Refresh in Long-Running Processes
For servers or long-running scripts, wrap the runtime in the refreshing context manager:
async with oauth_manager.refreshing(runtime):
# handle requests; tokens stay fresh automatically
await serve_requests()
This pattern keeps the access token valid indefinitely without blocking request handling.
Summary
- Kimi CLI OAuth authentication flow uses an OAuth 2.0 device-code grant implemented in
src/kimi_cli/auth/oauth.py. - The
request_device_authorizationfunction acquires a device code, whilelogin_kimi_codehandles user verification and token polling. - Tokens are persisted as atomic JSON files under
~/.kimi/credentials, and legacy keyring entries are automatically migrated on first read. OAuthManagerloads credentials, caches access tokens, and refreshes them in the background using a cross-process lock.- The refresh algorithm detects expiry, exchanges refresh tokens, handles 401/403 rejection, and writes new credentials back to disk.
Frequently Asked Questions
How does Kimi CLI handle OAuth token storage securely?
Kimi CLI writes tokens to a local JSON file using an atomic os.replace operation via _save_to_file in src/kimi_cli/auth/oauth.py. Earlier versions stored tokens in the OS keyring, but the current codebase treats keyring as deprecated and automatically migrates those entries to files on load.
What happens if the access token expires while the CLI is running?
The OAuthManager spawns a background refresh task when entered through the refreshing async context manager. It checks token age at regular intervals and invokes _refresh_tokens before the token expires. If the refresh fails with an unauthorized error, the token is marked rejected and the user must log in again.
Can multiple Kimi CLI processes refresh the same token simultaneously?
No. The _CrossProcessLock class in src/kimi_cli/auth/oauth.py uses fcntl.flock on Unix and msvcrt.locking on Windows to serialize refresh attempts. Only one process can perform the token refresh; others wait until the lock is released and then read the updated file.
Why does Kimi CLI use the OAuth device-code flow instead of a browser redirect?
The device-code flow is ideal for CLI applications because the client runs in a terminal without a local HTTP listener. The user completes authorization on a separate device or browser window by visiting the verification URI, while the CLI polls _request_device_token in the background until the server issues the access and refresh tokens.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →