# Video Input Processing and Media Handling in Kimi Code: Architecture and Implementation

> Explore Kimi Code's advanced video input processing and media handling. Discover its layered architecture for validation, storage, and provider capability negotiation.

- Repository: [Moonshot AI/kimi-code](https://github.com/MoonshotAI/kimi-code)
- Tags: architecture
- Published: 2026-07-26

---

**Kimi Code implements video input processing through a layered architecture that validates media via JSON schemas, stores content as session-global transcript attachments, and negotiates provider capabilities to ensure compatible model routing.**

MoonshotAI/kimi-code treats video as a first-class input modality across its entire stack, from protocol definitions to terminal UI rendering. The system validates video URLs, base64 data, and file references through strict schema enforcement while maintaining persistent attachment records that persist across conversation turns.

## Architecture Overview

The video handling system spans three distinct layers, each with specific responsibilities for managing media content:

- **Protocol Layer** – Defines JSON schema validation for video messages sent over the API. The `videoContentSchema` in [`packages/protocol/src/message.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/protocol/src/message.ts) validates objects containing `type: "video"` and `source` properties (URL, base64, or file references).

- **Transcript Model** – Stores video as session-global attachments via the `TranscriptAttachment` interface in [`packages/transcript/src/model/attachment.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/transcript/src/model/attachment.ts), enabling references across multiple conversation turns and streaming to UI components.

- **History Processing** – The `groupTurns` logic in [`packages/transcript/src/history/groupTurns.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/transcript/src/history/groupTurns.ts) classifies message parts as `video` types and includes them in turn payloads as `HistoryMediaSource` objects.

## End-to-End Video Processing Flow

The journey from API request to model consumption follows six distinct phases:

### 1. Protocol Validation

API requests containing video undergo strict validation against `videoContentSchema`. The schema accepts three source kinds: `url`, `base64` (with `media_type` and `data` fields), and `file` (referencing uploaded file IDs).

```typescript
// Validated video content structure
{
  "type": "video",
  "source": {
    "kind": "url",
    "url": "https://example.com/clip.mp4"
  }
}

```

### 2. SDK Normalization

The Node SDK ([`packages/node-sdk/src/types.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/node-sdk/src/types.ts)) maps raw video parts to `PromptPart` union types with `type: 'video_url'`. The SDK throws a validation error with the message **"Prompt input cannot contain empty video URLs"** when encountering invalid inputs.

### 3. Turn Construction

The `groupTurns` utility processes incoming messages and creates turn payloads where video parts are stored as `HistoryMediaSource` objects. This payload structure persists in the transcript journal for session continuity.

### 4. Attachment Storage

The transcript service creates `TranscriptAttachment` records for each video input. These records store the source metadata, enabling later retrieval for UI rendering, conversation replay, or cross-turn referencing.

### 5. Capability Negotiation

Provider compatibility is determined in [`packages/oauth/src/open-platform.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/oauth/src/open-platform.ts) and [`managed-kimi-code.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/managed-kimi-code.ts). The system checks for `supports_video_in` in provider metadata and adds `video_in` to the model's capability set, ensuring only compatible models receive video inputs.

### 6. Rendering Pipeline

In the terminal UI ([`packages/pi-tui/src/components/input.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/pi-tui/src/components/input.ts)), the cursor draws using reverse-video escape codes (`\x1b[7m…\x1b[27m`). The web interface consumes transcript attachment URLs to generate video preview components.

## Implementation Examples

### Creating Video Prompt Parts

Use the Node SDK to construct video inputs with strict type safety:

```typescript
import { PromptPart } from '@moonshot-ai/kimi-code-sdk';

const videoPart: PromptPart = {
  type: 'video_url',
  videoUrl: { url: 'https://example.com/clip.mp4' },
};

// Mixed content prompt
await client.sessions.prompt(sessionId, {
  content: [
    { type: 'text', text: 'Analyze this video' },
    videoPart,
  ],
});

```

### Retrieving Stored Attachments

Access persistent video metadata through the transcript service:

```typescript
import { TranscriptService } from '@moonshot-ai/kimi-code-sdk';

const transcript = await TranscriptService.get(sessionId);
const videoAttachment = transcript.attachments.find(
  a => a.type === 'video'
);

console.log('Video URL:', videoAttachment?.source?.url);

```

### Base64 and File Source Handling

The schema supports inline base64 encoding and file references:

```typescript
// Base64 encoded video
const base64Part: PromptPart = {
  type: 'video_url',
  videoUrl: {
    url: 'data:video/mp4;base64,<encoded-data>',
    media_type: 'video/mp4'
  }
};

// Referenced file upload
const filePart = {
  type: 'video',
  source: { kind: 'file', file_id: 'uploaded-file-uuid' }
};

```

## Edge Cases and Fallback Handling

**Unsupported Provider Fallback** – When a provider does not advertise `supports_video_in`, the system replaces video content with placeholder text: `"(video omitted: not supported by this provider)"` (as implemented in [`packages/kosong/test/openai-responses.test.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/kosong/test/openai-responses.test.ts)).

**Empty URL Prevention** – The SDK validates URLs before transmission, preventing empty strings from entering the transcript history and triggering provider errors.

**Source Type Flexibility** – Beyond HTTP URLs, the architecture handles `kind: "base64"` for inline data (requiring `media_type` and `data` fields) and `kind: "file"` for referencing previously uploaded content via file IDs.

## Summary

- **Validation Layer** – [`packages/protocol/src/message.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/protocol/src/message.ts) defines `videoContentSchema` for strict API input validation across URL, base64, and file source types.
- **Persistence Model** – `TranscriptAttachment` in [`packages/transcript/src/model/attachment.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/transcript/src/model/attachment.ts) enables session-wide video reference and UI streaming.
- **Capability Detection** – Provider metadata in [`packages/oauth/src/open-platform.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/oauth/src/open-platform.ts) determines `video_in` support before routing requests.
- **Error Prevention** – The Node SDK blocks empty video URLs and replaces unsupported content with descriptive placeholders.
- **Multi-Interface Support** – Videos render in terminal TUI via escape codes and in web interfaces through attachment URL consumption.

## Frequently Asked Questions

### How does Kimi Code validate video input sources?

The system validates video inputs through `videoContentSchema` in [`packages/protocol/src/message.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/protocol/src/message.ts), which enforces structure for `type: "video"` objects. The Node SDK ([`packages/node-sdk/src/types.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/node-sdk/src/types.ts)) performs additional runtime validation, rejecting empty URLs with specific error messages and ensuring only properly formatted URLs, base64 data, or file references enter the transcript.

### What happens when a model provider doesn't support video inputs?

When capability negotiation in [`packages/oauth/src/open-platform.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/oauth/src/open-platform.ts) detects that a provider lacks `supports_video_in` metadata, the system replaces the video content with the placeholder text `"(video omitted: not supported by this provider)"` before sending the request. This prevents provider-side errors while maintaining conversation context.

### How are video attachments stored across conversation turns?

Videos persist as `TranscriptAttachment` objects defined in [`packages/transcript/src/model/attachment.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/transcript/src/model/attachment.ts). The `groupTurns` logic in [`packages/transcript/src/history/groupTurns.ts`](https://github.com/MoonshotAI/kimi-code/blob/main/packages/transcript/src/history/groupTurns.ts) stores video references as `HistoryMediaSource` entries within turn payloads, making them accessible throughout the session for context windows, UI rendering, and conversation replay.

### What video source formats does Kimi Code support beyond standard URLs?

The architecture supports three source types: `kind: "url"` for HTTP/HTTPS endpoints, `kind: "base64"` for inline data (requiring `media_type` and `data` fields), and `kind: "file"` for referencing previously uploaded content via file IDs. This flexibility enables direct data embedding and file reference patterns without requiring external hosting.