How Video Input Processing Works for Screen Recordings in Kimi-Code

Kimi-Code processes screen recordings through a multi-stage pipeline that detects video files in the UI, uploads them to generate stable file_id references, stores them as TranscriptAttachment entities, encodes them as protocol-compliant VideoContent objects, and routes them to AI models only when the provider advertises video_in capability.

MoonshotAI's kimi-code treats screen recordings as first-class multimodal inputs alongside text and images. Understanding the video input processing architecture reveals how the system bridges browser-based file uploads with backend storage and AI provider APIs.

Frontend File Detection and Upload Handling

When users drag screen recordings into the chat interface, the system identifies video MIME types and initiates a secure upload sequence.

Drag-and-Drop Detection in the UI

In apps/kimi-code/src/components/AttachmentDropzone.tsx, the frontend detects dropped files whose MIME types start with video/. The component creates a standard File object and initiates a multipart/form-data POST request to the backend endpoint /api/v1/attachments.

Backend Storage and ID Generation

The server-side handler defined in packages/kap-server/src/routes/attachments.ts receives the video stream, persists it to the configured file store, and returns a JSON response containing a stable file_id. This identifier serves as the canonical reference for the video throughout the transcript lifecycle.

// packages/kap-server/src/routes/attachments.ts
router.post('/attachments', async (req, res) => {
  const { file } = req.files as { file: UploadedFile };
  const fileId = await storeFile(file);
  res.json({ file_id: fileId });
});

Transcript Attachment Model

Once uploaded, videos are represented internally using the TranscriptAttachment type defined in packages/transcript/src/model/attachment.ts.

This model stores critical metadata including the attachment type (video), MIME type, file size, and the file_id linking to the stored blob. By unifying videos with other media types under a single attachment interface, the transcript system maintains consistency across different input modalities.

Protocol Message Schema

Before transmission to AI models, video attachments convert to protocol-compliant message parts. In packages/protocol/src/message.ts, the videoContentSchema defines the structure for video content blocks:

{
  type: "video",
  source: { 
    kind: "file", 
    file_id: "<file_id>" 
  }
}

The schema supports multiple source kinds including url, base64, and file. Screen recordings typically use the file kind, referencing the backend-stored ID rather than embedding base64 data directly.

Capability Validation and Provider Routing

Not all AI providers accept video input. Kimi-Code validates capabilities before forwarding content to prevent API errors.

The session management logic in packages/node-sdk/src/session.ts inspects the model's advertised capabilities for video_in. If the provider supports video, the VideoContent object passes through unchanged. If unsupported, the system substitutes a placeholder text such as "(video omitted: not supported by this provider)", ensuring graceful degradation without breaking conversation continuity.

History Aggregation and Turn Management

When reconstructing conversation history for context windows, packages/transcript/src/history/groupTurns.ts merges video parts into their corresponding transcript turns. This aggregation ensures that screen recordings remain logically attached to the user messages that uploaded them, preserving temporal context across multi-turn interactions.

Implementation Examples

Uploading a Screen Recording from the Frontend

async function uploadScreenRecording(file: File): Promise<string> {
  const form = new FormData();
  form.append('file', file);

  const response = await fetch('/api/v1/attachments', {
    method: 'POST',
    body: form,
  });
  const { file_id } = await response.json();
  return file_id; // e.g., "file_video_01"
}

Sending Video Input via the Node SDK

import { Kimi } from '@moonshot-ai/kimi-code-sdk';

const kimi = new Kimi({ apiKey: 'YOUR_API_KEY' });

await kimi.prompt({
  content: [
    { 
      type: 'text', 
      text: 'Analyze this screen recording and identify the UI interactions.' 
    },
    {
      type: 'video',
      source: { 
        kind: 'file', 
        file_id: 'file_video_01' 
      }
    }
  ]
});

Summary

  • MIME Type Detection: The frontend identifies video files via AttachmentDropzone.tsx when users drag screen recordings into the interface, filtering by video/* MIME types.
  • Stable References: The backend generates persistent file_id values through the attachments API in packages/kap-server/src/routes/attachments.ts, decoupling storage from conversation state.
  • Type Safety: TranscriptAttachment in attachment.ts provides the internal data model, while videoContentSchema in message.ts defines the external protocol contract for AI providers.
  • Capability Gates: The Node SDK in session.ts validates video_in support before transmitting video content, preventing API errors with incompatible models.
  • History Preservation: The groupTurns.ts logic ensures video attachments remain attached to their originating user messages when building context windows.

Frequently Asked Questions

What video file formats does kimi-code support for screen recordings?

Kimi-code accepts any format with a MIME type starting with video/, including MP4, MOV, and WebM files commonly produced by screen recording software. The detection logic in AttachmentDropzone.tsx relies on browser-reported MIME types during the drag-and-drop event.

How does kimi-code handle screen recordings when the AI model doesn't support video input?

When packages/node-sdk/src/session.ts detects that the selected provider lacks the video_in capability, it automatically replaces the VideoContent object with a descriptive text placeholder. This ensures the conversation proceeds without API errors while transparently indicating that video analysis is unavailable.

Can video attachments reference external URLs instead of uploaded files?

Yes. The videoContentSchema in packages/protocol/src/message.ts supports a source object with kind: "url", allowing screen recordings hosted on external CDNs or storage services to be referenced directly without uploading through the /api/v1/attachments endpoint.

How are screen recordings rendered in the chat interface?

After upload, the UI fetches video metadata using the file_id and renders the attachment as a playable video element. The transcript model in attachment.ts maintains the reference, while the frontend component retrieves the actual binary stream from the backend storage API for playback.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →