Using Cosmos 3 Reasoner for Video Captioning and Temporal Event Localization

The Cosmos 3 Reasoner is a vision-language understanding surface that accepts text, images, or videos and returns structured text or JSON outputs, enabling detailed video captioning and precise temporal event localization via an OpenAI-compatible API.

The NVIDIA Cosmos repository provides a state-of-the-art omnimodal world model family with two distinct runtime surfaces. While the Generator handles multimodal content creation, the Reasoner specializes in vision-language understanding tasks. Built on a unified Mixture-of-Transformers (MoT) backbone with 3-D multi-dimensional rotary position embeddings (mRoPE), the Reasoner processes visual tokens through a dedicated vision encoder and language tokens through causal self-attention, making it ideal for analyzing video content and extracting temporal information.

Architecture and Key Components

The Reasoner shares transformer layers and multimodal attention mechanisms with the Generator but operates as a causal autoregressive decoder. According to the source code in README.md, the architecture processes language tokens with causal self-attention while handling visual tokens through a vision encoder, enabling the model to predict text outputs one token at a time until an end-of-sequence token is produced.

Vision Encoder and Token Processing

The vision encoder converts image or video frames into visual tokens. For video inputs, frame sampling is controlled via extra_body.media_io_kwargs, allowing you to specify parameters like "fps": 4 to balance detail against latency. This configuration is documented in cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md.

Sampling Parameters

The Reasoner uses conservative sampling defaults for standard tasks: top_p=0.8, top_k=20, and temperature=0.7. When requesting explicit reasoning chains, the repository recommends top_p=0.95 and temperature=0.6 to improve coherence.

Deployment Options

You can deploy the Cosmos 3 Reasoner using either vLLM for customizable setups or the pre-built NIM container for turnkey deployment.

vLLM Deployment

The vLLM launch supports both Nano (16B) and Super (64B) checkpoints. This method requires CUDA-specific setup but provides maximum flexibility for customizing inference parameters. The setup process is detailed in cookbooks/cosmos3/reasoner/README.md.

NIM Container

The NIM container offers a deployment option without CUDA-specific configuration on the host, providing a ready-to-use service that exposes an OpenAI-compatible endpoint on port 8000.

Video Captioning Implementation

To generate detailed video descriptions, you must structure the request with media appearing before the text prompt in the message list. The following example demonstrates captioning with frame sampling at 4 FPS to optimize processing speed.

from pathlib import Path
import openai

# Convert a local video file to a file-URI

video_path = Path("assets/video_caption.mp4").resolve()
video_uri = f"file://{video_path}"

client = openai.OpenAI(
    api_key="EMPTY",  # vLLM does not require an API key

    base_url="http://localhost:8000/v1"
)

response = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": video_uri}},
                {"type": "text", "text": "Describe the video in detail."}
            ]
        }
    ],
    max_tokens=1024,
    temperature=0.7,
    top_p=0.8,
    extra_body={"media_io_kwargs": {"video": {"fps": 4}}}
)

print(response.choices[0].message.content)

The extra_body.media_io_kwargs parameter controls video preprocessing, sampling at 4 FPS to reduce token count while preserving sufficient temporal information for accurate descriptions.

Temporal Event Localization

For extracting specific time intervals and events, prompt the model to return structured JSON. The cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md file provides examples showing how to request raw JSON outputs containing start times, end times, and captions for each detected event.

from pathlib import Path
import openai
import json

video_uri = f"file://{Path('assets/temporal_localization_1.mp4').resolve()}"

client = openai.OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

prompt = """List all action segments in the video.
Provide the result in JSON format with 'seconds' for time depiction for each event.
Use keys: "start", "end", "caption"."""

response = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": video_uri}},
                {"type": "text", "text": prompt}
            ]
        }
    ],
    max_tokens=2048,
    temperature=0.6,
    top_p=0.95,
    extra_body={"media_io_kwargs": {"video": {"fps": 4}}}
)

# Parse the JSON response

events = json.loads(response.choices[0].message.content)
print(json.dumps(events, indent=2))

This approach leverages the model's ability to generate structured outputs by explicitly requesting JSON formatting in the prompt, enabling direct integration with downstream analytics pipelines.

Enabling Chain-of-Thought Reasoning

To activate explicit reasoning before the final answer, include a <think> block in the user prompt. This instructs the model to emit a chain-of-thought trace prior to generating the structured output or caption. According to the prompt guide, use higher top_p (0.95) and lower temperature (0.6) when reasoning is enabled to maintain output stability.

NIM Container Usage

When using the NIM container, transmit media as base64 data URIs since the container cannot access host file paths directly.

VIDEO_B64=$(base64 -w 0 assets/video_caption.mp4)

curl -sS -X POST http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model":"nvidia/cosmos3-nano-reasoner",
    "messages":[
      {"role":"system","content":"You are a helpful assistant."},
      {"role":"user","content":[
        {"type":"video_url","video_url":{"url":"data:video/mp4;base64,'"${VIDEO_B64}"'"}},
        {"type":"text","text":"Describe the video in detail."}
      ]}
    ],
    "max_tokens":1024,
    "temperature":0.7,
    "top_p":0.8
  }' | jq -r '.choices[0].message.content'

The endpoint accepts the same OpenAI-compatible payload structure, maintaining consistency across deployment methods.

Summary

  • The Cosmos 3 Reasoner provides vision-language understanding through an OpenAI-compatible API, distinct from the Generator's multimodal generation capabilities.
  • Media ordering is critical: video or image URLs must precede text prompts in the message list.
  • Use extra_body.media_io_kwargs to control video frame sampling (e.g., 4 FPS) for latency optimization.
  • Deploy via vLLM for flexibility with Nano and Super checkpoints, or use the NIM container for simplified deployment without CUDA setup.
  • Request JSON outputs for temporal localization by explicitly formatting prompts to specify keys like "start", "end", and "caption".
  • Enable chain-of-thought reasoning by adding <think> blocks and adjusting sampling parameters to top_p=0.95 and temperature=0.6.

Frequently Asked Questions

How do I control the video frame sampling rate in Cosmos 3 Reasoner?

Pass a dictionary to extra_body.media_io_kwargs with a "video" key containing an "fps" parameter. For example, extra_body={"media_io_kwargs": {"video": {"fps": 4}}} samples the video at 4 frames per second, reducing token count and inference latency while preserving sufficient temporal detail for most captioning and localization tasks.

Can I use Cosmos 3 Reasoner without setting up CUDA on my host machine?

Yes. The NIM container provides a turnkey deployment option that does not require CUDA-specific setup on the host. Launch the container on a compatible GPU system, and it exposes an OpenAI-compatible endpoint on port 8000. Note that when using NIM, you must encode video files as base64 data URIs since the container cannot access host filesystem paths directly.

How do I enable chain-of-thought reasoning for complex video analysis?

Insert a <think> block into your user prompt. This instructs the model to generate a reasoning trace before producing the final answer. When using this feature, adjust sampling parameters to top_p=0.95 and temperature=0.6 as recommended in the reasoner_prompt_guide.md to ensure coherent reasoning chains while maintaining output diversity.

What is the difference between the Generator and Reasoner surfaces in Cosmos 3?

The Generator is designed for multimodal content creation, producing images or videos from text prompts. The Reasoner is optimized for vision-language understanding, accepting text, images, or videos as inputs and generating text or JSON outputs. The Reasoner uses the same MoT backbone and mRoPE as the Generator but operates as a causal autoregressive decoder for analyzing and describing visual content rather than synthesizing it.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →