# Using Cosmos 3 Reasoner for Video Captioning and Temporal Event Localization

> Leverage the Cosmos 3 Reasoner for advanced video captioning and temporal event localization. Get structured text or JSON outputs via an OpenAI-compatible API.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-13

---

**The Cosmos 3 Reasoner is a vision-language understanding surface that accepts text, images, or videos and returns structured text or JSON outputs, enabling detailed video captioning and precise temporal event localization via an OpenAI-compatible API.**

The NVIDIA Cosmos repository provides a state-of-the-art omnimodal world model family with two distinct runtime surfaces. While the Generator handles multimodal content creation, the **Reasoner** specializes in vision-language understanding tasks. Built on a unified Mixture-of-Transformers (MoT) backbone with 3-D multi-dimensional rotary position embeddings (mRoPE), the Reasoner processes visual tokens through a dedicated vision encoder and language tokens through causal self-attention, making it ideal for analyzing video content and extracting temporal information.

## Architecture and Key Components

The Reasoner shares transformer layers and multimodal attention mechanisms with the Generator but operates as a causal autoregressive decoder. According to the source code in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), the architecture processes language tokens with causal self-attention while handling visual tokens through a vision encoder, enabling the model to predict text outputs one token at a time until an end-of-sequence token is produced.

**Vision Encoder and Token Processing**

The vision encoder converts image or video frames into visual tokens. For video inputs, frame sampling is controlled via `extra_body.media_io_kwargs`, allowing you to specify parameters like `"fps": 4` to balance detail against latency. This configuration is documented in [`cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md).

**Sampling Parameters**

The Reasoner uses conservative sampling defaults for standard tasks: `top_p=0.8`, `top_k=20`, and `temperature=0.7`. When requesting explicit reasoning chains, the repository recommends `top_p=0.95` and `temperature=0.6` to improve coherence.

## Deployment Options

You can deploy the Cosmos 3 Reasoner using either vLLM for customizable setups or the pre-built NIM container for turnkey deployment.

**vLLM Deployment**

The vLLM launch supports both Nano (16B) and Super (64B) checkpoints. This method requires CUDA-specific setup but provides maximum flexibility for customizing inference parameters. The setup process is detailed in [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md).

**NIM Container**

The NIM container offers a deployment option without CUDA-specific configuration on the host, providing a ready-to-use service that exposes an OpenAI-compatible endpoint on port 8000.

## Video Captioning Implementation

To generate detailed video descriptions, you must structure the request with media appearing before the text prompt in the message list. The following example demonstrates captioning with frame sampling at 4 FPS to optimize processing speed.

```python
from pathlib import Path
import openai

# Convert a local video file to a file-URI

video_path = Path("assets/video_caption.mp4").resolve()
video_uri = f"file://{video_path}"

client = openai.OpenAI(
    api_key="EMPTY",  # vLLM does not require an API key

    base_url="http://localhost:8000/v1"
)

response = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": video_uri}},
                {"type": "text", "text": "Describe the video in detail."}
            ]
        }
    ],
    max_tokens=1024,
    temperature=0.7,
    top_p=0.8,
    extra_body={"media_io_kwargs": {"video": {"fps": 4}}}
)

print(response.choices[0].message.content)

```

The `extra_body.media_io_kwargs` parameter controls video preprocessing, sampling at 4 FPS to reduce token count while preserving sufficient temporal information for accurate descriptions.

## Temporal Event Localization

For extracting specific time intervals and events, prompt the model to return structured JSON. The [`cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md) file provides examples showing how to request raw JSON outputs containing start times, end times, and captions for each detected event.

```python
from pathlib import Path
import openai
import json

video_uri = f"file://{Path('assets/temporal_localization_1.mp4').resolve()}"

client = openai.OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

prompt = """List all action segments in the video.
Provide the result in JSON format with 'seconds' for time depiction for each event.
Use keys: "start", "end", "caption"."""

response = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": video_uri}},
                {"type": "text", "text": prompt}
            ]
        }
    ],
    max_tokens=2048,
    temperature=0.6,
    top_p=0.95,
    extra_body={"media_io_kwargs": {"video": {"fps": 4}}}
)

# Parse the JSON response

events = json.loads(response.choices[0].message.content)
print(json.dumps(events, indent=2))

```

This approach leverages the model's ability to generate structured outputs by explicitly requesting JSON formatting in the prompt, enabling direct integration with downstream analytics pipelines.

## Enabling Chain-of-Thought Reasoning

To activate explicit reasoning before the final answer, include a `<think>` block in the user prompt. This instructs the model to emit a chain-of-thought trace prior to generating the structured output or caption. According to the prompt guide, use higher `top_p` (0.95) and lower `temperature` (0.6) when reasoning is enabled to maintain output stability.

## NIM Container Usage

When using the NIM container, transmit media as base64 data URIs since the container cannot access host file paths directly.

```bash
VIDEO_B64=$(base64 -w 0 assets/video_caption.mp4)

curl -sS -X POST http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model":"nvidia/cosmos3-nano-reasoner",
    "messages":[
      {"role":"system","content":"You are a helpful assistant."},
      {"role":"user","content":[
        {"type":"video_url","video_url":{"url":"data:video/mp4;base64,'"${VIDEO_B64}"'"}},
        {"type":"text","text":"Describe the video in detail."}
      ]}
    ],
    "max_tokens":1024,
    "temperature":0.7,
    "top_p":0.8
  }' | jq -r '.choices[0].message.content'

```

The endpoint accepts the same OpenAI-compatible payload structure, maintaining consistency across deployment methods.

## Summary

- The **Cosmos 3 Reasoner** provides vision-language understanding through an OpenAI-compatible API, distinct from the Generator's multimodal generation capabilities.
- **Media ordering** is critical: video or image URLs must precede text prompts in the message list.
- Use **`extra_body.media_io_kwargs`** to control video frame sampling (e.g., 4 FPS) for latency optimization.
- Deploy via **vLLM** for flexibility with Nano and Super checkpoints, or use the **NIM container** for simplified deployment without CUDA setup.
- Request **JSON outputs** for temporal localization by explicitly formatting prompts to specify keys like `"start"`, `"end"`, and `"caption"`.
- Enable **chain-of-thought reasoning** by adding `<think>` blocks and adjusting sampling parameters to `top_p=0.95` and `temperature=0.6`.

## Frequently Asked Questions

### How do I control the video frame sampling rate in Cosmos 3 Reasoner?

Pass a dictionary to `extra_body.media_io_kwargs` with a `"video"` key containing an `"fps"` parameter. For example, `extra_body={"media_io_kwargs": {"video": {"fps": 4}}}` samples the video at 4 frames per second, reducing token count and inference latency while preserving sufficient temporal detail for most captioning and localization tasks.

### Can I use Cosmos 3 Reasoner without setting up CUDA on my host machine?

Yes. The **NIM container** provides a turnkey deployment option that does not require CUDA-specific setup on the host. Launch the container on a compatible GPU system, and it exposes an OpenAI-compatible endpoint on port 8000. Note that when using NIM, you must encode video files as base64 data URIs since the container cannot access host filesystem paths directly.

### How do I enable chain-of-thought reasoning for complex video analysis?

Insert a `<think>` block into your user prompt. This instructs the model to generate a reasoning trace before producing the final answer. When using this feature, adjust sampling parameters to `top_p=0.95` and `temperature=0.6` as recommended in the [`reasoner_prompt_guide.md`](https://github.com/NVIDIA/cosmos/blob/main/reasoner_prompt_guide.md) to ensure coherent reasoning chains while maintaining output diversity.

### What is the difference between the Generator and Reasoner surfaces in Cosmos 3?

The **Generator** is designed for multimodal content creation, producing images or videos from text prompts. The **Reasoner** is optimized for vision-language understanding, accepting text, images, or videos as inputs and generating text or JSON outputs. The Reasoner uses the same MoT backbone and mRoPE as the Generator but operates as a causal autoregressive decoder for analyzing and describing visual content rather than synthesizing it.