# How to Use NVIDIA Cosmos 3 Reasoner for Video Captioning and Temporal Localization

> Learn to use NVIDIA Cosmos 3 Reasoner for video captioning and temporal localization. Process video inputs and generate detailed captions or structured temporal events with this powerful transformer architecture.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**The Cosmos 3 Reasoner processes video inputs through a Mixture-of-Transformers architecture to generate detailed captions or structured temporal event localizations via an OpenAI-compatible API endpoint.**

The NVIDIA Cosmos repository provides a production-ready implementation of the Cosmos 3 Reasoner, a vision-language model designed specifically for multimodal understanding tasks. Unlike the Generator surface which handles multimodal generation, the Reasoner focuses on vision-language understanding by accepting text, images, and videos to produce text or JSON outputs. This guide demonstrates how to leverage the Reasoner surface for video captioning and temporal event localization using both vLLM and NIM deployment options.

## Understanding Cosmos 3 Reasoner Architecture

The Cosmos 3 Reasoner builds on a unified **Mixture-of-Transformers (MoT)** backbone that processes language tokens through causal self-attention while handling visual tokens via a dedicated vision encoder. According to the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), the architecture shares transformer layers, multimodal attention mechanisms, and 3‑D multi-dimensional rotary position embeddings (mRoPE) with the Generator variant, but operates as a causal autoregressive decoder that predicts tokens sequentially until reaching an end-of-sequence marker.

### Vision Encoder and Multimodal Processing

The **vision encoder** converts image or video frames into visual token sequences that integrate with language tokens within the same transformer stack. For video inputs, frame sampling is controlled through `extra_body.media_io_kwargs` parameters, allowing you to specify FPS values such as `"fps": 4` to balance temporal resolution against inference latency. This unified representation enables the model to understand spatial relationships across frames while maintaining temporal coherence.

### Causal Autoregressive Design

The Reasoner employs **causal self-attention** that restricts the model to attend only to previous tokens during generation. This design supports open-ended text generation including detailed video captions and step-by-step reasoning traces. The causal nature ensures that each output token depends solely on the input media and previously generated text, making the model suitable for streaming applications and complex reasoning tasks.

## Setting Up the Reasoner Inference Server

The Cosmos 3 Reasoner supports two deployment surfaces: vLLM for customizable environments and NIM containers for turnkey deployment.

### Deploying with vLLM

Launch the Reasoner server using vLLM to support both Nano (16 B) and Super (64 B) checkpoints. The vLLM deployment provides an OpenAI-compatible API endpoint without requiring CUDA-specific setup beyond standard GPU drivers. According to [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md), this approach offers maximum flexibility for development and benchmarking while maintaining full compatibility with the OpenAI Python client.

### Deploying the NIM Container

The pre-built NIM container provides a production-ready deployment path that eliminates manual CUDA configuration. As documented in the repository's inference benchmarks, the NIM container exposes the same REST API as vLLM but handles media encoding through base64 data URIs rather than file paths. This containerized approach is ideal for production environments requiring consistent performance across different infrastructure configurations.

## Implementing Video Captioning

To generate video captions, structure your OpenAI-compatible request with media elements preceding the text prompt, as required by the Reasoner's input protocol documented in [`cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md).

```python
from pathlib import Path
import openai

# Convert local video to file URI for vLLM access

video_path = Path("assets/video_caption.mp4").resolve()
video_uri = f"file://{video_path}"

client = openai.OpenAI(
    api_key="EMPTY",  # vLLM does not require API key authentication

    base_url="http://localhost:8000/v1"
)

response = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": video_uri}},
                {"type": "text", "text": "Describe the video in detail."}
            ]
        }
    ],
    max_tokens=1024,
    temperature=0.7,
    top_p=0.8,
    extra_body={"media_io_kwargs": {"video": {"fps": 4}}}  # Sample at 4 fps

)

print(response.choices[0].message.content)

```

This implementation demonstrates the required media-text ordering and uses `extra_body.media_io_kwargs` to sample video frames at 4 FPS, reducing computational load while preserving sufficient temporal information for accurate captioning.

## Temporal Event Localization Techniques

For temporal localization tasks, prompt the model to return structured JSON containing event timestamps and descriptions. The Reasoner can parse complex instructions requesting specific output formats, enabling direct integration with downstream analytics pipelines.

```python
from pathlib import Path
import openai, json

video_uri = f"file://{Path('assets/temporal_localization_1.mp4').resolve()}"
client = openai.OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

prompt = """List all action segments in the video.
Provide the result in JSON format with 'seconds' for time depiction for each event.
Use keys: "start", "end", "caption"."""

response = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": video_uri}},
                {"type": "text", "text": prompt}
            ]
        }
    ],
    max_tokens=2048,
    temperature=0.6,
    top_p=0.95,
    extra_body={"media_io_kwargs": {"video": {"fps": 4}}}
)

# Parse JSON response for structured event extraction

events = json.loads(response.choices[0].message.content)
print(json.dumps(events, indent=2))

```

This example follows the temporal localization patterns from the prompt guide, utilizing higher `top_p` (0.95) and lower `temperature` (0.6) settings appropriate for structured output generation. The model returns a JSON array of objects containing start times, end times, and captions for each detected event.

## Configuration and Optimization

### Sampling Parameters

The Reasoner uses distinct sampling defaults depending on the task complexity. For standard video captioning without explicit reasoning, apply `top_p=0.8`, `top_k=20`, and `temperature=0.7`. When requiring structured reasoning or complex temporal analysis, switch to `top_p=0.95` and `temperature=0.6` to maintain output coherence while allowing sufficient variability for accurate localization.

### Enabling Chain-of-Thought Reasoning

Insert a ` ` block within your user prompt to activate the model's reasoning mode. When enabled, the Reasoner emits a chain-of-thought trace before delivering the final answer, improving accuracy on complex temporal reasoning tasks at the cost of increased token generation. This toggle is particularly effective for ambiguous video segments requiring contextual interpretation.

## Summary

- **Cosmos 3 Reasoner** provides a vision-language understanding surface distinct from the Generator, built on a Mixture-of-Transformers architecture with causal autoregressive decoding.
- Input requests must place **media elements before text prompts** in the message array, with frame sampling controlled via `extra_body.media_io_kwargs`.
- **Video captioning** uses standard sampling parameters (`temperature=0.7`, `top_p=0.8`) and accepts file URIs through vLLM or base64 data through NIM containers.
- **Temporal localization** requires specific JSON formatting prompts and benefits from adjusted sampling (`temperature=0.6`, `top_p=0.95`) to ensure structured output reliability.
- Deployment options include **vLLM** for development flexibility and **NIM containers** for production deployment, both exposing OpenAI-compatible endpoints.

## Frequently Asked Questions

### What is the difference between Cosmos 3 Generator and Reasoner?

The Generator surface handles multimodal generation tasks including video and image synthesis, while the Reasoner surface focuses exclusively on vision-language understanding and text generation. The Reasoner uses the same MoT backbone and vision encoder but operates as a causal autoregressive decoder optimized for analyzing input media and producing descriptive or structured text outputs rather than generating new visual content.

### How do I control video frame sampling rate in Cosmos 3 Reasoner?

Specify the frame sampling rate through the `extra_body.media_io_kwargs` parameter in your API request, setting `"video": {"fps": 4}` to sample at 4 frames per second. Lower FPS values reduce computational overhead and latency while higher values preserve fine-grained temporal details for complex action recognition tasks.

### Can Cosmos 3 Reasoner output structured JSON for temporal localization?

Yes, the Reasoner can generate structured JSON outputs when explicitly instructed through prompt engineering. Include specific formatting instructions in your text prompt requesting JSON objects with keys such as `"start"`, `"end"`, and `"caption"`, then parse the response content using standard JSON parsers. The model reliably follows these formatting instructions when using appropriate sampling parameters for structured generation.

### What hardware requirements are needed for running Cosmos 3 Reasoner?

The Reasoner supports two model sizes: Nano (16 billion parameters) and Super (64 billion parameters). Deployment requires NVIDIA GPUs with sufficient VRAM to accommodate the selected checkpoint size and desired batch throughput. The vLLM deployment method provides flexible hardware utilization while the NIM container offers optimized configurations for specific GPU architectures as detailed in the repository's [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) documentation.