# How to Use Temporal Localization for Event Detection in Videos with NVIDIA Cosmos

> Learn temporal localization for event detection in videos using NVIDIA Cosmos. Explore multimodal rotary position embeddings and MoT backbone for timestamped JSON event segments.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Cosmos 3’s Reasoner surface enables temporal localization by encoding video frames with multimodal rotary position embeddings (mRoPE) and processing them through a Mixture-of-Transformers (MoT) backbone to generate timestamped JSON event segments.**

Temporal localization for event detection in videos allows AI systems to identify when specific actions occur within a timeline. NVIDIA Cosmos provides this capability through its Reasoner surface, which processes MP4 inputs using a unified architecture that jointly attends to spatial and temporal dimensions. By leveraging **multimodal rotary position embeddings (mRoPE)** and a **Mixture-of-Transformers (MoT)** backbone, Cosmos converts sampled frames into time-aware tokens that support precise event timestamping.

## How Temporal Localization Works in Cosmos 3

Cosmos 3 implements temporal localization through its Reasoner surface, which treats video understanding as a token prediction task. When you submit a video, the model extracts frames at a configurable sampling rate and tokenizes them using mRoPE to encode both spatial coordinates and temporal positions.

The MoT backbone—shared between the vision and language modalities—processes these tokens to maintain temporal coherence across the video sequence. This unified architecture allows the model to attend to specific time steps when generating responses about event timing, rather than treating the video as a static image collection.

## Configuring Frame Sampling Rates

The temporal resolution of event detection depends on how Cosmos samples frames from the input video. By default, the system extracts frames at **4 fps**, but you can adjust this via the `extra_body.mm_processor_kwargs` parameter.

To modify the sampling rate, include the following configuration in your API request:

```python
extra_body = {
    "mm_processor_kwargs": {
        "fps": 4  # Adjust for higher or lower temporal resolution

    }
}

```

Higher frame rates improve detection precision for brief events but increase token count and processing time, while lower rates reduce computational overhead for longer videos.

## Constructing Temporal Localization Requests

To detect events with temporal localization, structure your payload to place the video media before the text query. This ordering ensures the model processes the visual context prior to generating timestamped responses.

Include a specific prompt that requests temporal segmentation, such as "List all action segments in the video" or "When does each event start and end?" The model processes the video through the MoT backbone and responds with a JSON array containing precise timestamps.

Example request structure:

```python
import requests

payload = {
    "model": "cosmos-3-reasoner",
    "messages": [
        {
            "role": "user",
            "content": [
                {
                    "type": "video",
                    "video": "path/to/video.mp4"
                },
                {
                    "type": "text",
                    "text": "List all action segments in the video with start and end timestamps."
                }
            ]
        }
    ],
    "extra_body": {
        "mm_processor_kwargs": {
            "fps": 4
        }
    }
}

response = requests.post("https://api.nvidia.com/cosmos/v1/chat/completions", json=payload)

```

## Parsing Temporal Localization Results

When Cosmos identifies events, it returns a structured JSON response where each detected event includes temporal boundaries. Each entry contains:

- `start`: The beginning timestamp in seconds
- `end`: The ending timestamp in seconds
- `caption`: A brief description of the detected event

Example response format:

```json
[
  {
    "start": 12.5,
    "end": 18.2,
    "caption": "Person opens the door"
  },
  {
    "start": 24.0,
    "end": 29.5,
    "caption": "Person walks through hallway"
  }
]

```

This format enables direct integration with video editing tools, annotation pipelines, or search indexing systems that require precise temporal coordinates.

## Summary

- Cosmos 3 uses **multimodal rotary position embeddings (mRoPE)** to encode temporal information into video tokens alongside spatial data.
- Configure frame sampling rates via `extra_body.mm_processor_kwargs` with the `fps` parameter (default 4).
- Place video media before text queries in the payload to ensure proper temporal context processing.
- Temporal localization prompts return JSON arrays with `start`, `end`, and `caption` fields for each detected event.
- The **Mixture-of-Transformers (MoT)** backbone processes video and text through a unified architecture, enabling cross-modal temporal reasoning across the input sequence.

## Frequently Asked Questions

### What frame rate does Cosmos use for temporal localization?

By default, Cosmos samples video at **4 frames per second** (fps) when processing for temporal localization. You can override this by setting `extra_body.mm_processor_kwargs["fps"]` to a higher or lower value depending on your precision requirements and computational budget.

### How does Cosmos maintain temporal coherence across video frames?

Cosmos implements **multimodal rotary position embeddings (mRoPE)** that encode both spatial and temporal dimensions into each token. These tokens are processed by the **Mixture-of-Transformers (MoT)** backbone, which allows the model to attend across time steps and maintain temporal relationships throughout the video sequence.

### What format does Cosmos return for detected event timestamps?

Cosmos returns a JSON list where each event object contains three fields: `start` (seconds), `end` (seconds), and `caption` (string description). This structure provides precise temporal boundaries that integrate directly with video analysis pipelines and timestamp-based search systems.

### Can I adjust the temporal resolution for finer-grained event detection?

Yes. Increase the `fps` value in `mm_processor_kwargs` to sample more frames per second, which improves detection of brief or rapid events. Note that higher sampling rates increase the number of tokens processed by the MoT backbone, resulting in higher computational costs and latency.