How to Use Temporal Localization for Event Detection in Videos with NVIDIA Cosmos
Cosmos 3’s Reasoner surface enables temporal localization by encoding video frames with multimodal rotary position embeddings (mRoPE) and processing them through a Mixture-of-Transformers (MoT) backbone to generate timestamped JSON event segments.
Temporal localization for event detection in videos allows AI systems to identify when specific actions occur within a timeline. NVIDIA Cosmos provides this capability through its Reasoner surface, which processes MP4 inputs using a unified architecture that jointly attends to spatial and temporal dimensions. By leveraging multimodal rotary position embeddings (mRoPE) and a Mixture-of-Transformers (MoT) backbone, Cosmos converts sampled frames into time-aware tokens that support precise event timestamping.
How Temporal Localization Works in Cosmos 3
Cosmos 3 implements temporal localization through its Reasoner surface, which treats video understanding as a token prediction task. When you submit a video, the model extracts frames at a configurable sampling rate and tokenizes them using mRoPE to encode both spatial coordinates and temporal positions.
The MoT backbone—shared between the vision and language modalities—processes these tokens to maintain temporal coherence across the video sequence. This unified architecture allows the model to attend to specific time steps when generating responses about event timing, rather than treating the video as a static image collection.
Configuring Frame Sampling Rates
The temporal resolution of event detection depends on how Cosmos samples frames from the input video. By default, the system extracts frames at 4 fps, but you can adjust this via the extra_body.mm_processor_kwargs parameter.
To modify the sampling rate, include the following configuration in your API request:
extra_body = {
"mm_processor_kwargs": {
"fps": 4 # Adjust for higher or lower temporal resolution
}
}
Higher frame rates improve detection precision for brief events but increase token count and processing time, while lower rates reduce computational overhead for longer videos.
Constructing Temporal Localization Requests
To detect events with temporal localization, structure your payload to place the video media before the text query. This ordering ensures the model processes the visual context prior to generating timestamped responses.
Include a specific prompt that requests temporal segmentation, such as "List all action segments in the video" or "When does each event start and end?" The model processes the video through the MoT backbone and responds with a JSON array containing precise timestamps.
Example request structure:
import requests
payload = {
"model": "cosmos-3-reasoner",
"messages": [
{
"role": "user",
"content": [
{
"type": "video",
"video": "path/to/video.mp4"
},
{
"type": "text",
"text": "List all action segments in the video with start and end timestamps."
}
]
}
],
"extra_body": {
"mm_processor_kwargs": {
"fps": 4
}
}
}
response = requests.post("https://api.nvidia.com/cosmos/v1/chat/completions", json=payload)
Parsing Temporal Localization Results
When Cosmos identifies events, it returns a structured JSON response where each detected event includes temporal boundaries. Each entry contains:
start: The beginning timestamp in secondsend: The ending timestamp in secondscaption: A brief description of the detected event
Example response format:
[
{
"start": 12.5,
"end": 18.2,
"caption": "Person opens the door"
},
{
"start": 24.0,
"end": 29.5,
"caption": "Person walks through hallway"
}
]
This format enables direct integration with video editing tools, annotation pipelines, or search indexing systems that require precise temporal coordinates.
Summary
- Cosmos 3 uses multimodal rotary position embeddings (mRoPE) to encode temporal information into video tokens alongside spatial data.
- Configure frame sampling rates via
extra_body.mm_processor_kwargswith thefpsparameter (default 4). - Place video media before text queries in the payload to ensure proper temporal context processing.
- Temporal localization prompts return JSON arrays with
start,end, andcaptionfields for each detected event. - The Mixture-of-Transformers (MoT) backbone processes video and text through a unified architecture, enabling cross-modal temporal reasoning across the input sequence.
Frequently Asked Questions
What frame rate does Cosmos use for temporal localization?
By default, Cosmos samples video at 4 frames per second (fps) when processing for temporal localization. You can override this by setting extra_body.mm_processor_kwargs["fps"] to a higher or lower value depending on your precision requirements and computational budget.
How does Cosmos maintain temporal coherence across video frames?
Cosmos implements multimodal rotary position embeddings (mRoPE) that encode both spatial and temporal dimensions into each token. These tokens are processed by the Mixture-of-Transformers (MoT) backbone, which allows the model to attend across time steps and maintain temporal relationships throughout the video sequence.
What format does Cosmos return for detected event timestamps?
Cosmos returns a JSON list where each event object contains three fields: start (seconds), end (seconds), and caption (string description). This structure provides precise temporal boundaries that integrate directly with video analysis pipelines and timestamp-based search systems.
Can I adjust the temporal resolution for finer-grained event detection?
Yes. Increase the fps value in mm_processor_kwargs to sample more frames per second, which improves detection of brief or rapid events. Note that higher sampling rates increase the number of tokens processed by the MoT backbone, resulting in higher computational costs and latency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →