How Multi-Dimensional Rotary Position Embedding (mRoPE) Encodes Spatial-Temporal Structure Across Modalities in NVIDIA Cosmos
NVIDIA Cosmos uses a unified 3-D multi-dimensional rotary position embedding (mRoPE) that applies tensor-product rotations across height, width, and time axes to give every token—whether from image, video, audio, or action data—a consistent spatial-temporal coordinate in a shared representation space.
NVIDIA Cosmos employs a unified 3-D multi-dimensional rotary position embedding (mRoPE) to provide its transformer backbone with a consistent understanding of space and time across heterogeneous modalities. By encoding positional information through rotary transformations rather than additive sinusoidal vectors, the model maintains representational capacity while establishing relationships between tokens from images, video frames, audio waveforms, and action trajectories. This approach enables the architecture to reason about where and when events occur using a single positional encoding scheme shared between autoregressive reasoning and diffusion-based generation.
The Mechanics of mRoPE
Rotary Encoding Fundamentals
Unlike traditional position embeddings that add fixed sinusoidal vectors to token embeddings, rotary embeddings rotate the query and key vectors inside the self-attention module. This operation provides relative position bias while preserving the full representational capacity of the model's hidden sub-spaces.
Three-Dimensional Tensor Product
mRoPE extends the classic rotary scheme to three orthogonal axes: height (H), width (W), and time (T). For any token, the positional factor is computed as R_H ⊗ R_W ⊗ R_T, representing the tensor product of three rotation matrices. This computation simultaneously encodes where the token lies in the image plane and when it occurs in the video sequence.
Cross-Modal Unification
All modalities are projected into a common token space before entering the transformer. The same mRoPE matrices are applied to every token regardless of whether it originates from a still image, video frame, audio waveform segment, or action vector. This unification allows the model to reason jointly about spatial location, temporal occurrence, and semantic content without requiring separate positional encodings for each modality.
Spatial-Temporal Consistency
Because rotary rotations are linear and invertible, the model can efficiently relate tokens separated by large distances in space or time. This property supports long-range dependencies—such as tracking a robot's action across multiple frames—while maintaining linear memory complexity relative to sequence length. The embedding enables the model to understand relationships like "the robot lifted the cup after it passed the table" by directly attending to the encoded 3-D coordinates.
Implementation in the Cosmos Architecture
According to the NVIDIA Cosmos source code, mRoPE is injected before the attention layers of both the autoregressive (Reasoner) transformer and the diffusion (Generator) transformer. The core implementation resides in the Cosmos Framework repository at cosmos_framework/models/rotary_embedding.py, while architectural documentation appears in the main repository's README.md and visual diagrams in cosmos3-model-architecture.png.
Code Examples
Text-to-Video Generation
The following example demonstrates how the Cosmos 3 generator pipeline automatically applies mRoPE to encode spatial-temporal structure when generating video from text prompts, as shown in cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb.
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers import UniPCMultistepScheduler
from diffusers.utils import export_to_video
# Load the Nano checkpoint (the pipeline builds the mRoPE internally)
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
# Use the UniPC scheduler (still works with the same mRoPE)
pipe.scheduler = UniPCMultistepScheduler.from_config(
pipe.scheduler.config, flow_shift=10.0
)
# Prompt that requires both spatial and temporal understanding
prompt = (
"A small mobile robot drives down a hallway, stops in front of a shelf, "
"and then lifts a box while the camera pans left."
)
result = pipe(
prompt=prompt,
negative_prompt="blurred, low-quality",
image=None,
num_frames=120,
height=720,
width=1280,
fps=24,
num_inference_steps=35,
guidance_scale=6.0,
enable_sound=False,
)
# The returned video respects the spatial-temporal layout encoded by mRoPE
export_to_video(result.video, "robot_demo.mp4", fps=24)
Under the hood, each video frame is tokenized with frame indices feeding the time-axis rotation matrix while pixel coordinates feed the height and width rotation matrices.
Video-to-Text Reasoning
When using the Reasoner model via a vLLM endpoint, the same mRoPE embeddings allow the model to analyze spatial-temporal relationships in input video, as demonstrated in cookbooks/cosmos3/reasoner/run_with_vllm.ipynb.
import json
import requests
# Assume a vLLM-Omni server is running (cosmos-3-reasoner)
url = "http://localhost:8000/v1/chat/completions"
headers = {"Content-Type": "application/json"}
payload = {
"model": "nvidia/cosmos3-nano-reasoner",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [
{"type": "video_url",
"video_url": {"url": "https://example.com/robot_walk.mp4"}},
{"type": "text",
"text": "Explain the robot's motion and predict the next 2 seconds."}
]}
],
"max_tokens": 256,
"stream": False,
"extra_body": {"media_io_kwargs": {"video": {"fps": 4.0}}}
}
resp = requests.post(url, headers=headers, data=json.dumps(payload))
print(resp.json()["choices"][0]["message"]["content"])
The server applies the same mRoPE to video tokens, ensuring the reasoning module accurately references where (pixel coordinates) and when (frame timestamps) the robot appears.
Summary
- mRoPE uses tensor-product rotations across height, width, and time axes (
R_H ⊗ R_W ⊗ R_T) to encode 3-D positional information for every token. - Unified representation allows the same embedding scheme to work across images, video, audio, and action modalities in a common token space.
- Linear invertible operations enable efficient long-range spatial-temporal dependencies without quadratic memory growth.
- Shared architecture applies to both the autoregressive Reasoner and diffusion Generator transformers in Cosmos 3.
- Implementation is available in the Cosmos Framework at
cosmos_framework/models/rotary_embedding.py, with usage examples incookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynbandcookbooks/cosmos3/reasoner/run_with_vllm.ipynb.
Frequently Asked Questions
How does mRoPE differ from standard rotary position embeddings?
Standard rotary position embeddings operate on one-dimensional sequences by rotating query and key vectors based on relative position. mRoPE extends this to three dimensions by computing the tensor product of separate rotation matrices for height, width, and time, enabling simultaneous encoding of spatial and temporal relationships in a single operation.
Why is a shared positional encoding important for multimodal models?
Without a shared encoding, each modality would require separate positional representations, preventing the model from establishing direct spatial-temporal relationships between different data types. The unified mRoPE allows tokens from video frames and action trajectories to attend to each other based on their actual 3-D coordinates in space and time, as implemented in the Cosmos transformer architecture.
Where is mRoPE implemented in the NVIDIA Cosmos codebase?
The core implementation resides in the Cosmos Framework repository at cosmos_framework/models/rotary_embedding.py. The architectural documentation and conceptual overview are available in the main Cosmos repository's README.md and illustrated in cosmos3-model-architecture.png, while practical examples appear in the generator and reasoner cookbooks.
Can mRoPE handle variable frame rates and resolutions?
Yes, because mRoPE computes positional encodings based on the specific height, width, and temporal indices of each token, it naturally adapts to different resolutions and frame rates. The rotation matrices scale to accommodate varying sequence lengths and spatial dimensions while maintaining consistent relative positional relationships across modalities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →