What Is mRoPE and How Does It Encode Spatial-Temporal Structure in Cosmos 3?
mRoPE (multi-dimensional rotary position embedding) is a 3D extension of rotary embeddings that encodes spatial (X, Y) and temporal (T) coordinates directly into query and key vectors, allowing NVIDIA Cosmos 3 to reason across language, images, video, audio, and robot actions within a unified transformer backbone.
Cosmos 3 is an omnimodal world model developed by NVIDIA that processes multiple data modalities through a single transformer architecture. At the core of this system lies mRoPE, a positional encoding scheme that treats tokens as existing in three-dimensional space-time rather than on a linear sequence axis. According to the repository's README.md, this unified representation enables both the Reasoner (causal) and Generator (diffusion) modes to share the same attention mechanisms across perceptual and action domains.
What Is mRoPE?
mRoPE stands for multi-dimensional rotary position embedding, a variant of rotary position embedding (RoPE) that extends positional encoding from a single sequence dimension to three independent axes. While standard RoPE rotates query and key vectors based on a token's linear position in a sequence, mRoPE computes rotations based on spatial coordinates (X, Y) and temporal coordinates (T). This allows the model to understand where a token appears in physical space and when it occurs in time.
The mechanism works by applying separate sinusoidal frequencies to each axis. For a token at position [x, y, t], mRoPE generates distinct rotation matrices for the spatial and temporal dimensions, then combines them into a composite rotation applied to the query and key vectors in the attention mechanism.
How mRoPE Encodes Spatial-Temporal Structure
In Cosmos 3, every token—whether representing an image patch, video frame, audio window, or robot action—is assigned a 3D coordinate in the unified space. This design ensures that positional relationships remain consistent across modalities.
3D Coordinate Representation
The embedding treats each token as having three independent axes:
- X and Y: Spatial coordinates indicating horizontal and vertical position within a frame
- T: Temporal coordinate indicating the time step or sequence position
When a token is processed, its [x, y, t] coordinate determines a composite rotation matrix that is multiplied into the query and key vectors. Tokens that are close in space and time receive highly aligned query/key vectors, making them easy to attend to, while distant tokens are decorrelated.
Cross-Modal Position Sharing
The same mRoPE logic applies uniformly across all modalities:
- Images: Encoded as 2D grids of patches using
(x, y)coordinates with fixedt - Videos: Extended to 3D volumes using
(x, y, t)coordinates across frames - Audio: Treated as 1D temporal streams using
tcoordinates with dummy spatial axes - Robot Actions: Represented as trajectories using
tfor temporal ordering and optional spatial encoding for end-effector positions
Because all modalities share the same positional language, the transformer can attend to any token regardless of its source, enabling the model to align video scenes with spoken commands or plan robot motions conditioned on visual input.
Implementation Example
Below is a minimal Python implementation illustrating how mRoPE injects 3D positional information into self-attention layers. While the actual Cosmos 3 implementation optimizes these operations on the GPU, the logic remains identical.
import torch
import math
from transformers import PretrainedConfig
class MRopeConfig(PretrainedConfig):
"""Configuration for the multi-dimensional rotary embedding."""
def __init__(self, dim, max_x=1024, max_y=1024, max_t=2048, **kwargs):
super().__init__(**kwargs)
self.dim = dim
self.max_x = max_x
self.max_y = max_y
self.max_t = max_t
def mrope(position, config: MRopeConfig):
"""
Compute the 3-D rotary matrix for a given (x, y, t) coordinate.
Returns a tensor of shape (dim,) that will be applied to Q/K.
"""
dim = config.dim
inv_freq = 1.0 / (10000 ** (torch.arange(0, dim, 2, dtype=torch.float32) / dim))
# Sinusoidal components for each axis
sinusoid_x = position[0] * inv_freq
sinusoid_y = position[1] * inv_freq
sinusoid_t = position[2] * inv_freq
# Combine the three axes
sinusoid = torch.stack([sinusoid_x, sinusoid_y, sinusoid_t], dim=-1).flatten()[:dim]
cos = torch.cos(sinusoid)
sin = torch.sin(sinusoid)
# Create rotation matrix (cos, -sin; sin, cos)
rope = torch.stack([cos, -sin, sin, cos], dim=-1).view(dim // 2, 2, 2)
return rope
class MRopeSelfAttention(torch.nn.Module):
"""Self-attention that incorporates mRoPE."""
def __init__(self, hidden_size, num_heads, config: MRopeConfig):
super().__init__()
self.num_heads = num_heads
self.head_dim = hidden_size // num_heads
self.qkv = torch.nn.Linear(hidden_size, hidden_size * 3, bias=False)
self.out = torch.nn.Linear(hidden_size, hidden_size, bias=False)
self.config = config
def forward(self, x, positions):
"""
x : (B, N, hidden)
positions : (B, N, 3) - integer (x, y, t) for each token
"""
B, N, _ = x.shape
qkv = self.qkv(x).reshape(B, N, 3, self.num_heads, self.head_dim)
q, k, v = qkv.unbind(dim=2) # each (B, N, H, D)
# Apply mRoPE to Q and K
for b in range(B):
for n in range(N):
rope = mrope(positions[b, n], self.config)
q[b, n] = torch.einsum('hd,dh->h', q[b, n], rope)
k[b, n] = torch.einsum('hd,dh->h', k[b, n], rope)
attn = torch.nn.functional.scaled_dot_product_attention(
q, k, v, dropout_p=0.0, is_causal=False
)
attn = attn.reshape(B, N, -1)
return self.out(attn)
Example usage for video tokens:
# Generate tokens for a 4x4 spatial grid across 5 frames
B, T, H, W = 1, 5, 4, 4
hidden = 768
num_heads = 12
config = MRopeConfig(dim=hidden // num_heads, max_x=W, max_y=H, max_t=T)
# Dummy patch embeddings
tokens = torch.randn(B, T * H * W, hidden)
# Build (x, y, t) coordinates
coords = []
for t in range(T):
for y in range(H):
for x in range(W):
coords.append([x, y, t])
coords = torch.tensor(coords).unsqueeze(0).repeat(B, 1, 1) # (B, N, 3)
attn = MRopeSelfAttention(hidden, num_heads, config)
output = attn(tokens, coords) # (B, N, hidden) with spatial-temporal context
Source Files and Architecture References
The mRoPE implementation is documented across several key files in the NVIDIA Cosmos repository:
README.md: Contains the architectural overview stating that "both modes share the same transformer architecture, multimodal attention layers, and a unified 3D multi-dimensional rotary position embedding (mRoPE) representation that encodes spatial and temporal structure across modalities"cookbooks/cosmos3/cosmos3-model-architecture.png: Visual diagram illustrating how rotary embeddings sit between token embeddings and attention layerscookbooks/cosmos3/reasoner/README.md: Documents the Reasoner mode (causal self-attention) which utilizes mRoPE for autoregressive predictioncookbooks/cosmos3/generator/audiovisual/README.md: Describes the Generator mode (full-attention diffusion) which also relies on mRoPE for joint denoising of multimodal streams
Summary
- mRoPE extends standard rotary position embeddings to three dimensions (X, Y, T), encoding both spatial and temporal relationships directly into attention query/key vectors.
- The mechanism enables cross-modal consistency by applying the same positional encoding to image patches, video frames, audio windows, and robot action tokens.
- By embedding 3D coordinates into the attention mechanism, Cosmos 3 maintains a unified world model across its Reasoner (causal) and Generator (diffusion) operating modes.
- Implementation details are specified in
README.mdand practical examples are provided in thecookbooks/cosmos3/directories for both reasoning and generation tasks.
Frequently Asked Questions
What does mRoPE stand for?
mRoPE stands for multi-dimensional rotary position embedding. It is a 3D extension of the rotary position embedding (RoPE) technique that encodes token positions using three independent axes (spatial X, Y and temporal T) rather than a single sequence index.
How does mRoPE differ from standard RoPE?
Standard RoPE encodes position using a single linear index along the sequence dimension, rotating query/key vectors based on token order. mRoPE extends this to three dimensions, computing separate rotations for spatial coordinates (X, Y) and temporal coordinates (T), then combining them into a composite rotation matrix that preserves spatial-temporal relationships.
Why is 3D positioning necessary for world models?
World models must reason about entities that exist in physical space and evolve over time. By encoding (x, y, t) coordinates directly into the attention mechanism, mRoPE allows the transformer to naturally understand spatial proximity (objects near each other in a frame) and temporal continuity (events progressing across frames) without requiring separate encoding schemes for each modality.
Where is mRoPE implemented in the Cosmos 3 codebase?
The mRoPE architecture is defined at the repository level in README.md as part of the unified transformer backbone. Practical usage examples appear in cookbooks/cosmos3/reasoner/README.md for causal prediction tasks and cookbooks/cosmos3/generator/audiovisual/README.md for multimodal generation, with architectural diagrams available in cookbooks/cosmos3/cosmos3-model-architecture.png.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →