# What Is mRoPE and How Does It Encode Spatial-Temporal Structure in Cosmos 3?

> Discover mRoPE in NVIDIA Cosmos 3 a 3D rotary position embedding that encodes spatial temporal coordinates to unify language images video audio and robot actions in a single transformer.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: deep-dive
- Published: 2026-06-06

---

**mRoPE (multi-dimensional rotary position embedding) is a 3D extension of rotary embeddings that encodes spatial (X, Y) and temporal (T) coordinates directly into query and key vectors, allowing NVIDIA Cosmos 3 to reason across language, images, video, audio, and robot actions within a unified transformer backbone.**

Cosmos 3 is an omnimodal world model developed by NVIDIA that processes multiple data modalities through a single transformer architecture. At the core of this system lies **mRoPE**, a positional encoding scheme that treats tokens as existing in three-dimensional space-time rather than on a linear sequence axis. According to the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), this unified representation enables both the Reasoner (causal) and Generator (diffusion) modes to share the same attention mechanisms across perceptual and action domains.

## What Is mRoPE?

**mRoPE** stands for multi-dimensional rotary position embedding, a variant of rotary position embedding (RoPE) that extends positional encoding from a single sequence dimension to three independent axes. While standard RoPE rotates query and key vectors based on a token's linear position in a sequence, mRoPE computes rotations based on spatial coordinates (X, Y) and temporal coordinates (T). This allows the model to understand where a token appears in physical space and when it occurs in time.

The mechanism works by applying **separate sinusoidal frequencies** to each axis. For a token at position `[x, y, t]`, mRoPE generates distinct rotation matrices for the spatial and temporal dimensions, then combines them into a composite rotation applied to the query and key vectors in the attention mechanism.

## How mRoPE Encodes Spatial-Temporal Structure

In Cosmos 3, every token—whether representing an image patch, video frame, audio window, or robot action—is assigned a 3D coordinate in the unified space. This design ensures that positional relationships remain consistent across modalities.

### 3D Coordinate Representation

The embedding treats each token as having three independent axes:

- **X and Y**: Spatial coordinates indicating horizontal and vertical position within a frame
- **T**: Temporal coordinate indicating the time step or sequence position

When a token is processed, its `[x, y, t]` coordinate determines a composite rotation matrix that is multiplied into the query and key vectors. Tokens that are close in space and time receive highly aligned query/key vectors, making them easy to attend to, while distant tokens are decorrelated.

### Cross-Modal Position Sharing

The same mRoPE logic applies uniformly across all modalities:

- **Images**: Encoded as 2D grids of patches using `(x, y)` coordinates with fixed `t`
- **Videos**: Extended to 3D volumes using `(x, y, t)` coordinates across frames
- **Audio**: Treated as 1D temporal streams using `t` coordinates with dummy spatial axes
- **Robot Actions**: Represented as trajectories using `t` for temporal ordering and optional spatial encoding for end-effector positions

Because all modalities share the same positional language, the transformer can attend to any token regardless of its source, enabling the model to align video scenes with spoken commands or plan robot motions conditioned on visual input.

## Implementation Example

Below is a minimal Python implementation illustrating how mRoPE injects 3D positional information into self-attention layers. While the actual Cosmos 3 implementation optimizes these operations on the GPU, the logic remains identical.

```python
import torch
import math
from transformers import PretrainedConfig

class MRopeConfig(PretrainedConfig):
    """Configuration for the multi-dimensional rotary embedding."""
    def __init__(self, dim, max_x=1024, max_y=1024, max_t=2048, **kwargs):
        super().__init__(**kwargs)
        self.dim = dim
        self.max_x = max_x
        self.max_y = max_y
        self.max_t = max_t

def mrope(position, config: MRopeConfig):
    """
    Compute the 3-D rotary matrix for a given (x, y, t) coordinate.
    Returns a tensor of shape (dim,) that will be applied to Q/K.
    """
    dim = config.dim
    inv_freq = 1.0 / (10000 ** (torch.arange(0, dim, 2, dtype=torch.float32) / dim))
    
    # Sinusoidal components for each axis

    sinusoid_x = position[0] * inv_freq
    sinusoid_y = position[1] * inv_freq
    sinusoid_t = position[2] * inv_freq
    
    # Combine the three axes

    sinusoid = torch.stack([sinusoid_x, sinusoid_y, sinusoid_t], dim=-1).flatten()[:dim]
    cos = torch.cos(sinusoid)
    sin = torch.sin(sinusoid)
    
    # Create rotation matrix (cos, -sin; sin, cos)

    rope = torch.stack([cos, -sin, sin, cos], dim=-1).view(dim // 2, 2, 2)
    return rope

class MRopeSelfAttention(torch.nn.Module):
    """Self-attention that incorporates mRoPE."""
    def __init__(self, hidden_size, num_heads, config: MRopeConfig):
        super().__init__()
        self.num_heads = num_heads
        self.head_dim = hidden_size // num_heads
        self.qkv = torch.nn.Linear(hidden_size, hidden_size * 3, bias=False)
        self.out = torch.nn.Linear(hidden_size, hidden_size, bias=False)
        self.config = config

    def forward(self, x, positions):
        """
        x         : (B, N, hidden)
        positions : (B, N, 3) - integer (x, y, t) for each token
        """
        B, N, _ = x.shape
        qkv = self.qkv(x).reshape(B, N, 3, self.num_heads, self.head_dim)
        q, k, v = qkv.unbind(dim=2)  # each (B, N, H, D)

        
        # Apply mRoPE to Q and K

        for b in range(B):
            for n in range(N):
                rope = mrope(positions[b, n], self.config)
                q[b, n] = torch.einsum('hd,dh->h', q[b, n], rope)
                k[b, n] = torch.einsum('hd,dh->h', k[b, n], rope)
        
        attn = torch.nn.functional.scaled_dot_product_attention(
            q, k, v, dropout_p=0.0, is_causal=False
        )
        attn = attn.reshape(B, N, -1)
        return self.out(attn)

```

**Example usage for video tokens:**

```python

# Generate tokens for a 4x4 spatial grid across 5 frames

B, T, H, W = 1, 5, 4, 4
hidden = 768
num_heads = 12
config = MRopeConfig(dim=hidden // num_heads, max_x=W, max_y=H, max_t=T)

# Dummy patch embeddings

tokens = torch.randn(B, T * H * W, hidden)

# Build (x, y, t) coordinates

coords = []
for t in range(T):
    for y in range(H):
        for x in range(W):
            coords.append([x, y, t])
coords = torch.tensor(coords).unsqueeze(0).repeat(B, 1, 1)  # (B, N, 3)

attn = MRopeSelfAttention(hidden, num_heads, config)
output = attn(tokens, coords)  # (B, N, hidden) with spatial-temporal context

```

## Source Files and Architecture References

The mRoPE implementation is documented across several key files in the NVIDIA Cosmos repository:

- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)**: Contains the architectural overview stating that "both modes share the same transformer architecture, multimodal attention layers, and a unified 3D multi-dimensional rotary position embedding (mRoPE) representation that encodes spatial and temporal structure across modalities"
- **`cookbooks/cosmos3/cosmos3-model-architecture.png`**: Visual diagram illustrating how rotary embeddings sit between token embeddings and attention layers
- **[`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md)**: Documents the Reasoner mode (causal self-attention) which utilizes mRoPE for autoregressive prediction
- **[`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md)**: Describes the Generator mode (full-attention diffusion) which also relies on mRoPE for joint denoising of multimodal streams

## Summary

- **mRoPE** extends standard rotary position embeddings to three dimensions (X, Y, T), encoding both spatial and temporal relationships directly into attention query/key vectors.
- The mechanism enables **cross-modal consistency** by applying the same positional encoding to image patches, video frames, audio windows, and robot action tokens.
- By embedding 3D coordinates into the attention mechanism, Cosmos 3 maintains a **unified world model** across its Reasoner (causal) and Generator (diffusion) operating modes.
- Implementation details are specified in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) and practical examples are provided in the `cookbooks/cosmos3/` directories for both reasoning and generation tasks.

## Frequently Asked Questions

### What does mRoPE stand for?

**mRoPE** stands for **multi-dimensional rotary position embedding**. It is a 3D extension of the rotary position embedding (RoPE) technique that encodes token positions using three independent axes (spatial X, Y and temporal T) rather than a single sequence index.

### How does mRoPE differ from standard RoPE?

Standard **RoPE** encodes position using a single linear index along the sequence dimension, rotating query/key vectors based on token order. **mRoPE** extends this to three dimensions, computing separate rotations for spatial coordinates (X, Y) and temporal coordinates (T), then combining them into a composite rotation matrix that preserves spatial-temporal relationships.

### Why is 3D positioning necessary for world models?

World models must reason about entities that exist in physical space and evolve over time. By encoding **(x, y, t)** coordinates directly into the attention mechanism, mRoPE allows the transformer to naturally understand spatial proximity (objects near each other in a frame) and temporal continuity (events progressing across frames) without requiring separate encoding schemes for each modality.

### Where is mRoPE implemented in the Cosmos 3 codebase?

The mRoPE architecture is defined at the repository level in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) as part of the unified transformer backbone. Practical usage examples appear in [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md) for causal prediction tasks and [`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md) for multimodal generation, with architectural diagrams available in `cookbooks/cosmos3/cosmos3-model-architecture.png`.