Understanding mRoPE Position Embedding for Multimodal Reasoning in Cosmos 3

mRoPE (multi-dimensional Rotary Position Embedding) extends standard 2-D RoPE to three axes—height, width, and time—enabling NVIDIA Cosmos 3 to encode spatial-temporal relationships across vision, audio, and robot actions using a unified transformer backbone.

Cosmos 3 is an omnimodal world model that unifies autoregressive reasoning and diffusion-based generation under a single transformer architecture. Central to this capability is the mRoPE position embedding for multimodal reasoning, a generalized rotary embedding implemented in the NVIDIA/cosmos repository that simultaneously encodes spatial and temporal dimensions for all input modalities. This article examines the configuration and implementation details found in the evaluation notebooks and cookbooks to demonstrate how developers can leverage this mechanism for high-resolution inference.

From Standard RoPE to 3-D mRoPE

Standard Rotary Position Embedding (RoPE) rotates token embeddings by sinusoidal functions of the token's position, allowing self-attention to be position-aware without explicit absolute embeddings. While effective for 2-D vision-language models, standard RoPE cannot natively handle temporal sequences or cross-modal alignment.

mRoPE generalizes this to three dimensions (H, W, T). Each token receives a tensor of positional factors that are multiplied into query and key vectors along each dimension, yielding a single attention score that respects spatial-temporal continuity. According to the Cosmos source code, this is enabled via the configuration flag unified_3d_mrope as described in README.md (lines 74-78).

Configuration Parameters in Cosmos 3

The concrete implementation appears in evaluation/cosmos3/generator/paibench_c/run_with_cosmos_framework.ipynb (lines 3527-3541), which exposes several critical parameters for the PaIBench-C evaluation:

{
  "position_embedding_type": "unified_3d_mrope",
  "rope_h_extrapolation_ratio": 1.0,
  "rope_w_extrapolation_ratio": 1.0,
  "rope_t_extrapolation_ratio": 1.0,
  "unified_3d_mrope_reset_spatial_ids": true,
  "unified_3d_mrope_temporal_modality_margin": 15000
}
  • Extrapolation ratios: rope_h_extrapolation_ratio, rope_w_extrapolation_ratio, and rope_t_extrapolation_ratio control how far the sinusoid can be extrapolated beyond training resolution. Set to 1.0 by default, these can be tuned for higher-resolution inference.
  • Spatial-ID reset: unified_3d_mrope_reset_spatial_ids=True re-indexes spatial IDs when new scenes start, preventing position information leakage between unrelated clips.
  • Temporal-modality margin: unified_3d_mrope_temporal_modality_margin=15000 reserves token IDs for temporal tokens (audio or actions), distinguishing them from pure spatial tokens.

Implementing mRoPE with the Cosmos Framework

Below are minimal, self-contained examples demonstrating how to enable mRoPE when loading Cosmos 3 models with the public Cosmos Framework API.

Loading a Pre-trained Model

from cosmos_framework import CosmosModel, CosmosConfig

# Load the default config and override the position embedding type

cfg = CosmosConfig.from_pretrained("nvidia/cosmos3")
cfg.position_embedding_type = "unified_3d_mrope"
cfg.rope_h_extrapolation_ratio = 1.0
cfg.rope_w_extrapolation_ratio = 1.0
cfg.rope_t_extrapolation_ratio = 1.0
cfg.unified_3d_mrope_reset_spatial_ids = True
cfg.unified_3d_mrope_temporal_modality_margin = 15000

model = CosmosModel.from_pretrained("nvidia/cosmos3", config=cfg)

This configuration mirrors the JSON block found in lines 3527-3541 of the PaIBench-C evaluation notebook.

Processing Multimodal Inputs


# Assume video is (B, T, C, H, W), audio is (B, T, A), actions (B, T, D)

inputs = {
    "video": video_tensor,
    "audio": audio_tensor,
    "actions": actions_tensor,
}

# The model internally flattens dimensions and applies mRoPE

output = model(**inputs)

The model's internal tokenization projects each modality onto the same (H, W, T) lattice, allowing attention layers to automatically respect spatio-temporal geometry encoded by mRoPE.

Configuring High-Resolution Extrapolation


# Double the spatial frequency for 4K video inference

cfg.rope_h_extrapolation_ratio = 2.0
cfg.rope_w_extrapolation_ratio = 2.0
model = CosmosModel.from_pretrained("nvidia/cosmos3", config=cfg)

Increasing these ratios allows the sinusoidal basis to cover larger coordinate ranges without retraining, as implemented in the Cosmos 3 generator pipeline.

Why mRoPE Matters for Multimodal Reasoning

The mRoPE position embedding provides four key advantages for Cosmos 3's omnimodal capabilities, as reflected in the architecture diagram cookbooks/cosmos3/cosmos3-model-architecture.png and the reasoner documentation:

  1. Unified token space: By treating frames as 3-D grids (H × W × T), the same attention kernels attend across space and time without separate spatial-only and temporal-only pathways.
  2. Consistent geometry: Rotational invariance enables the model to learn relative relationships (e.g., "the robot's left hand is to the left of the block") rather than absolute pixel coordinates, which is essential for reasoning about dynamic scenes.
  3. Scalable resolution: Extrapolation ratios allow the model to handle higher-resolution video or longer audio streams beyond training dimensions without retraining.
  4. Modality-agnostic design: Audio and robot action tokens occupy the T-axis, enabling cross-modal attention (e.g., "the sound of a servo coincides with the robot's grasp") through the same positional mechanism.

As documented in cookbooks/cosmos3/reasoner/README.md, this architecture enables the reasoner mode (causal self-attention) to consume mRoPE-augmented tokens for coherent spatial-temporal reasoning over heterogeneous streams.

Summary

  • mRoPE extends 2-D rotary embeddings to three dimensions (height, width, time) for unified spatial-temporal reasoning in NVIDIA/cosmos.
  • Configuration in evaluation/cosmos3/generator/paibench_c/run_with_cosmos_framework.ipynb (lines 3527-3541) uses unified_3d_mrope with extrapolation ratios and spatial-ID reset controls.
  • The mechanism supports multimodal reasoning across video, audio, and robot actions by projecting all modalities onto a shared (H, W, T) lattice.
  • Extrapolation ratios (rope_h_extrapolation_ratio, rope_w_extrapolation_ratio, rope_t_extrapolation_ratio) enable inference at resolutions beyond training data.

Frequently Asked Questions

What is the difference between standard RoPE and mRoPE?

Standard RoPE encodes 2-D spatial positions using rotational sinusoids, suitable for static images. mRoPE adds a temporal dimension (T), creating a 3-D positional encoding that simultaneously handles height, width, and time. This allows the transformer to compute attention across video frames, audio sequences, and action trajectories using the same kernel, as required for the omnimodal world modeling in Cosmos 3.

How do I enable mRoPE in Cosmos 3?

Set position_embedding_type to "unified_3d_mrope" in your CosmosConfig before model initialization. As documented in lines 3527-3541 of the PaIBench-C evaluation notebook, you should also configure unified_3d_mrope_reset_spatial_ids and unified_3d_mrope_temporal_modality_margin to handle scene transitions and multimodal token separation.

What are extrapolation ratios used for?

The parameters rope_h_extrapolation_ratio, rope_w_extrapolation_ratio, and rope_t_extrapolation_ratio scale the sinusoidal frequency basis. Values greater than 1.0 allow the model to generalize to larger spatial resolutions or longer temporal sequences than seen during training, enabling 4K video inference without retraining the NVIDIA Cosmos model.

Why does Cosmos 3 reset spatial IDs between scenes?

When unified_3d_mrope_reset_spatial_ids is set to True, the model re-indexes spatial coordinates at scene boundaries. This prevents positional information from leaking between unrelated video clips, ensuring that the first frame of a new scene receives fresh spatial indexing rather than continuing the coordinate stream from the previous scene.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →