# How the Mixture-of-Transformers (MoT) Architecture in Cosmos 3 Unifies Autoregressive Reasoning and Diffusion Generation

> Discover how NVIDIA Cosmos 3's Mixture-of-Transformers (MoT) unifies autoregressive reasoning and diffusion generation using a single transformer for advanced multimodal applications.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: deep-dive
- Published: 2026-06-14

---

**The Mixture-of-Transformers (MoT) architecture in NVIDIA Cosmos 3 employs a single unified transformer that switches between autoregressive reasoning (causal attention) and diffusion generation (bidirectional attention) while sharing identical layers, multimodal attention mechanisms, and 3-D rotary position embeddings.**

The NVIDIA Cosmos repository introduces a paradigm shift in multimodal AI by eliminating the traditional separation between perception and generation models. The **Mixture-of-Transformers (MoT)** architecture enables one model to operate in two distinct modes—Reasoner and Generator—using the same underlying transformer weights and latent representations. This design ensures that reasoning capabilities directly inform synthesis tasks without requiring separate network pipelines or incompatible latent spaces.

## Dual-Mode Architecture: Reasoner vs. Generator

The MoT architecture toggles between two complementary operational modes within the same `OmniMoTModel` instance:

- **Reasoner Mode**: Implements **autoregressive (AR)** processing with **causal self-attention**, sequentially predicting next tokens for language understanding, visual perception, and action planning tasks.
- **Generator Mode**: Implements **diffusion (DM)** processing with **full bidirectional attention**, iteratively denoising multimodal tokens to synthesize high-fidelity images, video, audio, and action trajectories.

Both modes utilize the complete transformer stack—including multi-head attention, MLP blocks, and layer normalization—ensuring that knowledge acquired during reasoning tasks directly improves generation quality.

## Shared Architecture Components

### Identical Transformer Layers

According to the source code in [`cosmos_framework/model/vfm/omni_mot_model.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/model/vfm/omni_mot_model.py), the `OmniMoTModel` class maintains the same stack of transformer layers regardless of operational mode. When switching between Reasoner and Generator modes, the model does not swap weights or architecture; it merely alters the attention masking strategy and forward processing pathway.

### Multimodal Attention with Modality Tags

Each token processed by the MoT carries a discrete modality tag, enabling **multimodal attention** across language, vision, audio, and action tokens. This mechanism allows the transformer to maintain cross-modal relationships whether performing robot action planning or generating video sequences from textual descriptions.

### Unified 3-D Multi-dimensional Rotary Position Embeddings (mRoPE)

The architecture implements **unified 3-D mRoPE** to encode spatial (x-y), temporal (t), and modality-specific dimensions consistently. This single positional encoding scheme captures geometric relationships across all modalities, ensuring that temporal reasoning in Reasoner Mode translates directly to temporal coherence in Generator Mode.

## Implementing Mode Switching in Practice

The Cosmos framework exposes mode selection through the `mode` parameter in the `OmniMoTModel` class. Users instantiate a single model and toggle between pathways for different tasks without reloading checkpoints.

### Running Reasoner Mode (Autoregressive)

To perform next-token prediction for multimodal reasoning tasks such as video question answering or robotics planning, configure the model with `mode="reasoner"`:

```python
from cosmos_framework.model.vfm.omni_mot_model import OmniMoTModel
from cosmos_framework.configs.base.defaults import model_config

# Load unified MoT configuration containing both AR and Diffusion settings

cfg = model_config.OmniMoTModelConfig.from_json_file(
    "cosmos_framework/configs/omni_mot/qwen3_vl_mot.json"
)
model = OmniMoTModel.from_config(cfg, torch_dtype="float16").to("cuda")

# Prepare multimodal prompt for reasoning

prompt = "<|startoftext|>User: Describe the next action of the robot given the video.<|endoftext|>"
tokens = model.tokenizer.encode(prompt, return_tensors="pt").to("cuda")

# Autoregressive generation with causal self-attention

output_ids = model.generate(
    tokens,
    max_new_tokens=128,
    do_sample=False,
    attention_mask=None,   # Causal mask applied automatically

    mode="reasoner"        # Selects the AR pathway

)
print(model.tokenizer.decode(output_ids[0]))

```

### Running Generator Mode (Diffusion)

For synthesis tasks, the same model instance switches to `mode="generator"` to utilize the diffusion pathway with full bidirectional attention:

```python
from cosmos_framework.pipeline.diffusion import DiffusionPipeline

# Reuse existing model; switch to diffusion mode

pipeline = DiffusionPipeline(model, scheduler="ddpm", num_steps=50)

# Condition on textual description (potentially from Reasoner output)

condition = "A robot moving toward a red cube on a wooden table."
generated_video = pipeline.generate(
    condition,
    height=256,
    width=256,
    num_frames=16,
    mode="generator"      # Selects the diffusion pathway

)

# Save output tensor of shape (C, T, H, W)

import torchvision.io as io
io.write_video("robot_action.mp4", generated_video, fps=8)

```

## Configuration and Model Weights

The unified architecture relies on shared configuration management through [`cosmos_framework/configs/base/defaults/model_config.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/configs/base/defaults/model_config.py). The `OmniMoTModelConfig` class loads both autoregressive and diffusion hyperparameters from single JSON files such as [`cosmos_framework/configs/omni_mot/qwen3_vl_mot.json`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/configs/omni_mot/qwen3_vl_mot.json), ensuring that one checkpoint contains compatible weights for both operational modes.

For complete implementation examples, reference the official notebooks:
- `cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb` demonstrates Reasoner Mode on video understanding and planning tasks.
- `cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb` demonstrates Generator Mode for multimodal audio-visual synthesis.

## Summary

- **Single unified transformer**: The MoT architecture uses identical layers for both reasoning and generation, eliminating the need for separate perception and synthesis models.
- **Dual attention mechanisms**: Causal self-attention for autoregressive Reasoner Mode and full bidirectional attention for diffusion Generator Mode.
- **Modality-agnostic processing**: Multimodal attention layers and unified 3-D mRoPE embeddings handle language, vision, audio, and action tokens interchangeably.
- **Runtime mode switching**: The `OmniMoTModel` class in [`cosmos_framework/model/vfm/omni_mot_model.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/model/vfm/omni_mot_model.py) toggles between pathways via the `mode` parameter without reloading weights.
- **Coherent latent space**: Shared representations ensure that reasoning outputs directly condition generation tasks without pipeline stitching or latent space translation.

## Frequently Asked Questions

### What is the Mixture-of-Transformers (MoT) architecture in Cosmos 3?

The Mixture-of-Transformers (MoT) architecture is a unified transformer design in NVIDIA Cosmos 3 that combines autoregressive reasoning and diffusion generation within a single network. It operates in two modes—Reasoner (causal attention) and Generator (bidirectional attention)—while sharing all transformer layers, multimodal attention mechanisms, and 3-D rotary position embeddings.

### How does Cosmos 3 switch between reasoning and generation modes?

Cosmos 3 switches modes at runtime through the `mode` parameter in the `OmniMoTModel.generate()` method. Setting `mode="reasoner"` enables causal self-attention for next-token prediction, while `mode="generator"` activates the diffusion pathway with full bidirectional attention. Both modes use the same checkpoint and transformer weights stored in [`cosmos_framework/model/vfm/omni_mot_model.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/model/vfm/omni_mot_model.py).

### What file contains the core implementation of the MoT architecture?

The core implementation resides in [`cosmos_framework/model/vfm/omni_mot_model.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/model/vfm/omni_mot_model.py), which defines the `OmniMoTModel` class. This file handles the mode switching logic, attention masking, and shared transformer layers. Configuration management is located in [`cosmos_framework/configs/base/defaults/model_config.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/configs/base/defaults/model_config.py).

### Why does the MoT architecture use unified 3-D mRoPE embeddings?

Unified 3-D multi-dimensional rotary position embeddings (mRoPE) provide a consistent coordinate system for spatial (x-y), temporal (t), and modality-specific dimensions across all tokens. This ensures that geometric relationships learned during autoregressive reasoning remain valid when the model switches to diffusion generation, maintaining coherence between understanding and synthesis tasks.