# How the Mixture-of-Transformers (MoT) Architecture in Cosmos 3 Unifies Reasoning and Generation

> Explore the Mixture-of-Transformers (MoT) in Cosmos 3 uniting causal reasoning and multimodal generation with AR and DM transformer heads via a shared backbone and mRoPE.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: architecture
- Published: 2026-06-06

---

**Cosmos 3 uses a unified Mixture-of-Transformers (MoT) architecture that pairs an autoregressive (AR) transformer head for causal reasoning with a diffusion transformer (DM) head for bidirectional multimodal generation, all backed by a single shared transformer backbone and 3-D multi-dimensional rotary position embeddings (mRoPE).**

The **NVIDIA/cosmos** repository implements Cosmos 3 as an omnimodal world model built on this **Mixture-of-Transformers (MoT) architecture**. As stated in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 74), “Cosmos 3 is an omnimodal world model built on a unified Mixture-of-Transformers (MoT) architecture that combines an autoregressive (AR) transformer for reasoning with a diffusion transformer (DM) for multimodal generation.” This design enables one model to handle both step-by-step inference and creative synthesis across language, vision, audio, and action modalities.

## Core Components of the MoT Architecture

The Cosmos 3 MoT architecture is not a mixture of separate models, but a **single backbone** with two specialized heads. Each head is optimized for a distinct computational pattern—causal autoregression for reasoning and full-attention diffusion for generation—while drawing on the same shared parameters.

### Autoregressive Transformer for Reasoning

The **autoregressive (AR) transformer** head drives **reasoning mode** by applying **causal (unidirectional) self-attention** over input tokens. This next-token-prediction mechanism processes sequences from language and visual inputs sequentially, making it ideal for perception, planning, world-state inference, and action selection. Because attention is restricted to prior tokens only, the AR head naturally handles step-by-step tasks such as answering questions about a video frame or predicting future states.

### Diffusion Transformer for Generation

The **diffusion (DM) transformer** head powers **generation mode** through a denoising process that relies on **full (bidirectional) attention** across the entire token sequence. Input tokens are first corrupted with noise and then iteratively denoised to produce coherent multimodal outputs, including images, video, audio, and action trajectories. This bidirectional context allows the model to jointly synthesize complex content, which is why it is used for creative tasks like text-to-image and video synthesis.

## Shared Backbone and Position Encoding

Both the AR and DM heads in Cosmos 3 sit on top of the **same stack of transformer layers and multimodal attention blocks**, allowing both reasoning and generation to use identical weights. The architecture encodes spatial-temporal structure with a **3-D multi-dimensional rotary position embedding (mRoPE)** that is applied uniformly across all modalities. According to the architecture description in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 53–74), this unified backbone lets the model reason over and generate images, video, audio, and action data without maintaining separate parameters or retraining when switching heads.

## Invoking Reasoning and Generation Modes

The NVIDIA/cosmos cookbooks provide concrete CLI and Python examples for running each head. The entry points are located in `cookbooks/cosmos3/reasoner/` for the AR head and `cookbooks/cosmos3/generator/audiovisual/` for the DM head.

### CLI Commands for Reasoner Mode

To launch the autoregressive reasoner, follow the instructions in [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md) and the notebook `cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb`. The recommended CLI invocation is:

```bash
vllm serve nvidia/Cosmos3-Nano \
  --omni \
  --model-class-name Cosmos3OmniPipeline \
  --reasoner

```

This command instantiates the `Cosmos3OmniPipeline` class and starts a server for perception and planning tasks.

### CLI Commands for Generator Mode

For diffusion-based generation, refer to [`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md) and the notebook `cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb`. Launch the generator with:

```bash
vllm serve nvidia/Cosmos3-Nano \
  --omni \
  --model-class-name Cosmos3OmniDiffusersPipeline \
  --generator

```

This starts an endpoint through the `Cosmos3OmniDiffusersPipeline` class for synthesizing images, video, audio, or action trajectories.

### Switching Between Modes in Python

You can also instantiate either pipeline directly in Python to toggle between reasoning and generation programmatically. Both load the same pretrained backbone and differ only in the active head:

```python
from cosmos import Cosmos3OmniPipeline, Cosmos3OmniDiffusersPipeline

# Reasoner (AR) — ask a question about a video frame

reasoner = Cosmos3OmniPipeline.from_pretrained("nvidia/Cosmos3-Nano")
response = reasoner.predict(
    "Describe the scene in the third frame of this video.",
    video=video_bytes
)

# Generator (Diffusion) — synthesize a short video from text

generator = Cosmos3OmniDiffusersPipeline.from_pretrained("nvidia/Cosmos3-Nano")
video = generator.generate(
    "A robot pouring water into a glass.",
    num_frames=16
)

```

Because `Cosmos3OmniPipeline` and `Cosmos3OmniDiffusersPipeline` share the underlying model weights, you can switch modes without loading a separate checkpoint.

## Source Files and Architecture Diagrams

The canonical description of the MoT design resides in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 53–74), which defines the AR and DM heads and the role of 3-D mRoPE. Supporting implementation guides and visual references are located at the following paths:

- [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md) — setup and CLI instructions for reasoning.
- `cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb` — interactive notebook for AR mode.
- [`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md) — setup and CLI instructions for generation.
- `cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb` — interactive notebook for DM mode.
- `cookbooks/cosmos3/cosmos3-model-architecture.png` — diagram of the unified AR + DM architecture.

## Summary

- The **Mixture-of-Transformers (MoT) architecture in Cosmos 3** fuses an autoregressive head for reasoning and a diffusion head for generation into one model.
- The **AR head** uses **causal self-attention** for next-token prediction, enabling step-by-step perception, planning, and inference.
- The **DM head** uses **full bidirectional attention** to denoise corrupt token streams and generate multimodal outputs like images, video, and audio.
- Both heads share the **same transformer backbone** and **3-D mRoPE**, eliminating the need for separate weights or retraining when switching tasks.
- Users select the active mode through dedicated pipeline classes—`Cosmos3OmniPipeline` for reasoning and `Cosmos3OmniDiffusersPipeline` for generation—via CLI or Python.

## Frequently Asked Questions

### What is the Mixture-of-Transformers (MoT) architecture in Cosmos 3?

The MoT architecture is a unified transformer design in the NVIDIA/cosmos repository that merges an autoregressive head for causal reasoning with a diffusion transformer head for bidirectional generation. Both heads run on a single shared backbone equipped with 3-D mRoPE, enabling one model to process and synthesize language, vision, audio, and action data.

### How does Cosmos 3 handle reasoning tasks?

Cosmos 3 handles reasoning through its autoregressive head, which applies causal self-attention to perform next-token prediction over sequential inputs. This mode is exposed via the `Cosmos3OmniPipeline` class and is suited for tasks such as perception, planning, and question answering about visual content.

### How does Cosmos 3 generate multimodal content?

Cosmos 3 generates content through its diffusion transformer head, which corrupts tokens with noise and then denoises them using full bidirectional attention. This generator mode is accessed via the `Cosmos3OmniDiffusersPipeline` class and can produce images, video, audio, and action trajectories from text or other prompts.

### Do reasoning and generation use separate model weights?

No, both modes share the same transformer layers, multimodal attention blocks, and 3-D mRoPE embeddings. The AR and DM heads are specialized branches on top of this shared backbone, so you can switch between reasoning and generation without loading separate checkpoints.