Understanding the Mixture-of-Transformers Architecture in Cosmos 3
Cosmos 3 implements a unified Mixture-of-Transformers (MoT) architecture that combines a single transformer backbone with dual functional heads—an Autoregressive Reasoner and a Diffusion Generator—to simultaneously handle text reasoning and multimodal generation across images, video, audio, and actions.
The NVIDIA/cosmos repository introduces Cosmos 3 as a foundation model built around this novel Mixture-of-Transformers design. By sharing transformer weights between reasoning and generation tasks, the architecture reduces memory footprint while enabling seamless cross-modal transfer learning.
Core Components of the MoT Architecture
According to the README.md in the NVIDIA/cosmos repository (lines 70-76), the Mixture-of-Transformers architecture rests on a single transformer stack reused for all modalities. This unified backbone projects tokens from text, images, video, audio, and action spaces into a common embedding space before processing them through shared attention layers.
The Dual-Head Design
The architecture splits into two functional heads after the shared transformer backbone:
-
Autoregressive (AR) Reasoner head: Processes language and visual-understanding tokens using causal self-attention. This head enables next-token prediction for tasks such as captioning, planning, and world-reasoning.
-
Diffusion (DM) Generator head: Handles noisy multimodal tokens with full-attention denoising. This head produces coherent outputs for image, video, audio, and action generation tasks.
Both heads share the same transformer weights, multimodal attention mechanisms, and positional encoding infrastructure, allowing the model to switch seamlessly between reasoning and generation modes as described in the repository's architecture documentation.
3-D Multi-Dimensional Rotary Position Embedding (mRoPE)
Cosmos 3 employs mRoPE, a unified positional encoding that captures spatial (X-Y), temporal (frame), and modality-specific axes. This 3-D rotary embedding enables consistent reasoning about space-time structures across images, video frames, and action trajectories, ensuring that the transformer understands both where objects appear and when events occur.
Multimodal Tokenization Strategy
The architecture relies on mode-agnostic tokenization to handle diverse inputs. Text tokens, visual patches, audio spectrogram patches, and action vectors are all flattened into a single sequence format. This approach allows the Mixture-of-Transformers to process heterogeneous modalities without architectural changes, feeding everything through the same attention layers and mRoPE encodings.
Implementing Cosmos 3: Inference Patterns
The NVIDIA/cosmos repository provides distinct implementation paths for the two functional heads, while the underlying Cosmos3Omni classes share the same MoT weights.
Text and Image Reasoning with Transformers
For AR reasoning tasks, use Cosmos3OmniForConditionalGeneration with the Transformers library. This path utilizes the causal self-attention head for next-token prediction on interleaved text and image inputs.
from pathlib import Path
import torch
from transformers import AutoProcessor, Cosmos3OmniForConditionalGeneration
model_id = "nvidia/Cosmos3-Nano"
image_path = Path("cookbooks/cosmos3/reasoner/assets/robot_153.jpg")
processor = AutoProcessor.from_pretrained(model_id)
model = Cosmos3OmniForConditionalGeneration.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto"
)
messages = [
{
"role": "user",
"content": [
{"type": "image", "path": str(image_path)},
{"type": "text", "text": "Give a detailed caption of this scene."},
],
}
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device, torch.bfloat16)
output_ids = model.generate(**inputs, do_sample=False, max_new_tokens=512)
output_text = processor.batch_decode(
output_ids[:, inputs.input_ids.shape[-1] :],
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0]
print(output_text)
Reference implementation available in cookbooks/cosmos3/reasoner/run_with_transformers.ipynb.
Multimodal Generation with Diffusers
For generation tasks, invoke the Diffusion head using Cosmos3OmniPipeline with a UniPCMultistepScheduler. This path handles the full-attention denoising process for video and image synthesis.
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
from diffusers.utils import export_to_video
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano", torch_dtype=torch.bfloat16, device_map="cuda"
)
pipe.scheduler = UniPCMultistepScheduler.from_config(pipe.scheduler.config, flow_shift=10.0)
result = pipe(
prompt="A delivery drone flies over a futuristic city at sunset.",
num_frames=189, # generate a 7-second video @ 24 fps
height=720,
width=1280,
fps=24,
num_inference_steps=35,
guidance_scale=6.0,
generator=torch.Generator(device="cuda").manual_seed(1234),
)
# Save video output
export_to_video(result.video, "drone_city.mp4", fps=24)
See cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb for the complete notebook implementation.
Unified Serving with vLLM-Omni
For production deployment, the vLLM-Omni server provides an OpenAI-compatible API that routes requests to either the AR or Diffusion head automatically based on the input context.
Start the server:
docker run --gpus all -p 8000:8000 \
-v "$(pwd):/workspace" \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Nano \
--omni --model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / --port 8000 --init-timeout 1800
Query the unified endpoint:
import requests, json, base64
# Encode an input image as a data URI
with open("cookbooks/cosmos3/reasoner/assets/robot_153.jpg", "rb") as f:
img_b64 = base64.b64encode(f.read()).decode()
image_uri = f"data:image/jpeg;base64,{img_b64}"
payload = {
"model": "nvidia/Cosmos3-Nano",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": image_uri}},
{"type": "text", "text": "Summarize the scene in two sentences."}
]}
],
"max_tokens": 256,
}
resp = requests.post("http://localhost:8000/v1/chat/completions", json=payload)
print(resp.json()["choices"][0]["message"]["content"])
The vLLM-Omni implementation is documented in cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb.
Summary
- Cosmos 3 uses a unified Mixture-of-Transformers architecture with a single backbone shared between reasoning and generation.
- Dual functional heads—Autoregressive for text reasoning and Diffusion for multimodal generation—operate on the same transformer weights.
- mRoPE provides 3-D positional encoding across spatial, temporal, and modality axes.
- Mode-agnostic tokenization flattens text, images, video, audio, and actions into a common sequence format.
- Implementation files in the NVIDIA/cosmos repository include
cookbooks/cosmos3/reasoner/run_with_transformers.ipynbfor AR tasks andcookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynbfor generation.
Frequently Asked Questions
What distinguishes the Mixture-of-Transformers from standard mixture-of-experts models?
While Mixture-of-Experts (MoE) architectures route tokens to different neural network experts, the Mixture-of-Transformers in Cosmos 3 uses a single transformer backbone for all modalities. The "mixture" refers to the dual-head design—AR and Diffusion—sharing the same attention layers and mRoPE embeddings, not to sparse expert selection. This design reduces parameters and improves cross-modal transfer.
How does mRoPE handle different modalities simultaneously?
The 3-D multi-dimensional rotary position embedding (mRoPE) encodes three distinct axes: spatial coordinates (X-Y) for visual patches, temporal frame indices for video sequences, and modality-specific identifiers. This allows the transformer to attend to relationships like "the robot arm moves left in frame 5" consistently across text descriptions and visual inputs.
Can Cosmos 3 process text, images, and video in a single forward pass?
Yes. The mode-agnostic tokenization converts all inputs—text tokens, image patches, video frames, and audio spectrograms—into a flat sequence. The unified transformer processes these tokens together, enabling in-context reasoning such as generating a video continuation from a text description and starter frames. The vLLM-Omni server automates this routing through the appropriate head.
Where can I find the architecture diagram for Cosmos 3?
The official architecture visualization is located at cookbooks/cosmos3/cosmos3-model-architecture.png in the NVIDIA/cosmos repository. This diagram illustrates the unified transformer stack, the split between AR and Diffusion heads, and the mRoPE embedding scheme referenced in README.md lines 70-76.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →