# Reasoner vs Generator Runtime Surfaces in NVIDIA Cosmos: Architecture and Usage

> Understand the difference between NVIDIA Cosmos Reasoner and Generator runtimes. Explore their distinct attention mechanisms and diverse output capabilities for multimodal applications.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: architecture
- Published: 2026-06-12

---

**The Reasoner and Generator runtime surfaces in Cosmos 3 share the same Mixture-of-Transformers (MoT) backbone but operate through fundamentally different attention mechanisms: the Reasoner employs autoregressive processing to generate text outputs from vision and language inputs, while the Generator utilizes diffusion-based denoising to produce vision, audio, and action outputs from multimodal prompts.**

NVIDIA Cosmos exposes two distinct runtime surfaces through a unified architecture that enables both multimodal understanding and world generation capabilities. According to the source code in the `NVIDIA/cosmos` repository, both surfaces utilize identical transformer backbones and multi-dimensional rotary position embeddings (mRoPE), yet they activate different computational sub-graphs depending on whether the task requires reasoning or synthesis.

## Input and Output Differences Between Surfaces

The primary distinction between these runtime surfaces lies in their accepted input modalities and resulting output types. As documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 55-61), the surfaces differ as follows:

| Surface | Accepted Inputs | Produced Outputs | Typical Use Cases |
|---------|----------------|------------------|-------------------|
| **Reasoner** | Text + vision (image or video) | Text | World understanding, grounding, physical reasoning, task planning, action forecasting, embodied-agent reasoning, autonomous-system decision making |
| **Generator** | Text + vision + sound + action | Vision, sound, action (e.g., images, videos, audio streams, action trajectories) | World generation, world simulation, future prediction, synthetic data generation, policy learning, robot training |

## Architectural Distinctions: Autoregressive vs Diffusion

While both surfaces share underlying infrastructure, they implement divergent attention mechanisms optimized for their specific tasks.

### Reasoner Surface – Autoregressive Mode

The Reasoner operates in **autoregressive mode**, processing language and visual tokens through causal self-attention mechanisms. This architecture enables next-token prediction optimized for **understanding** tasks such as captioning, temporal localization, and 2-D grounding. When you invoke the Reasoner surface, the model processes input sequences sequentially, attending only to previous tokens to generate coherent textual responses.

### Generator Surface – Diffusion Mode

The Generator surface operates in **diffusion mode**, utilizing full-attention diffusion steps to denoise multimodal tokens. Unlike the causal approach of the Reasoner, the Generator processes noisy representations of image, video, audio, and action tokens simultaneously through the complete attention layers. This enables the synthesis of high-fidelity **generative** outputs including text-to-image/video conversion, image-to-video generation, audio-synchronized video creation, and action trajectory rollouts.

## Shared Components and Unified Backbone

Despite their operational differences, both runtime surfaces leverage identical core components:

* **The same transformer backbone** – Both surfaces utilize the shared MoT architecture without requiring separate model weights.
* **Unified multimodal attention layers** – The attention mechanisms are repurposed depending on the runtime mode (causal masking for Reasoner, full attention for Generator).
* **3-D multi-dimensional rotary position embedding (mRoPE)** – This encodes spatial-temporal structure consistently across all modalities, ensuring coherent handling of vision, audio, and action data.
* **Single checkpoint compatibility** – A single model checkpoint can serve either surface; the choice of runtime determines which sub-graph of the model is activated.

## Practical Implementation Examples

The `NVIDIA/cosmos` repository provides specific implementations demonstrating how to invoke each surface through different APIs.

### Reasoner Inference with vLLM

To generate text outputs from visual inputs using the Reasoner surface, implement the vLLM API following the schema described in the repository. The payload follows the Qwen-3-VL message format, returning text strings because the Reasoner surface produces only textual answers:

```python
import requests, json, base64, pathlib

# Load an image (or video) and encode it as base64

with open("example.jpg", "rb") as f:
    img_b64 = base64.b64encode(f.read()).decode()

payload = {
    "model": "cosmos3-reasoner-nano",
    "messages": [
        {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
        {"role": "user", "content": [
            {"type": "image", "image": img_b64},
            {"type": "text", "text": "Describe the scene in detail."}
        ]}
    ],
    "max_tokens": 256,
    "temperature": 0.6,
    "top_p": 0.95
}

resp = requests.post("http://localhost:8000/v1/chat/completions", json=payload)
print(resp.json()["choices"][0]["message"]["content"])

```

For a complete end-to-end implementation, reference `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` in the repository.

### Text-to-Image Generation with Diffusers

To invoke the Generator surface for vision synthesis, use the Diffusers pipeline which activates the diffusion sub-graph via the wrapper implemented in the source:

```python
from diffusers import Cosmos3OmniPipeline
import torch

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    variant="diffusion"
).to("cuda")

prompt = "A futuristic laboratory with robots assembling microchips"
image = pipe(prompt, height=720, width=1280, max_new_tokens=1024).images[0]
image.save("generated.png")

```

This implementation denoises image tokens to output vision artifacts (PNG files) rather than text. See `cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb` for comprehensive Generator usage patterns.

### Multimodal Video and Audio Generation

The Generator surface supports multi-stream output when the request includes audio parameters. The following example demonstrates video generation with synchronized sound:

```python
payload = {
    "model": "cosmos3-super",
    "messages": [
        {"role": "user", "content": [
            {"type": "text", "text": "Create a short video of a robot pouring water, with realistic sound."}
        ]}
    ],
    "max_tokens": 20000,
    "height": 720,
    "width": 1280,
    "num_frames": 60,
    "audio": True
}
resp = requests.post("http://localhost:8001/v1/chat/completions", json=payload)
video_bytes = resp.json()["choices"][0]["message"]["content"]["video"]
with open("robot_water.mp4", "wb") as f:
    f.write(base64.b64decode(video_bytes))

```

When `audio=True` is specified, the Generator emits video and audio streams, returning a base64-encoded MP4 suitable for playback or training data generation.

## Performance Characteristics and Benchmarking

Latency and throughput characteristics differ between surfaces due to their distinct computational modes. The Reasoner's autoregressive nature typically generates tokens sequentially, while the Generator's diffusion process requires multiple denoising steps. For detailed performance tables comparing Reasoner versus Generator latency and throughput metrics, consult [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) in the repository root.

## Summary

* **Reasoner and Generator runtime surfaces** in Cosmos 3 activate different sub-graphs of the same Mixture-of-Transformers model.
* The **Reasoner** uses autoregressive attention to produce text outputs from vision and language inputs, suitable for understanding and reasoning tasks.
* The **Generator** employs diffusion-based denoising to create vision, audio, and action outputs, enabling world generation and simulation.
* Both surfaces share the **mRoPE** embedding scheme and transformer backbone, allowing a single checkpoint to serve both modes.
* Implementation differs by API: **vLLM** for Reasoner text generation, **Diffusers** for Generator synthesis.

## Frequently Asked Questions

### Can the same model checkpoint be used for both Reasoner and Generator tasks?

Yes. According to the implementation in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), a single checkpoint can serve either surface because both runtime surfaces share the same transformer backbone and multimodal attention layers. The choice of runtime (Reasoner or Generator) determines which sub-graph and attention mode (autoregressive or diffusion) is activated during inference.

### What input modalities does each runtime surface support?

The Reasoner surface accepts text and vision inputs (images or videos) and produces text outputs. The Generator surface accepts text, vision, sound, and action inputs, and generates vision, sound, and action outputs including images, videos, audio streams, and action trajectories. This distinction is documented in the repository's comparison table (lines 55-61).

### How do I choose between the Reasoner and Generator surfaces for my application?

Select the **Reasoner** when your task requires understanding, grounding, physical reasoning, or text-based responses to visual queries. Select the **Generator** when you need to synthesize visual content, simulate environments, generate training data, or produce multimodal outputs including audio and action trajectories. The Reasoner excels at captioning and temporal localization, while the Generator handles text-to-image, image-to-video, and audio-synchronized generation.

### Where can I find benchmark data comparing Reasoner and Generator performance?

Performance metrics for both surfaces are documented in [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) located in the repository root. This file contains latency and throughput tables comparing the Reasoner and Generator surfaces across different model configurations, providing guidance for optimizing inference pipelines based on your specific latency and quality requirements.