# Difference Between Reasoner and Generator Runtime Surfaces in Cosmos 3

> Discover the key differences between Cosmos 3's Reasoner and Generator runtime surfaces. Understand how Reasoner generates text and Generator creates vision, audio, and action outputs.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: deep-dive
- Published: 2026-06-06

---

**The Reasoner surface produces text via autoregressive decoding for multimodal understanding, while the Generator surface produces vision, audio, and action outputs via diffusion denoising, yet both share the same Mixture-of-Transformers (MoT) backbone in NVIDIA Cosmos 3.**

Cosmos 3 is a unified multimodal foundation model developed by NVIDIA that exposes two distinct runtime surfaces through a single checkpoint. The primary difference between the Reasoner and Generator runtime surfaces in Cosmos 3 lies in their attention modes, accepted inputs, and output modalities. As documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) at lines 55–61, both surfaces are built on top of the same transformer architecture but activate different sub-graphs depending on the task.

## How the Two Surfaces Differ

The Reasoner and Generator surfaces diverge across three dimensions: input types, output types, and internal attention mechanisms.

- **Reasoner** accepts **text and vision** inputs—such as images or videos—and returns **text-only** outputs. Typical applications include world understanding, physical reasoning, task planning, action forecasting, grounded captioning, temporal localization, and embodied-agent decision making.

- **Generator** accepts **text, vision, sound, and action** inputs, and produces **vision, sound, and action** outputs—for example, images, videos, audio streams, and action trajectories. Typical use cases include world generation, world simulation, future prediction, synthetic data generation, policy learning, and robot training.

These distinctions are enumerated in the repository’s primary documentation at [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)【/cache/repos/github.com/NVIDIA/cosmos/main/README.md#L55-L61】.

## Architectural Distinction

Both runtime surfaces share the unified transformer backbone and multimodal attention layers within the MoT architecture. However, they employ fundamentally different forward passes.

### Reasoner: Autoregressive Understanding

The Reasoner surface operates in **autoregressive** mode. Language and visual tokens flow through **causal self-attention**, enabling next-token prediction optimized for understanding tasks. When you query the model with an image or video and a text prompt, the Reasoner decodes a textual response by attending only to previous tokens in the sequence.

This mode powers capabilities such as 2-D grounding, detailed captioning, and temporal localization. Because the surface restricts outputs to text, it functions as a multimodal large language model that reasons about visual content without generating new pixel or audio data.

### Generator: Diffusion-Based Creation

The Generator surface operates in **diffusion** mode. It accepts noisy multimodal tokens representing image, video, audio, or action modalities, then denoises them through **full-attention diffusion steps**. Unlike the Reasoner’s sequential decoding, the Generator iteratively refines a complete latent representation to produce high-fidelity generative outputs.

This surface handles text-to-image, text-to-video, image-to-video, audio-synchronized video, and action-rollout generation. The Diffusers pipeline `Cosmos3OmniPipeline` and the vLLM-Omni server both invoke this Generator sub-graph when executing creation tasks.

### Shared Components

Despite their different operational modes, both surfaces rely on identical core infrastructure:

- The same **Mixture-of-Transformers (MoT)** backbone and multimodal attention layers.
- The unified **3-D multi-dimensional rotary position embedding (mRoPE)** that encodes spatial-temporal structure across all modalities.

Consequently, a single model checkpoint can serve either surface; the runtime selection simply determines which sub-graph of the shared model receives activation.

## Practical Code Examples

The NVIDIA Cosmos repository provides cookbook notebooks and API wrappers that demonstrate each surface in action.

### Reasoner: Text Output from an Image via vLLM

The following example targets the Reasoner surface through a local vLLM API server. It encodes an image as base64, constructs a Qwen-3-VL-compatible message payload, and receives a text-only response.

```python
import requests, json, base64, pathlib

# Load an image (or video) and encode it as base64

with open("example.jpg", "rb") as f:
    img_b64 = base64.b64encode(f.read()).decode()

payload = {
    "model": "cosmos3-reasoner-nano",
    "messages": [
        {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
        {"role": "user", "content": [
            {"type": "image", "image": img_b64},
            {"type": "text", "text": "Describe the scene in detail."}
        ]}
    ],
    "max_tokens": 256,
    "temperature": 0.6,
    "top_p": 0.95
}

resp = requests.post("http://localhost:8000/v1/chat/completions", json=payload)
print(resp.json()["choices"][0]["message"]["content"])

```

The payload schema follows the Qwen-3-VL message format described in the README. Because the request targets the Reasoner, `resp` contains a **text** string suitable for captioning or reasoning tasks.

### Generator: Text-to-Image with Diffusers

This example invokes the Generator surface using the `Cosmos3OmniPipeline` from Hugging Face Diffusers. The `variant="diffusion"` flag activates the diffusion sub-graph for vision generation.

```python
from diffusers import Cosmos3OmniPipeline
import torch

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    variant="diffusion"
).to("cuda")

prompt = "A futuristic laboratory with robots assembling microchips"
image = pipe(prompt, height=720, width=1280, max_new_tokens=1024).images[0]
image.save("generated.png")

```

Here, `pipe(...)` denoises image tokens to emit a **vision** artifact. The Generator surface processes the text prompt through full-attention diffusion rather than autoregressive causal attention.

### Generator: Video and Audio via vLLM-Omni

The Generator surface can also emit multimodal streams such as synchronized video and audio. The following request targets a vLLM-Omni server with `audio=True` and a high `max_new_tokens` budget suitable for long video generation.

```python
payload = {
    "model": "cosmos3-super",
    "messages": [
        {"role": "user", "content": [
            {"type": "text", "text": "Create a short video of a robot pouring water, with realistic sound."}
        ]}
    ],
    "max_tokens": 20000,
    "height": 720,
    "width": 1280,
    "num_frames": 60,
    "audio": True
}
resp = requests.post("http://localhost:8001/v1/chat/completions", json=payload)
video_bytes = resp.json()["choices"][0]["message"]["content"]["video"]
with open("robot_water.mp4", "wb") as f:
    f.write(base64.b64decode(video_bytes))

```

The server returns a base64-encoded MP4 because the Generator surface is permitted to produce **vision and audio** outputs. This is impossible on the Reasoner surface, which is strictly text-only.

## Key Files in the NVIDIA Cosmos Repository

Several files in the `NVIDIA/cosmos` repository document and implement these two surfaces:

- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** — Defines the dual-surface architecture, input/output matrices, and the shared MoT backbone at lines 55–61.
- **[`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md)** — Provides latency and throughput tables comparing Reasoner against Generator inference.
- **`cookbooks/cosmos3/reasoner/run_with_vllm.ipynb`** — End-to-end notebook demonstrating Reasoner inference via vLLM.
- **`cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb`** — Practical guide to Generator inference using the Diffusers pipeline for audiovisual output.

## Summary

- The **Reasoner** runtime surface in Cosmos 3 runs in **autoregressive** mode, accepts text and vision inputs, and produces **text-only** outputs for understanding and reasoning tasks.
- The **Generator** runtime surface runs in **diffusion** mode, accepts text, vision, sound, and action inputs, and produces **vision, sound, and action** outputs for world generation and simulation.
- Both surfaces share the same **MoT backbone**, multimodal attention layers, and **3-D mRoPE** position embeddings; a single checkpoint activates either sub-graph at runtime.
- The **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** and dedicated cookbooks provide canonical examples for each surface via vLLM, vLLM-Omni, and Diffusers APIs.

## Frequently Asked Questions

### What inputs can the Cosmos 3 Reasoner surface accept?

The Reasoner surface accepts **text and vision** inputs, including images and videos. It does not process audio or action tokens; those modalities are reserved for the Generator surface.

### Can the Generator surface produce text outputs like the Reasoner?

No. The Generator surface is optimized for **generative** outputs such as images, videos, audio streams, and action trajectories. It operates via diffusion rather than causal autoregressive text decoding, so it does not return text answers.

### Is a separate model checkpoint required for each surface?

No. Both the Reasoner and Generator runtime surfaces share the **same underlying checkpoint** and transformer weights. The runtime API or pipeline selection—such as vLLM for Reasoner or Diffusers for Generator—determines which attention mode and sub-graph are activated.

### Where can I find performance benchmarks comparing Reasoner and Generator inference?

Performance tables for latency and throughput are documented in **[`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md)** within the `NVIDIA/cosmos` repository. This file provides side-by-side metrics for both runtime surfaces across different hardware configurations.