Difference Between Reasoner and Generator Runtime Surfaces in Cosmos 3
The Reasoner surface produces text via autoregressive decoding for multimodal understanding, while the Generator surface produces vision, audio, and action outputs via diffusion denoising, yet both share the same Mixture-of-Transformers (MoT) backbone in NVIDIA Cosmos 3.
Cosmos 3 is a unified multimodal foundation model developed by NVIDIA that exposes two distinct runtime surfaces through a single checkpoint. The primary difference between the Reasoner and Generator runtime surfaces in Cosmos 3 lies in their attention modes, accepted inputs, and output modalities. As documented in README.md at lines 55–61, both surfaces are built on top of the same transformer architecture but activate different sub-graphs depending on the task.
How the Two Surfaces Differ
The Reasoner and Generator surfaces diverge across three dimensions: input types, output types, and internal attention mechanisms.
-
Reasoner accepts text and vision inputs—such as images or videos—and returns text-only outputs. Typical applications include world understanding, physical reasoning, task planning, action forecasting, grounded captioning, temporal localization, and embodied-agent decision making.
-
Generator accepts text, vision, sound, and action inputs, and produces vision, sound, and action outputs—for example, images, videos, audio streams, and action trajectories. Typical use cases include world generation, world simulation, future prediction, synthetic data generation, policy learning, and robot training.
These distinctions are enumerated in the repository’s primary documentation at README.md【/cache/repos/github.com/NVIDIA/cosmos/main/README.md#L55-L61】.
Architectural Distinction
Both runtime surfaces share the unified transformer backbone and multimodal attention layers within the MoT architecture. However, they employ fundamentally different forward passes.
Reasoner: Autoregressive Understanding
The Reasoner surface operates in autoregressive mode. Language and visual tokens flow through causal self-attention, enabling next-token prediction optimized for understanding tasks. When you query the model with an image or video and a text prompt, the Reasoner decodes a textual response by attending only to previous tokens in the sequence.
This mode powers capabilities such as 2-D grounding, detailed captioning, and temporal localization. Because the surface restricts outputs to text, it functions as a multimodal large language model that reasons about visual content without generating new pixel or audio data.
Generator: Diffusion-Based Creation
The Generator surface operates in diffusion mode. It accepts noisy multimodal tokens representing image, video, audio, or action modalities, then denoises them through full-attention diffusion steps. Unlike the Reasoner’s sequential decoding, the Generator iteratively refines a complete latent representation to produce high-fidelity generative outputs.
This surface handles text-to-image, text-to-video, image-to-video, audio-synchronized video, and action-rollout generation. The Diffusers pipeline Cosmos3OmniPipeline and the vLLM-Omni server both invoke this Generator sub-graph when executing creation tasks.
Shared Components
Despite their different operational modes, both surfaces rely on identical core infrastructure:
- The same Mixture-of-Transformers (MoT) backbone and multimodal attention layers.
- The unified 3-D multi-dimensional rotary position embedding (mRoPE) that encodes spatial-temporal structure across all modalities.
Consequently, a single model checkpoint can serve either surface; the runtime selection simply determines which sub-graph of the shared model receives activation.
Practical Code Examples
The NVIDIA Cosmos repository provides cookbook notebooks and API wrappers that demonstrate each surface in action.
Reasoner: Text Output from an Image via vLLM
The following example targets the Reasoner surface through a local vLLM API server. It encodes an image as base64, constructs a Qwen-3-VL-compatible message payload, and receives a text-only response.
import requests, json, base64, pathlib
# Load an image (or video) and encode it as base64
with open("example.jpg", "rb") as f:
img_b64 = base64.b64encode(f.read()).decode()
payload = {
"model": "cosmos3-reasoner-nano",
"messages": [
{"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
{"role": "user", "content": [
{"type": "image", "image": img_b64},
{"type": "text", "text": "Describe the scene in detail."}
]}
],
"max_tokens": 256,
"temperature": 0.6,
"top_p": 0.95
}
resp = requests.post("http://localhost:8000/v1/chat/completions", json=payload)
print(resp.json()["choices"][0]["message"]["content"])
The payload schema follows the Qwen-3-VL message format described in the README. Because the request targets the Reasoner, resp contains a text string suitable for captioning or reasoning tasks.
Generator: Text-to-Image with Diffusers
This example invokes the Generator surface using the Cosmos3OmniPipeline from Hugging Face Diffusers. The variant="diffusion" flag activates the diffusion sub-graph for vision generation.
from diffusers import Cosmos3OmniPipeline
import torch
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.bfloat16,
variant="diffusion"
).to("cuda")
prompt = "A futuristic laboratory with robots assembling microchips"
image = pipe(prompt, height=720, width=1280, max_new_tokens=1024).images[0]
image.save("generated.png")
Here, pipe(...) denoises image tokens to emit a vision artifact. The Generator surface processes the text prompt through full-attention diffusion rather than autoregressive causal attention.
Generator: Video and Audio via vLLM-Omni
The Generator surface can also emit multimodal streams such as synchronized video and audio. The following request targets a vLLM-Omni server with audio=True and a high max_new_tokens budget suitable for long video generation.
payload = {
"model": "cosmos3-super",
"messages": [
{"role": "user", "content": [
{"type": "text", "text": "Create a short video of a robot pouring water, with realistic sound."}
]}
],
"max_tokens": 20000,
"height": 720,
"width": 1280,
"num_frames": 60,
"audio": True
}
resp = requests.post("http://localhost:8001/v1/chat/completions", json=payload)
video_bytes = resp.json()["choices"][0]["message"]["content"]["video"]
with open("robot_water.mp4", "wb") as f:
f.write(base64.b64decode(video_bytes))
The server returns a base64-encoded MP4 because the Generator surface is permitted to produce vision and audio outputs. This is impossible on the Reasoner surface, which is strictly text-only.
Key Files in the NVIDIA Cosmos Repository
Several files in the NVIDIA/cosmos repository document and implement these two surfaces:
README.md— Defines the dual-surface architecture, input/output matrices, and the shared MoT backbone at lines 55–61.inference_benchmarks.md— Provides latency and throughput tables comparing Reasoner against Generator inference.cookbooks/cosmos3/reasoner/run_with_vllm.ipynb— End-to-end notebook demonstrating Reasoner inference via vLLM.cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb— Practical guide to Generator inference using the Diffusers pipeline for audiovisual output.
Summary
- The Reasoner runtime surface in Cosmos 3 runs in autoregressive mode, accepts text and vision inputs, and produces text-only outputs for understanding and reasoning tasks.
- The Generator runtime surface runs in diffusion mode, accepts text, vision, sound, and action inputs, and produces vision, sound, and action outputs for world generation and simulation.
- Both surfaces share the same MoT backbone, multimodal attention layers, and 3-D mRoPE position embeddings; a single checkpoint activates either sub-graph at runtime.
- The
README.mdand dedicated cookbooks provide canonical examples for each surface via vLLM, vLLM-Omni, and Diffusers APIs.
Frequently Asked Questions
What inputs can the Cosmos 3 Reasoner surface accept?
The Reasoner surface accepts text and vision inputs, including images and videos. It does not process audio or action tokens; those modalities are reserved for the Generator surface.
Can the Generator surface produce text outputs like the Reasoner?
No. The Generator surface is optimized for generative outputs such as images, videos, audio streams, and action trajectories. It operates via diffusion rather than causal autoregressive text decoding, so it does not return text answers.
Is a separate model checkpoint required for each surface?
No. Both the Reasoner and Generator runtime surfaces share the same underlying checkpoint and transformer weights. The runtime API or pipeline selection—such as vLLM for Reasoner or Diffusers for Generator—determines which attention mode and sub-graph are activated.
Where can I find performance benchmarks comparing Reasoner and Generator inference?
Performance tables for latency and throughput are documented in inference_benchmarks.md within the NVIDIA/cosmos repository. This file provides side-by-side metrics for both runtime surfaces across different hardware configurations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →