Reasoner vs Generator Runtime Surfaces in NVIDIA Cosmos: Architecture and Usage
The Reasoner and Generator runtime surfaces in Cosmos 3 share the same Mixture-of-Transformers (MoT) backbone but operate through fundamentally different attention mechanisms: the Reasoner employs autoregressive processing to generate text outputs from vision and language inputs, while the Generator utilizes diffusion-based denoising to produce vision, audio, and action outputs from multimodal prompts.
NVIDIA Cosmos exposes two distinct runtime surfaces through a unified architecture that enables both multimodal understanding and world generation capabilities. According to the source code in the NVIDIA/cosmos repository, both surfaces utilize identical transformer backbones and multi-dimensional rotary position embeddings (mRoPE), yet they activate different computational sub-graphs depending on whether the task requires reasoning or synthesis.
Input and Output Differences Between Surfaces
The primary distinction between these runtime surfaces lies in their accepted input modalities and resulting output types. As documented in README.md (lines 55-61), the surfaces differ as follows:
| Surface | Accepted Inputs | Produced Outputs | Typical Use Cases |
|---|---|---|---|
| Reasoner | Text + vision (image or video) | Text | World understanding, grounding, physical reasoning, task planning, action forecasting, embodied-agent reasoning, autonomous-system decision making |
| Generator | Text + vision + sound + action | Vision, sound, action (e.g., images, videos, audio streams, action trajectories) | World generation, world simulation, future prediction, synthetic data generation, policy learning, robot training |
Architectural Distinctions: Autoregressive vs Diffusion
While both surfaces share underlying infrastructure, they implement divergent attention mechanisms optimized for their specific tasks.
Reasoner Surface – Autoregressive Mode
The Reasoner operates in autoregressive mode, processing language and visual tokens through causal self-attention mechanisms. This architecture enables next-token prediction optimized for understanding tasks such as captioning, temporal localization, and 2-D grounding. When you invoke the Reasoner surface, the model processes input sequences sequentially, attending only to previous tokens to generate coherent textual responses.
Generator Surface – Diffusion Mode
The Generator surface operates in diffusion mode, utilizing full-attention diffusion steps to denoise multimodal tokens. Unlike the causal approach of the Reasoner, the Generator processes noisy representations of image, video, audio, and action tokens simultaneously through the complete attention layers. This enables the synthesis of high-fidelity generative outputs including text-to-image/video conversion, image-to-video generation, audio-synchronized video creation, and action trajectory rollouts.
Shared Components and Unified Backbone
Despite their operational differences, both runtime surfaces leverage identical core components:
- The same transformer backbone – Both surfaces utilize the shared MoT architecture without requiring separate model weights.
- Unified multimodal attention layers – The attention mechanisms are repurposed depending on the runtime mode (causal masking for Reasoner, full attention for Generator).
- 3-D multi-dimensional rotary position embedding (mRoPE) – This encodes spatial-temporal structure consistently across all modalities, ensuring coherent handling of vision, audio, and action data.
- Single checkpoint compatibility – A single model checkpoint can serve either surface; the choice of runtime determines which sub-graph of the model is activated.
Practical Implementation Examples
The NVIDIA/cosmos repository provides specific implementations demonstrating how to invoke each surface through different APIs.
Reasoner Inference with vLLM
To generate text outputs from visual inputs using the Reasoner surface, implement the vLLM API following the schema described in the repository. The payload follows the Qwen-3-VL message format, returning text strings because the Reasoner surface produces only textual answers:
import requests, json, base64, pathlib
# Load an image (or video) and encode it as base64
with open("example.jpg", "rb") as f:
img_b64 = base64.b64encode(f.read()).decode()
payload = {
"model": "cosmos3-reasoner-nano",
"messages": [
{"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
{"role": "user", "content": [
{"type": "image", "image": img_b64},
{"type": "text", "text": "Describe the scene in detail."}
]}
],
"max_tokens": 256,
"temperature": 0.6,
"top_p": 0.95
}
resp = requests.post("http://localhost:8000/v1/chat/completions", json=payload)
print(resp.json()["choices"][0]["message"]["content"])
For a complete end-to-end implementation, reference cookbooks/cosmos3/reasoner/run_with_vllm.ipynb in the repository.
Text-to-Image Generation with Diffusers
To invoke the Generator surface for vision synthesis, use the Diffusers pipeline which activates the diffusion sub-graph via the wrapper implemented in the source:
from diffusers import Cosmos3OmniPipeline
import torch
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.bfloat16,
variant="diffusion"
).to("cuda")
prompt = "A futuristic laboratory with robots assembling microchips"
image = pipe(prompt, height=720, width=1280, max_new_tokens=1024).images[0]
image.save("generated.png")
This implementation denoises image tokens to output vision artifacts (PNG files) rather than text. See cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb for comprehensive Generator usage patterns.
Multimodal Video and Audio Generation
The Generator surface supports multi-stream output when the request includes audio parameters. The following example demonstrates video generation with synchronized sound:
payload = {
"model": "cosmos3-super",
"messages": [
{"role": "user", "content": [
{"type": "text", "text": "Create a short video of a robot pouring water, with realistic sound."}
]}
],
"max_tokens": 20000,
"height": 720,
"width": 1280,
"num_frames": 60,
"audio": True
}
resp = requests.post("http://localhost:8001/v1/chat/completions", json=payload)
video_bytes = resp.json()["choices"][0]["message"]["content"]["video"]
with open("robot_water.mp4", "wb") as f:
f.write(base64.b64decode(video_bytes))
When audio=True is specified, the Generator emits video and audio streams, returning a base64-encoded MP4 suitable for playback or training data generation.
Performance Characteristics and Benchmarking
Latency and throughput characteristics differ between surfaces due to their distinct computational modes. The Reasoner's autoregressive nature typically generates tokens sequentially, while the Generator's diffusion process requires multiple denoising steps. For detailed performance tables comparing Reasoner versus Generator latency and throughput metrics, consult inference_benchmarks.md in the repository root.
Summary
- Reasoner and Generator runtime surfaces in Cosmos 3 activate different sub-graphs of the same Mixture-of-Transformers model.
- The Reasoner uses autoregressive attention to produce text outputs from vision and language inputs, suitable for understanding and reasoning tasks.
- The Generator employs diffusion-based denoising to create vision, audio, and action outputs, enabling world generation and simulation.
- Both surfaces share the mRoPE embedding scheme and transformer backbone, allowing a single checkpoint to serve both modes.
- Implementation differs by API: vLLM for Reasoner text generation, Diffusers for Generator synthesis.
Frequently Asked Questions
Can the same model checkpoint be used for both Reasoner and Generator tasks?
Yes. According to the implementation in README.md, a single checkpoint can serve either surface because both runtime surfaces share the same transformer backbone and multimodal attention layers. The choice of runtime (Reasoner or Generator) determines which sub-graph and attention mode (autoregressive or diffusion) is activated during inference.
What input modalities does each runtime surface support?
The Reasoner surface accepts text and vision inputs (images or videos) and produces text outputs. The Generator surface accepts text, vision, sound, and action inputs, and generates vision, sound, and action outputs including images, videos, audio streams, and action trajectories. This distinction is documented in the repository's comparison table (lines 55-61).
How do I choose between the Reasoner and Generator surfaces for my application?
Select the Reasoner when your task requires understanding, grounding, physical reasoning, or text-based responses to visual queries. Select the Generator when you need to synthesize visual content, simulate environments, generate training data, or produce multimodal outputs including audio and action trajectories. The Reasoner excels at captioning and temporal localization, while the Generator handles text-to-image, image-to-video, and audio-synchronized generation.
Where can I find benchmark data comparing Reasoner and Generator performance?
Performance metrics for both surfaces are documented in inference_benchmarks.md located in the repository root. This file contains latency and throughput tables comparing the Reasoner and Generator surfaces across different model configurations, providing guidance for optimizing inference pipelines based on your specific latency and quality requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →