Cosmos3-Nano vs Cosmos3-Super vs Cosmos3-Super-Text2Image: Complete Model Comparison
Cosmos3-Nano is a 16B parameter omni-modal world model for single-GPU deployment, Cosmos3-Super is a 64B parameter frontier-scale model requiring tensor-parallel inference across multiple GPUs, and Cosmos3-Super-Text2Image is a specialized 64B parameter variant optimized exclusively for high-fidelity text-to-image generation.
The NVIDIA Cosmos repository provides three distinct checkpoints built on the same Mixture-of-Transformers (MoT) architecture, each targeting different trade-offs between scale, capability, and hardware constraints. Understanding the specific differences between Cosmos3-Nano, Cosmos3-Super, and Cosmos3-Super-Text2Image is essential for selecting the right model for your multimodal AI or Physical AI workload.
Model Specifications at a Glance
All three models share the underlying Mixture-of-Transformers (MoT) architecture but differ in parameter count, primary capabilities, and supported modalities. According to the model-family table in README.md (lines 78-84), the specifications break down as follows:
- Cosmos3-Nano: 16B parameters, compact omni-modal world model supporting Generator and Reasoner modes across text, vision, audio, and action modalities.
- Cosmos3-Super: 64B parameters, frontier-scale omni-modal world model with the same full modality support as Nano but at higher fidelity.
- Cosmos3-Super-Text2Image: 64B parameters, specialized generator optimized exclusively for text-to-image synthesis, omitting the Reasoner surface and other modalities.
Architectural and Practical Differences
Scale and Memory Requirements
The Nano checkpoint (16B) fits on a single GPU with modest VRAM (e.g., 24GB) and can be served on a single-GPU vLLM-Omni endpoint on port 8000. In contrast, the Super checkpoint (64B) typically requires tensor-parallel inference across two or more GPUs; the official recipe recommends four GPUs with --tensor-parallel-size 4 for the default server launch, as documented in README.md (lines 312-322).
Capability Breadth
Cosmos3-Nano functions as a general-purpose omni-model capable of generating images, videos, sound, and action trajectories while also serving as a reasoner for text-only outputs. Cosmos3-Super provides identical breadth but delivers higher fidelity and longer context windows due to increased transformer depth. Cosmos3-Super-Text2Image is task-specific: it optimizes the text-to-image diffusion path only, omitting the reasoner surface and multimodal diffusion branches for video, audio, and action to reduce inference overhead.
Performance Characteristics
Benchmarks in inference_benchmarks.md (lines 136-146) demonstrate that Super incurs longer diffusion latency than Nano for identical resolutions due to the larger model size. The Super-Text2Image checkpoint maintains latency comparable to Super generator mode but eliminates the computational overhead of unused video and audio pathways.
Recommended Use Cases
- Cosmos3-Nano: Ideal for rapid prototyping, multimodal reasoning research, single-GPU demos, and memory-constrained environments.
- Cosmos3-Super: Suited for production-grade workloads requiring maximum visual-temporal fidelity, large-scale robotics simulations, and high-resolution multi-modal generation.
- Cosmos3-Super-Text2Image: Optimized for dedicated image-generation services such as art generation and visual content creation where video, audio, and action capabilities are unnecessary.
Loading and Inference Examples
The following Python snippets demonstrate how to load each checkpoint using the Diffusers pipeline. All examples assume a CUDA-enabled environment with PyTorch 2.0 or later.
Cosmos3-Nano (Omni-Modal)
import torch
from diffusers import Cosmos3OmniPipeline
pipe_nano = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
# Text-to-image generation
image = pipe_nano(
prompt="A futuristic city skyline at sunset",
num_frames=1,
height=720,
width=1280,
num_inference_steps=35,
guidance_scale=6.0,
).images[0]
image.save("nano_t2i.png")
Cosmos3-Super (Frontier-Scale)
import torch
from diffusers import Cosmos3OmniPipeline
pipe_super = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Super",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
# Text-to-video generation (30 fps, 189 frames)
video = pipe_super(
prompt="A drone flies over a forest and descends into a canyon",
num_frames=189,
height=720,
width=1280,
fps=30,
num_inference_steps=35,
guidance_scale=7.0,
).video
Cosmos3-Super-Text2Image (Image-Only)
import torch
from diffusers import Cosmos3OmniPipeline
pipe_super_t2i = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Super-Text2Image",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
# High-fidelity image generation
hi_res_image = pipe_super_t2i(
prompt="A photorealistic portrait of an astronaut in a nebula",
num_frames=1,
height=1024,
width=1024,
num_inference_steps=45,
guidance_scale=8.0,
).images[0]
hi_res_image.save("super_t2i.png")
Note: The Super variants require significantly more GPU memory. On single-GPU systems, enable tensor-parallelism or layer-wise offload as described in the vLLM-Omni recipes.
Key Source Files
For authoritative implementation details, reference these files in the NVIDIA Cosmos repository:
README.md(lines 78-84): Model family table comparing sizes and capabilities.inference_benchmarks.md(lines 136-146): Latency benchmarks comparing Nano and Super generator performance.cookbooks/cosmos3/generator/audiovisual/README.md: Practical examples for switching between checkpoints and launching vLLM-Omni endpoints.cookbooks/cosmos3/reasoner/README.md: Configuration details for the Reasoner surface on Nano (port 8000) versus Super (port 8001).
Summary
- Cosmos3-Nano (16B) provides full omni-modal capabilities (Generator + Reasoner) optimized for single-GPU inference with modest VRAM requirements.
- Cosmos3-Super (64B) delivers frontier-scale fidelity across all modalities but requires tensor-parallel inference across multiple GPUs (recommended: 4 GPUs with
--tensor-parallel-size 4). - Cosmos3-Super-Text2Image (64B) offers the same parameter count as Super but specializes exclusively in text-to-image generation, omitting video, audio, action, and reasoning capabilities for streamlined inference.
Frequently Asked Questions
Can Cosmos3-Nano generate video or only images?
Yes, Cosmos3-Nano is a full omni-modal model capable of generating images, videos, audio, and action trajectories. According to the repository documentation, it supports both Generator and Reasoner modes across all modalities including text, vision, audio, and action, making it suitable for world simulation and Physical AI applications beyond static image generation.
Why does Cosmos3-Super require multiple GPUs while Nano runs on a single GPU?
The Cosmos3-Super checkpoint contains 64 billion parameters compared to Nano's 16 billion, exceeding the memory capacity of single consumer GPUs. As implemented in the NVIDIA Cosmos inference recipes, Super requires tensor-parallel distribution across at least two GPUs, with the official documentation recommending four GPUs using --tensor-parallel-size 4 for optimal server deployment.
Is Cosmos3-Super-Text2Image capable of video generation?
No, Cosmos3-Super-Text2Image is exclusively a text-to-image generator. While it shares the same 64B parameter scale as the full Super model, it omits the diffusion branches for video, audio, and action generation, as well as the Reasoner surface. This specialization reduces inference overhead for image-only workloads but prevents it from processing temporal or audio modalities.
Which model should I choose for rapid prototyping on limited hardware?
Choose Cosmos3-Nano for rapid prototyping and single-GPU environments. Its 16B parameter architecture fits comfortably on GPUs with 24GB VRAM, supports the full range of omni-modal capabilities, and can be served via vLLM-Omni on a single GPU without requiring tensor-parallel infrastructure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →