# Cosmos3-Nano vs Cosmos3-Super vs Cosmos3-Super-Text2Image: Complete Model Comparison

> Compare Cosmos3-Nano, Cosmos3-Super, and Cosmos3-Super-Text2Image models from NVIDIA/cosmos. Understand parameter counts, GPU requirements, and use cases to choose the best fit for your project.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: deep-dive
- Published: 2026-06-14

---

**Cosmos3-Nano is a 16B parameter omni-modal world model for single-GPU deployment, Cosmos3-Super is a 64B parameter frontier-scale model requiring tensor-parallel inference across multiple GPUs, and Cosmos3-Super-Text2Image is a specialized 64B parameter variant optimized exclusively for high-fidelity text-to-image generation.**

The NVIDIA Cosmos repository provides three distinct checkpoints built on the same **Mixture-of-Transformers (MoT)** architecture, each targeting different trade-offs between scale, capability, and hardware constraints. Understanding the specific differences between **Cosmos3-Nano**, **Cosmos3-Super**, and **Cosmos3-Super-Text2Image** is essential for selecting the right model for your multimodal AI or Physical AI workload.

## Model Specifications at a Glance

All three models share the underlying **Mixture-of-Transformers (MoT)** architecture but differ in parameter count, primary capabilities, and supported modalities. According to the model-family table in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 78-84), the specifications break down as follows:

- **Cosmos3-Nano**: 16B parameters, compact omni-modal world model supporting Generator and Reasoner modes across text, vision, audio, and action modalities.
- **Cosmos3-Super**: 64B parameters, frontier-scale omni-modal world model with the same full modality support as Nano but at higher fidelity.
- **Cosmos3-Super-Text2Image**: 64B parameters, specialized generator optimized exclusively for text-to-image synthesis, omitting the Reasoner surface and other modalities.

## Architectural and Practical Differences

### Scale and Memory Requirements

The **Nano** checkpoint (16B) fits on a single GPU with modest VRAM (e.g., 24GB) and can be served on a single-GPU vLLM-Omni endpoint on port 8000. In contrast, the **Super** checkpoint (64B) typically requires tensor-parallel inference across two or more GPUs; the official recipe recommends four GPUs with `--tensor-parallel-size 4` for the default server launch, as documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 312-322).

### Capability Breadth

**Cosmos3-Nano** functions as a general-purpose omni-model capable of generating images, videos, sound, and action trajectories while also serving as a reasoner for text-only outputs. **Cosmos3-Super** provides identical breadth but delivers higher fidelity and longer context windows due to increased transformer depth. **Cosmos3-Super-Text2Image** is task-specific: it optimizes the text-to-image diffusion path only, omitting the reasoner surface and multimodal diffusion branches for video, audio, and action to reduce inference overhead.

### Performance Characteristics

Benchmarks in [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) (lines 136-146) demonstrate that **Super** incurs longer diffusion latency than **Nano** for identical resolutions due to the larger model size. The **Super-Text2Image** checkpoint maintains latency comparable to **Super** generator mode but eliminates the computational overhead of unused video and audio pathways.

### Recommended Use Cases

- **Cosmos3-Nano**: Ideal for rapid prototyping, multimodal reasoning research, single-GPU demos, and memory-constrained environments.
- **Cosmos3-Super**: Suited for production-grade workloads requiring maximum visual-temporal fidelity, large-scale robotics simulations, and high-resolution multi-modal generation.
- **Cosmos3-Super-Text2Image**: Optimized for dedicated image-generation services such as art generation and visual content creation where video, audio, and action capabilities are unnecessary.

## Loading and Inference Examples

The following Python snippets demonstrate how to load each checkpoint using the Diffusers pipeline. All examples assume a CUDA-enabled environment with PyTorch 2.0 or later.

### Cosmos3-Nano (Omni-Modal)

```python
import torch
from diffusers import Cosmos3OmniPipeline

pipe_nano = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

# Text-to-image generation

image = pipe_nano(
    prompt="A futuristic city skyline at sunset",
    num_frames=1,
    height=720,
    width=1280,
    num_inference_steps=35,
    guidance_scale=6.0,
).images[0]
image.save("nano_t2i.png")

```

### Cosmos3-Super (Frontier-Scale)

```python
import torch
from diffusers import Cosmos3OmniPipeline

pipe_super = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Super",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

# Text-to-video generation (30 fps, 189 frames)

video = pipe_super(
    prompt="A drone flies over a forest and descends into a canyon",
    num_frames=189,
    height=720,
    width=1280,
    fps=30,
    num_inference_steps=35,
    guidance_scale=7.0,
).video

```

### Cosmos3-Super-Text2Image (Image-Only)

```python
import torch
from diffusers import Cosmos3OmniPipeline

pipe_super_t2i = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Super-Text2Image",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

# High-fidelity image generation

hi_res_image = pipe_super_t2i(
    prompt="A photorealistic portrait of an astronaut in a nebula",
    num_frames=1,
    height=1024,
    width=1024,
    num_inference_steps=45,
    guidance_scale=8.0,
).images[0]
hi_res_image.save("super_t2i.png")

```

**Note**: The **Super** variants require significantly more GPU memory. On single-GPU systems, enable tensor-parallelism or layer-wise offload as described in the vLLM-Omni recipes.

## Key Source Files

For authoritative implementation details, reference these files in the NVIDIA Cosmos repository:

- [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 78-84): Model family table comparing sizes and capabilities.
- [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) (lines 136-146): Latency benchmarks comparing Nano and Super generator performance.
- [`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md): Practical examples for switching between checkpoints and launching vLLM-Omni endpoints.
- [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md): Configuration details for the Reasoner surface on Nano (port 8000) versus Super (port 8001).

## Summary

- **Cosmos3-Nano** (16B) provides full omni-modal capabilities (Generator + Reasoner) optimized for single-GPU inference with modest VRAM requirements.
- **Cosmos3-Super** (64B) delivers frontier-scale fidelity across all modalities but requires tensor-parallel inference across multiple GPUs (recommended: 4 GPUs with `--tensor-parallel-size 4`).
- **Cosmos3-Super-Text2Image** (64B) offers the same parameter count as Super but specializes exclusively in text-to-image generation, omitting video, audio, action, and reasoning capabilities for streamlined inference.

## Frequently Asked Questions

### Can Cosmos3-Nano generate video or only images?

Yes, **Cosmos3-Nano** is a full omni-modal model capable of generating images, videos, audio, and action trajectories. According to the repository documentation, it supports both Generator and Reasoner modes across all modalities including text, vision, audio, and action, making it suitable for world simulation and Physical AI applications beyond static image generation.

### Why does Cosmos3-Super require multiple GPUs while Nano runs on a single GPU?

The **Cosmos3-Super** checkpoint contains 64 billion parameters compared to Nano's 16 billion, exceeding the memory capacity of single consumer GPUs. As implemented in the NVIDIA Cosmos inference recipes, Super requires tensor-parallel distribution across at least two GPUs, with the official documentation recommending four GPUs using `--tensor-parallel-size 4` for optimal server deployment.

### Is Cosmos3-Super-Text2Image capable of video generation?

No, **Cosmos3-Super-Text2Image** is exclusively a text-to-image generator. While it shares the same 64B parameter scale as the full Super model, it omits the diffusion branches for video, audio, and action generation, as well as the Reasoner surface. This specialization reduces inference overhead for image-only workloads but prevents it from processing temporal or audio modalities.

### Which model should I choose for rapid prototyping on limited hardware?

Choose **Cosmos3-Nano** for rapid prototyping and single-GPU environments. Its 16B parameter architecture fits comfortably on GPUs with 24GB VRAM, supports the full range of omni-modal capabilities, and can be served via vLLM-Omni on a single GPU without requiring tensor-parallel infrastructure.