Cosmos 3 vs Cosmos 3 Super: Choosing the Right Model Size for Your Use Case
Cosmos 3 Nano (16B parameters) runs on a single GPU for rapid prototyping, while Cosmos 3 Super (64B parameters) requires four GPUs but delivers production-grade visual fidelity for high-resolution video generation.
The NVIDIA Cosmos repository provides two checkpoints built on the identical Mixture-of-Transformers (MoT) architecture, differing only in parameter count and compute requirements. Whether you are testing pipelines or deploying world-model generation at scale, understanding the trade-offs between Cosmos 3 Nano and Cosmos 3 Super ensures you allocate the right hardware for your specific multimodal or reasoning workload.
Architecture Overview: Why the Same Codebase Powers Both Models
Both Cosmos 3 Nano and Cosmos 3 Super share the exact same model architecture implemented in the NVIDIA Cosmos codebase. According to the architecture diagram in cookbooks/cosmos3/cosmos3-model-architecture.png, the MoT design combines an autoregressive transformer for reasoning with a diffusion transformer for generation, unified by 3-D rotary positional embeddings (mRoPE) that encode spatial-temporal structure across modalities.
Because the core graph is identical, you can swap between Nano and Super checkpoints without changing application code. The only adjustments required are resource allocation—specifically GPU count, tensor-parallelism settings, and memory offloading strategies.
Cosmos 3 Nano vs Cosmos 3 Super: Key Differences
| Aspect | Cosmos 3 Nano | Cosmos 3 Super |
|---|---|---|
| Parameters | 16 billion | 64 billion |
| Typical Hardware | 1 GPU (e.g., RTX 3090) | 4 GPUs with tensor-parallel size = 4 |
| Inference Latency | Faster per-frame generation | 2–3× higher latency than Nano |
| Memory Footprint | Fits in single GPU VRAM | Requires layer-wise offload or multi-GPU distribution |
| Visual Fidelity | Good for 256p–480p outputs | Highest fidelity at 720p, better temporal coherence |
| Recommended Use | Prototyping, low-latency endpoints | Production content, physical AI tasks |
The benchmark tables in inference_benchmarks.md quantify these differences, showing that while Nano prioritizes speed, Super excels at nuanced action reasoning and high-resolution generation.
When to Use Cosmos 3 Nano
Choose Cosmos 3 Nano when you need rapid iteration and have limited GPU memory. The 16B parameter model loads efficiently on single-GPU setups using Diffusers or vLLM-Omni without tensor parallelism.
Ideal scenarios include:
- Early-stage research and development
- Applications where response time matters more than absolute image quality
- Text-captioning endpoints or low-resolution video previews (480p and below)
- Development environments with consumer-grade hardware
When to Use Cosmos 3 Super
Select Cosmos 3 Super when visual realism and physical plausibility are critical. The 64B parameter checkpoint enables richer world-model representations necessary for complex action-policy and forward-dynamics tasks.
Deploy Super for:
- Production-grade 720p video generation with sound synchronization
- Fine-grained robot-policy generation requiring precise physical reasoning
- High-fidelity image-to-video tasks where temporal coherence is essential
- Evaluations requiring top Physics IQ scores (documented in
evaluation/cosmos3/Physics_IQ/README.md)
Deployment Examples
Loading Checkpoints with Diffusers
The Cosmos3OmniPipeline class loads either checkpoint by changing the pretrained model name. Both use the same tokenizer and scheduler configuration:
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
# Switch between Nano and Super by changing the checkpoint string
checkpoint = "nvidia/Cosmos3-Super" # or "nvidia/Cosmos3-Nano"
pipe = Cosmos3OmniPipeline.from_pretrained(
checkpoint,
torch_dtype=torch.bfloat16,
device_map="cuda",
)
pipe.scheduler = UniPCMultistepScheduler.from_config(
pipe.scheduler.config,
flow_shift=10.0
)
result = pipe(
prompt="A futuristic city skyline at sunset.",
num_frames=189,
height=720,
width=1280,
fps=24,
num_inference_steps=35,
guidance_scale=6.0,
generator=torch.Generator(device="cuda").manual_seed(1234),
)
The only difference between deployments is the checkpoint path; all other initialization code remains identical as documented in cookbooks/cosmos3/generator/audiovisual/README.md.
Serving with vLLM-Omni
For Super deployment across multiple GPUs, use tensor parallelism with layer-wise offloading:
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-v "$(pwd):/workspace" -p 8000:8000 --ipc=host \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Super \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--port 8000 \
--init-timeout 1800
Client requests use the standard OpenAI-compatible API:
import requests
payload = {
"prompt": "A robot arm assembles a wooden chair.",
"size": "1280x720",
"num_frames": 189,
"fps": 24,
"num_inference_steps": 35,
"guidance_scale": 6.0,
"flow_shift": 10.0,
"seed": 0
}
resp = requests.post("http://localhost:8000/v1/videos/sync", json=payload)
Running the NIM Container
For text-only reasoning with Super, deploy the NIM container with the model size environment variable:
docker run -it --rm --gpus all \
-e NGC_API_KEY=$NGC_API_KEY \
-e NIM_MODEL_SIZE=super \
-p 8000:8000 nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
response = client.chat.completions.create(
model="nvidia/cosmos3-super-reasoner",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [
{"type": "video_url", "video_url": {"url": "https://example.com/demo.mp4"}},
{"type": "text", "text": "Explain the robot's next action."}
]}
],
max_tokens=256,
)
Performance Trade-offs
Latency: The Super checkpoint incurs roughly 2–3× higher per-frame latency compared to Nano on equivalent hardware, as documented in inference_benchmarks.md. This reflects the increased attention heads and feed-forward layers processed per token in the 64B model.
Memory: A single Super checkpoint exceeds the capacity of modern single-GPU configurations. The official launch scripts in the repository recommend four GPUs with --tensor-parallel-size 4 or enabling --enable-layerwise-offload to distribute weights across GPU and CPU memory.
Quality: Qualitative evaluations in evaluation/cosmos3/Physics_IQ/README.md consistently favor Super for high-resolution image-to-video tasks and complex action predictions, while Nano provides sufficient quality for proof-of-concept development.
Summary
- Cosmos 3 Nano (16B) and Cosmos 3 Super (64B) share identical MoT architecture and tokenizers, differing only in parameter count and compute requirements.
- Use Nano for single-GPU prototyping, low-latency endpoints, and 480p or lower resolution outputs.
- Use Super for multi-GPU production deployments requiring 720p video, sound synchronization, or advanced physical reasoning.
- Switch between models by changing the checkpoint path in
Cosmos3OmniPipeline.from_pretrained()or the--tensor-parallel-sizeflag in vLLM-Omni—no code refactoring required. - Reference
inference_benchmarks.mdfor detailed latency metrics andcookbooks/cosmos3/reasoner/README.mdfor scaling configurations.
Frequently Asked Questions
What is the parameter difference between Cosmos 3 Nano and Super?
Cosmos 3 Nano contains 16 billion parameters, while Cosmos 3 Super contains 64 billion parameters. This 4× increase in model size enables the Super variant to learn richer latent representations for complex multimodal tasks, but requires significantly more GPU memory and compute resources during inference.
Can I run Cosmos 3 Super on a single GPU?
No, a single Super checkpoint exceeds the memory capacity of individual modern GPUs. According to the deployment guides in cookbooks/cosmos3/reasoner/README.md, Super requires four GPUs with tensor parallelism (--tensor-parallel-size 4) or layer-wise offloading (--enable-layerwise-offload) to distribute the model weights across devices and CPU memory.
How do I switch between Nano and Super in my code?
You switch models by changing the checkpoint identifier string. In Diffusers, update the from_pretrained() call to use either "nvidia/Cosmos3-Nano" or "nvidia/Cosmos3-Super". For vLLM-Omni, change the model argument in the serve command. The architecture, tokenizer, and API endpoints remain identical, so no other code changes are necessary.
Does Cosmos 3 Super produce better quality than Nano?
Yes. The 64B Super model delivers higher visual fidelity, better temporal coherence in video generation, and superior performance on physical reasoning benchmarks (Physics IQ) compared to the 16B Nano model. However, for rapid prototyping or applications where latency matters more than absolute quality, Nano provides sufficient performance with significantly faster inference speeds.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →