Cosmos 3 vs Cosmos 3 Super: Choosing the Right Model Size for Your Use Case

Cosmos 3 Nano (16B parameters) runs on a single GPU for rapid prototyping, while Cosmos 3 Super (64B parameters) requires four GPUs but delivers production-grade visual fidelity for high-resolution video generation.

The NVIDIA Cosmos repository provides two checkpoints built on the identical Mixture-of-Transformers (MoT) architecture, differing only in parameter count and compute requirements. Whether you are testing pipelines or deploying world-model generation at scale, understanding the trade-offs between Cosmos 3 Nano and Cosmos 3 Super ensures you allocate the right hardware for your specific multimodal or reasoning workload.

Architecture Overview: Why the Same Codebase Powers Both Models

Both Cosmos 3 Nano and Cosmos 3 Super share the exact same model architecture implemented in the NVIDIA Cosmos codebase. According to the architecture diagram in cookbooks/cosmos3/cosmos3-model-architecture.png, the MoT design combines an autoregressive transformer for reasoning with a diffusion transformer for generation, unified by 3-D rotary positional embeddings (mRoPE) that encode spatial-temporal structure across modalities.

Because the core graph is identical, you can swap between Nano and Super checkpoints without changing application code. The only adjustments required are resource allocation—specifically GPU count, tensor-parallelism settings, and memory offloading strategies.

Cosmos 3 Nano vs Cosmos 3 Super: Key Differences

Aspect Cosmos 3 Nano Cosmos 3 Super
Parameters 16 billion 64 billion
Typical Hardware 1 GPU (e.g., RTX 3090) 4 GPUs with tensor-parallel size = 4
Inference Latency Faster per-frame generation 2–3× higher latency than Nano
Memory Footprint Fits in single GPU VRAM Requires layer-wise offload or multi-GPU distribution
Visual Fidelity Good for 256p–480p outputs Highest fidelity at 720p, better temporal coherence
Recommended Use Prototyping, low-latency endpoints Production content, physical AI tasks

The benchmark tables in inference_benchmarks.md quantify these differences, showing that while Nano prioritizes speed, Super excels at nuanced action reasoning and high-resolution generation.

When to Use Cosmos 3 Nano

Choose Cosmos 3 Nano when you need rapid iteration and have limited GPU memory. The 16B parameter model loads efficiently on single-GPU setups using Diffusers or vLLM-Omni without tensor parallelism.

Ideal scenarios include:

  • Early-stage research and development
  • Applications where response time matters more than absolute image quality
  • Text-captioning endpoints or low-resolution video previews (480p and below)
  • Development environments with consumer-grade hardware

When to Use Cosmos 3 Super

Select Cosmos 3 Super when visual realism and physical plausibility are critical. The 64B parameter checkpoint enables richer world-model representations necessary for complex action-policy and forward-dynamics tasks.

Deploy Super for:

  • Production-grade 720p video generation with sound synchronization
  • Fine-grained robot-policy generation requiring precise physical reasoning
  • High-fidelity image-to-video tasks where temporal coherence is essential
  • Evaluations requiring top Physics IQ scores (documented in evaluation/cosmos3/Physics_IQ/README.md)

Deployment Examples

Loading Checkpoints with Diffusers

The Cosmos3OmniPipeline class loads either checkpoint by changing the pretrained model name. Both use the same tokenizer and scheduler configuration:

import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler

# Switch between Nano and Super by changing the checkpoint string

checkpoint = "nvidia/Cosmos3-Super"  # or "nvidia/Cosmos3-Nano"

pipe = Cosmos3OmniPipeline.from_pretrained(
    checkpoint,
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, 
    flow_shift=10.0
)

result = pipe(
    prompt="A futuristic city skyline at sunset.",
    num_frames=189,
    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    generator=torch.Generator(device="cuda").manual_seed(1234),
)

The only difference between deployments is the checkpoint path; all other initialization code remains identical as documented in cookbooks/cosmos3/generator/audiovisual/README.md.

Serving with vLLM-Omni

For Super deployment across multiple GPUs, use tensor parallelism with layer-wise offloading:

docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v "$(pwd):/workspace" -p 8000:8000 --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Super \
  --omni \
  --model-class-name Cosmos3OmniDiffusersPipeline \
  --tensor-parallel-size 4 \
  --enable-layerwise-offload \
  --port 8000 \
  --init-timeout 1800

Client requests use the standard OpenAI-compatible API:

import requests

payload = {
    "prompt": "A robot arm assembles a wooden chair.",
    "size": "1280x720",
    "num_frames": 189,
    "fps": 24,
    "num_inference_steps": 35,
    "guidance_scale": 6.0,
    "flow_shift": 10.0,
    "seed": 0
}

resp = requests.post("http://localhost:8000/v1/videos/sync", json=payload)

Running the NIM Container

For text-only reasoning with Super, deploy the NIM container with the model size environment variable:

docker run -it --rm --gpus all \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NIM_MODEL_SIZE=super \
  -p 8000:8000 nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")

response = client.chat.completions.create(
    model="nvidia/cosmos3-super-reasoner",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": [
            {"type": "video_url", "video_url": {"url": "https://example.com/demo.mp4"}},
            {"type": "text", "text": "Explain the robot's next action."}
        ]}
    ],
    max_tokens=256,
)

Performance Trade-offs

Latency: The Super checkpoint incurs roughly 2–3× higher per-frame latency compared to Nano on equivalent hardware, as documented in inference_benchmarks.md. This reflects the increased attention heads and feed-forward layers processed per token in the 64B model.

Memory: A single Super checkpoint exceeds the capacity of modern single-GPU configurations. The official launch scripts in the repository recommend four GPUs with --tensor-parallel-size 4 or enabling --enable-layerwise-offload to distribute weights across GPU and CPU memory.

Quality: Qualitative evaluations in evaluation/cosmos3/Physics_IQ/README.md consistently favor Super for high-resolution image-to-video tasks and complex action predictions, while Nano provides sufficient quality for proof-of-concept development.

Summary

  • Cosmos 3 Nano (16B) and Cosmos 3 Super (64B) share identical MoT architecture and tokenizers, differing only in parameter count and compute requirements.
  • Use Nano for single-GPU prototyping, low-latency endpoints, and 480p or lower resolution outputs.
  • Use Super for multi-GPU production deployments requiring 720p video, sound synchronization, or advanced physical reasoning.
  • Switch between models by changing the checkpoint path in Cosmos3OmniPipeline.from_pretrained() or the --tensor-parallel-size flag in vLLM-Omni—no code refactoring required.
  • Reference inference_benchmarks.md for detailed latency metrics and cookbooks/cosmos3/reasoner/README.md for scaling configurations.

Frequently Asked Questions

What is the parameter difference between Cosmos 3 Nano and Super?

Cosmos 3 Nano contains 16 billion parameters, while Cosmos 3 Super contains 64 billion parameters. This 4× increase in model size enables the Super variant to learn richer latent representations for complex multimodal tasks, but requires significantly more GPU memory and compute resources during inference.

Can I run Cosmos 3 Super on a single GPU?

No, a single Super checkpoint exceeds the memory capacity of individual modern GPUs. According to the deployment guides in cookbooks/cosmos3/reasoner/README.md, Super requires four GPUs with tensor parallelism (--tensor-parallel-size 4) or layer-wise offloading (--enable-layerwise-offload) to distribute the model weights across devices and CPU memory.

How do I switch between Nano and Super in my code?

You switch models by changing the checkpoint identifier string. In Diffusers, update the from_pretrained() call to use either "nvidia/Cosmos3-Nano" or "nvidia/Cosmos3-Super". For vLLM-Omni, change the model argument in the serve command. The architecture, tokenizer, and API endpoints remain identical, so no other code changes are necessary.

Does Cosmos 3 Super produce better quality than Nano?

Yes. The 64B Super model delivers higher visual fidelity, better temporal coherence in video generation, and superior performance on physical reasoning benchmarks (Physics IQ) compared to the 16B Nano model. However, for rapid prototyping or applications where latency matters more than absolute quality, Nano provides sufficient performance with significantly faster inference speeds.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →