Why NVIDIA Cosmos First Run Inference Takes Long (And Why It's Not a Hang)
NVIDIA Cosmos first run inference delays are expected behavior caused by multi-gigabyte model downloads and compute-intensive diffusion iterations—not a pipeline hang.
When executing a Cosmos 3 generator (e.g., text-to-video) from the NVIDIA/cosmos repository for the first time, the pipeline appears to freeze for several minutes. This article explains the two technical reasons behind these NVIDIA Cosmos first run inference delays and demonstrates how to verify the system is functioning correctly.
Model Download Overhead
The first invocation triggers a download of the Cosmos 3 checkpoint (e.g., nvidia/Cosmos3-Nano). The model weighs dozens of gigabytes, so the download can take a few minutes depending on network speed. As documented in README.md at lines 334-336, the initialization phase requires transferring the entire checkpoint to local storage before any computation begins.
Diffusion Computation Requirements
Unlike a pure language model that generates tokens sequentially, Cosmos 3’s generator runs a diffusion process that iterates over all diffusion steps before any frame is produced. When you call the pipeline with num_inference_steps=35, the system executes all 35 steps across the latent space before outputting video frames. This compute-intensive loop dominates the runtime on the first run and remains the primary cost in subsequent runs after the model is cached.
How to Verify the Pipeline Is Not Hung
The NVIDIA/cosmos source code explicitly notes that “a text-to-video run takes a while: the first run downloads Cosmos3-Nano, and diffusion is compute-heavy, running through every inference step before producing output. Long step times are expected, not a hang.” Because the model is cached after the initial download, subsequent runs start much faster (only the diffusion computation remains).
To confirm the pipeline is processing and not frozen:
- First run: Expect
from_pretrained()to display a download progress bar (tens of seconds to a few minutes), followed by thepipe()call taking additional time (typically 30–90 seconds on a single-GPU RTX 4090) to perform 35 diffusion steps for each frame. - Subsequent runs: No download occurs; the same
pipe()call completes noticeably faster (≈ 10–20 seconds) since only the diffusion computation remains.
Code Example: Reproducing First-Run Behavior
Run this minimal end-to-end script in a fresh environment to observe the download and diffusion latency:
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
from diffusers.utils import export_to_video
# Load the model – the first call will download the checkpoint.
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
# Use the UniPC scheduler (recommended for Cosmos 3).
pipe.scheduler = UniPCMultistepScheduler.from_config(
pipe.scheduler.config, flow_shift=10.0
)
# Run the diffusion pipeline (this step is compute-heavy).
result = pipe(
prompt="A mobile robot navigates a warehouse aisle and stops at a shelf.",
negative_prompt="",
image=None,
num_frames=189,
height=720,
width=1280,
fps=24,
num_inference_steps=35,
guidance_scale=6.0,
enable_sound=False,
add_resolution_template=False,
add_duration_template=False,
generator=torch.Generator(device="cuda").manual_seed(1234),
)
# Save the generated video.
export_to_video(result.video, "cosmos3_t2v.mp4", fps=24, macro_block_size=1)
print("✅ Video written to cosmos3_t2v.mp4")
Key Source Files Explaining the Behavior
The following files in the NVIDIA/cosmos repository document the first-run latency behavior:
README.md– Contains the primary documentation at lines 334-336 explaining that long step times are expected, not a hang.cookbooks/cosmos3/README.md– Quick-start guide for the Diffusers pipeline that reiterates the first-run latency note.cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb– Runnable notebook demonstrating the pipeline with expected timing behavior.cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb– Shows the same workflow via the Cosmos Framework entry point with identical initialization characteristics.
Summary
- First run downloads: The
nvidia/Cosmos3-Nanocheckpoint (dozens of gigabytes) downloads automatically viaCosmos3OmniPipeline.from_pretrained(). - Diffusion is iterative: The pipeline executes all
num_inference_steps(default 35) through the UniPC scheduler before generating frames, unlike autoregressive language models. - Caching applies: Model weights are cached locally after first download; only diffusion computation repeats in subsequent runs.
- Expected timing: 30–90 seconds for initial diffusion pass on high-end GPUs, reducing to 10–20 seconds for cached runs.
Frequently Asked Questions
Is the NVIDIA Cosmos pipeline hung during first run?
No. The pipeline is either downloading the multi-gigabyte checkpoint or executing the full diffusion sampling loop. According to the source code in README.md lines 334-336, extended initialization periods are expected behavior, not a hang.
How long should the first inference take?
The first run typically takes several minutes total: potentially 1–5 minutes for the model download depending on bandwidth, plus 30–90 seconds for the diffusion computation on a single RTX 4090. Subsequent runs eliminate the download phase and complete in approximately 10–20 seconds.
Does the model download every time?
No. The from_pretrained() method caches the checkpoint locally after the first successful download. Only the diffusion computation repeats on subsequent runs, though this remains compute-intensive due to the iterative denoising process.
Why is diffusion slower than LLM inference?
Cosmos 3 uses a diffusion model that iteratively refines latent noise over multiple steps (e.g., 35) before producing output frames, whereas autoregressive LLMs generate tokens sequentially in a single forward pass per token. This multi-step diffusion loop requires significantly more compute time per output unit than transformer-based text generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →