# Why NVIDIA Cosmos First Run Inference Takes Long (And Why It's Not a Hang)

> NVIDIA Cosmos first run inference takes time due to large model downloads and diffusion. Learn why this expected behavior isn't a pipeline hang.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: performance
- Published: 2026-06-06

---

**NVIDIA Cosmos first run inference delays are expected behavior caused by multi-gigabyte model downloads and compute-intensive diffusion iterations—not a pipeline hang.**

When executing a Cosmos 3 generator (e.g., `text-to-video`) from the `NVIDIA/cosmos` repository for the first time, the pipeline appears to freeze for several minutes. This article explains the two technical reasons behind these **NVIDIA Cosmos first run inference** delays and demonstrates how to verify the system is functioning correctly.

## Model Download Overhead

The first invocation triggers a download of the Cosmos 3 checkpoint (e.g., `nvidia/Cosmos3-Nano`). The model weighs dozens of gigabytes, so the download can take a few minutes depending on network speed. As documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) at lines 334-336, the initialization phase requires transferring the entire checkpoint to local storage before any computation begins.

## Diffusion Computation Requirements

Unlike a pure language model that generates tokens sequentially, Cosmos 3’s generator runs a **diffusion process** that iterates over all diffusion steps before any frame is produced. When you call the pipeline with `num_inference_steps=35`, the system executes all 35 steps across the latent space before outputting video frames. This compute-intensive loop dominates the runtime on the first run and remains the primary cost in subsequent runs after the model is cached.

## How to Verify the Pipeline Is Not Hung

The `NVIDIA/cosmos` source code explicitly notes that *“a `text-to-video` run takes a while: the first run downloads `Cosmos3-Nano`, and diffusion is compute-heavy, running through every inference step before producing output. Long step times are expected, not a hang.”* Because the model is cached after the initial download, subsequent runs start much faster (only the diffusion computation remains).

To confirm the pipeline is processing and not frozen:

- **First run:** Expect `from_pretrained()` to display a download progress bar (tens of seconds to a few minutes), followed by the `pipe()` call taking additional time (typically 30–90 seconds on a single-GPU RTX 4090) to perform 35 diffusion steps for each frame.
- **Subsequent runs:** No download occurs; the same `pipe()` call completes noticeably faster (≈ 10–20 seconds) since only the diffusion computation remains.

## Code Example: Reproducing First-Run Behavior

Run this minimal end-to-end script in a fresh environment to observe the download and diffusion latency:

```python
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
from diffusers.utils import export_to_video

# Load the model – the first call will download the checkpoint.

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

# Use the UniPC scheduler (recommended for Cosmos 3).

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

# Run the diffusion pipeline (this step is compute-heavy).

result = pipe(
    prompt="A mobile robot navigates a warehouse aisle and stops at a shelf.",
    negative_prompt="",
    image=None,
    num_frames=189,
    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    enable_sound=False,
    add_resolution_template=False,
    add_duration_template=False,
    generator=torch.Generator(device="cuda").manual_seed(1234),
)

# Save the generated video.

export_to_video(result.video, "cosmos3_t2v.mp4", fps=24, macro_block_size=1)
print("✅ Video written to cosmos3_t2v.mp4")

```

## Key Source Files Explaining the Behavior

The following files in the `NVIDIA/cosmos` repository document the first-run latency behavior:

- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** – Contains the primary documentation at lines 334-336 explaining that long step times are expected, not a hang.
- **[`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md)** – Quick-start guide for the Diffusers pipeline that reiterates the first-run latency note.
- **`cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb`** – Runnable notebook demonstrating the pipeline with expected timing behavior.
- **`cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb`** – Shows the same workflow via the Cosmos Framework entry point with identical initialization characteristics.

## Summary

- **First run downloads:** The `nvidia/Cosmos3-Nano` checkpoint (dozens of gigabytes) downloads automatically via `Cosmos3OmniPipeline.from_pretrained()`.
- **Diffusion is iterative:** The pipeline executes all `num_inference_steps` (default 35) through the UniPC scheduler before generating frames, unlike autoregressive language models.
- **Caching applies:** Model weights are cached locally after first download; only diffusion computation repeats in subsequent runs.
- **Expected timing:** 30–90 seconds for initial diffusion pass on high-end GPUs, reducing to 10–20 seconds for cached runs.

## Frequently Asked Questions

### Is the NVIDIA Cosmos pipeline hung during first run?

No. The pipeline is either downloading the multi-gigabyte checkpoint or executing the full diffusion sampling loop. According to the source code in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) lines 334-336, extended initialization periods are expected behavior, not a hang.

### How long should the first inference take?

The first run typically takes several minutes total: potentially 1–5 minutes for the model download depending on bandwidth, plus 30–90 seconds for the diffusion computation on a single RTX 4090. Subsequent runs eliminate the download phase and complete in approximately 10–20 seconds.

### Does the model download every time?

No. The `from_pretrained()` method caches the checkpoint locally after the first successful download. Only the diffusion computation repeats on subsequent runs, though this remains compute-intensive due to the iterative denoising process.

### Why is diffusion slower than LLM inference?

Cosmos 3 uses a diffusion model that iteratively refines latent noise over multiple steps (e.g., 35) before producing output frames, whereas autoregressive LLMs generate tokens sequentially in a single forward pass per token. This multi-step diffusion loop requires significantly more compute time per output unit than transformer-based text generation.