# Text-to-Video Generation with Cosmos 3 Using Diffusers: Setup and Inference Guide

> Generate videos from text using Cosmos 3 and Diffusers. This guide covers setup and inference, allowing you to create videos locally on your GPU with simple text prompts.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-05

---

**Use the Hugging Face Diffusers library to load the `Cosmos3OmniPipeline` from the `nvidia/Cosmos3-Nano` checkpoint, configure the `UniPCMultistepScheduler`, and call the pipeline with a text prompt to generate a video locally on a GPU.**

The `NVIDIA/cosmos` repository provides a unified **Generator** surface for synthesizing videos from pure text prompts. The most accessible way to experiment with this capability is through the **Hugging Face Diffusers** library, which loads the full checkpoint and runs the diffusion process locally. This guide covers the exact steps needed to run **text-to-video generation with Cosmos 3 using Diffusers**, from environment setup to exporting the final MP4.

## Environment Setup

According to the `NVIDIA/cosmos` README, you must install Diffusers from the upstream GitHub repository to pull the latest Cosmos 3 support. You will also need PyTorch, media utilities, and the guardrail package.

### Install Dependencies

Create a virtual environment and install the required runtime dependencies:

```bash
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
uv pip install --torch-backend=auto \
  "diffusers @ git+https://github.com/huggingface/diffusers.git" \
  accelerate av cosmos_guardrail huggingface_hub \
  imageio imageio-ffmpeg torch torchvision transformers

```

This installs `diffusers` from source along with `accelerate`, `av`, `imageio`, and `torchvision`, which the pipeline relies on for inference and MP4 serialization.

## Build the Diffusers Pipeline

The pipeline automatically loads the tokenizer, multimodal tokenizers, and diffusion model from the model hub. You should initialize it in `bfloat16` on a CUDA device to match the training precision.

### Load the Cosmos 3 Checkpoint

Instantiate the pipeline with `from_pretrained`:

```python
import torch
from diffusers import Cosmos3OmniPipeline

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

```

You can swap `nvidia/Cosmos3-Nano` for `nvidia/Cosmos3-Super` if you need higher quality and have sufficient VRAM. Expect a roughly 7 GB download for the Nano checkpoint on first run.

### Configure the Scheduler for Smooth Motion

Replace the default scheduler with `UniPCMultistepScheduler` and set `flow_shift` to improve temporal dynamics:

```python
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

```

## Run Text-to-Video Inference

With the pipeline ready, call it with a descriptive prompt and export the video tensor to an MP4 using `export_to_video`. The example below generates a 7.9-second clip at 720p and 24 FPS.

```python
from diffusers.utils import export_to_video

result = pipe(
    prompt="A mobile robot navigates a warehouse aisle and stops at a shelf.",
    negative_prompt="",
    image=None,                     # No image conditioning

    num_frames=189,                 # ≈7.9 s at 24 FPS

    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    enable_sound=False,             # Set True to get synchronized audio

    add_resolution_template=False,
    add_duration_template=False,
    generator=torch.Generator(device="cuda").manual_seed(1234),
)

export_to_video(result.video, "cosmos3_text2video.mp4", fps=24, macro_block_size=1)
print("Video saved to cosmos3_text2video.mp4")

```

### What Happens During Diffusion

As implemented in the `NVIDIA/cosmos` source code, the pipeline invokes the diffusion branch of a shared **Mixture-of-Transformers (MoT)** backbone. This branch repeatedly denoises a latent video tensor across `num_inference_steps` while conditioning on the tokenized prompt. The architecture uses **3-D Rotary Position Embeddings (mRoPE)** to encode spatial and temporal relationships, which helps maintain consistent motion and geometry across all 189 frames.

By default, the pipeline enforces safety **guardrails** that filter unsafe prompts and blur recognizable faces. You can disable these per-request via `extra_params` if your use case requires it.

### Expected Performance

The total wall-clock time is dominated by the diffusion steps. On a single RTX 4090, a full 35-step inference run takes several minutes. The resulting output is a `result.video` tensor that `export_to_video` writes directly to disk.

## Alternative Entry Point

If you prefer an interactive environment, a ready-to-run Jupyter notebook reproduces the same steps and includes inline visualization of intermediate frames. You can find it at:

```text
cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb

```

This notebook is maintained in the `NVIDIA/cosmos` repository alongside the canonical README instructions.

## Summary

- Install Diffusers from source along with `torch`, `accelerate`, `av`, and `imageio` to obtain the latest Cosmos 3 pipeline code.
- Load the `Cosmos3OmniPipeline` from `nvidia/Cosmos3-Nano` in `torch.bfloat16` with `device_map="cuda"`.
- Swap the default scheduler for `UniPCMultistepScheduler` and set `flow_shift=10.0` for smoother temporal dynamics.
- Call the pipeline with a text prompt, `num_frames`, resolution, and `guidance_scale`, then export the result with `export_to_video`.
- Safety guardrails run by default; disable them via `extra_params` if required.

## Frequently Asked Questions

### What hardware is required to run Cosmos 3 text-to-video with Diffusers?

A CUDA-capable GPU is required. The Nano checkpoint downloads roughly 7 GB of weights, and the diffusion process is compute-intensive; expect a single RTX 4090 to take several minutes for a standard 35-step generation.

### Can I generate synchronized audio along with the video?

Yes. Pass `enable_sound=True` in the pipeline call to request synchronized audio generation. The default example sets this to `False` to produce a silent video output.

### Where is the official Diffusers example located in the repository?

The canonical snippet lives in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) under the *Generator with Diffusers* section. A complete end-to-end Jupyter notebook is also provided at `cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb`.

### How does the pipeline maintain consistent motion across frames?

The diffusion branch uses a shared **Mixture-of-Transformers** architecture and **3-D Rotary Position Embeddings (mRoPE)** to encode spatial and temporal relationships across the latent video tensor during denoising, which preserves coherent motion and geometry throughout the clip.