Text-to-Video Generation with Cosmos 3 Using Diffusers: Setup and Inference Guide

Use the Hugging Face Diffusers library to load the Cosmos3OmniPipeline from the nvidia/Cosmos3-Nano checkpoint, configure the UniPCMultistepScheduler, and call the pipeline with a text prompt to generate a video locally on a GPU.

The NVIDIA/cosmos repository provides a unified Generator surface for synthesizing videos from pure text prompts. The most accessible way to experiment with this capability is through the Hugging Face Diffusers library, which loads the full checkpoint and runs the diffusion process locally. This guide covers the exact steps needed to run text-to-video generation with Cosmos 3 using Diffusers, from environment setup to exporting the final MP4.

Environment Setup

According to the NVIDIA/cosmos README, you must install Diffusers from the upstream GitHub repository to pull the latest Cosmos 3 support. You will also need PyTorch, media utilities, and the guardrail package.

Install Dependencies

Create a virtual environment and install the required runtime dependencies:

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
uv pip install --torch-backend=auto \
  "diffusers @ git+https://github.com/huggingface/diffusers.git" \
  accelerate av cosmos_guardrail huggingface_hub \
  imageio imageio-ffmpeg torch torchvision transformers

This installs diffusers from source along with accelerate, av, imageio, and torchvision, which the pipeline relies on for inference and MP4 serialization.

Build the Diffusers Pipeline

The pipeline automatically loads the tokenizer, multimodal tokenizers, and diffusion model from the model hub. You should initialize it in bfloat16 on a CUDA device to match the training precision.

Load the Cosmos 3 Checkpoint

Instantiate the pipeline with from_pretrained:

import torch
from diffusers import Cosmos3OmniPipeline

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

You can swap nvidia/Cosmos3-Nano for nvidia/Cosmos3-Super if you need higher quality and have sufficient VRAM. Expect a roughly 7 GB download for the Nano checkpoint on first run.

Configure the Scheduler for Smooth Motion

Replace the default scheduler with UniPCMultistepScheduler and set flow_shift to improve temporal dynamics:

from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

Run Text-to-Video Inference

With the pipeline ready, call it with a descriptive prompt and export the video tensor to an MP4 using export_to_video. The example below generates a 7.9-second clip at 720p and 24 FPS.

from diffusers.utils import export_to_video

result = pipe(
    prompt="A mobile robot navigates a warehouse aisle and stops at a shelf.",
    negative_prompt="",
    image=None,                     # No image conditioning

    num_frames=189,                 # ≈7.9 s at 24 FPS

    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    enable_sound=False,             # Set True to get synchronized audio

    add_resolution_template=False,
    add_duration_template=False,
    generator=torch.Generator(device="cuda").manual_seed(1234),
)

export_to_video(result.video, "cosmos3_text2video.mp4", fps=24, macro_block_size=1)
print("Video saved to cosmos3_text2video.mp4")

What Happens During Diffusion

As implemented in the NVIDIA/cosmos source code, the pipeline invokes the diffusion branch of a shared Mixture-of-Transformers (MoT) backbone. This branch repeatedly denoises a latent video tensor across num_inference_steps while conditioning on the tokenized prompt. The architecture uses 3-D Rotary Position Embeddings (mRoPE) to encode spatial and temporal relationships, which helps maintain consistent motion and geometry across all 189 frames.

By default, the pipeline enforces safety guardrails that filter unsafe prompts and blur recognizable faces. You can disable these per-request via extra_params if your use case requires it.

Expected Performance

The total wall-clock time is dominated by the diffusion steps. On a single RTX 4090, a full 35-step inference run takes several minutes. The resulting output is a result.video tensor that export_to_video writes directly to disk.

Alternative Entry Point

If you prefer an interactive environment, a ready-to-run Jupyter notebook reproduces the same steps and includes inline visualization of intermediate frames. You can find it at:

cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb

This notebook is maintained in the NVIDIA/cosmos repository alongside the canonical README instructions.

Summary

  • Install Diffusers from source along with torch, accelerate, av, and imageio to obtain the latest Cosmos 3 pipeline code.
  • Load the Cosmos3OmniPipeline from nvidia/Cosmos3-Nano in torch.bfloat16 with device_map="cuda".
  • Swap the default scheduler for UniPCMultistepScheduler and set flow_shift=10.0 for smoother temporal dynamics.
  • Call the pipeline with a text prompt, num_frames, resolution, and guidance_scale, then export the result with export_to_video.
  • Safety guardrails run by default; disable them via extra_params if required.

Frequently Asked Questions

What hardware is required to run Cosmos 3 text-to-video with Diffusers?

A CUDA-capable GPU is required. The Nano checkpoint downloads roughly 7 GB of weights, and the diffusion process is compute-intensive; expect a single RTX 4090 to take several minutes for a standard 35-step generation.

Can I generate synchronized audio along with the video?

Yes. Pass enable_sound=True in the pipeline call to request synchronized audio generation. The default example sets this to False to produce a silent video output.

Where is the official Diffusers example located in the repository?

The canonical snippet lives in README.md under the Generator with Diffusers section. A complete end-to-end Jupyter notebook is also provided at cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb.

How does the pipeline maintain consistent motion across frames?

The diffusion branch uses a shared Mixture-of-Transformers architecture and 3-D Rotary Position Embeddings (mRoPE) to encode spatial and temporal relationships across the latent video tensor during denoising, which preserves coherent motion and geometry throughout the clip.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →