Text-to-Video Generation with Cosmos 3 Using Diffusers: Setup and Inference Guide
Use the Hugging Face Diffusers library to load the Cosmos3OmniPipeline from the nvidia/Cosmos3-Nano checkpoint, configure the UniPCMultistepScheduler, and call the pipeline with a text prompt to generate a video locally on a GPU.
The NVIDIA/cosmos repository provides a unified Generator surface for synthesizing videos from pure text prompts. The most accessible way to experiment with this capability is through the Hugging Face Diffusers library, which loads the full checkpoint and runs the diffusion process locally. This guide covers the exact steps needed to run text-to-video generation with Cosmos 3 using Diffusers, from environment setup to exporting the final MP4.
Environment Setup
According to the NVIDIA/cosmos README, you must install Diffusers from the upstream GitHub repository to pull the latest Cosmos 3 support. You will also need PyTorch, media utilities, and the guardrail package.
Install Dependencies
Create a virtual environment and install the required runtime dependencies:
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
uv pip install --torch-backend=auto \
"diffusers @ git+https://github.com/huggingface/diffusers.git" \
accelerate av cosmos_guardrail huggingface_hub \
imageio imageio-ffmpeg torch torchvision transformers
This installs diffusers from source along with accelerate, av, imageio, and torchvision, which the pipeline relies on for inference and MP4 serialization.
Build the Diffusers Pipeline
The pipeline automatically loads the tokenizer, multimodal tokenizers, and diffusion model from the model hub. You should initialize it in bfloat16 on a CUDA device to match the training precision.
Load the Cosmos 3 Checkpoint
Instantiate the pipeline with from_pretrained:
import torch
from diffusers import Cosmos3OmniPipeline
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
You can swap nvidia/Cosmos3-Nano for nvidia/Cosmos3-Super if you need higher quality and have sufficient VRAM. Expect a roughly 7 GB download for the Nano checkpoint on first run.
Configure the Scheduler for Smooth Motion
Replace the default scheduler with UniPCMultistepScheduler and set flow_shift to improve temporal dynamics:
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
pipe.scheduler = UniPCMultistepScheduler.from_config(
pipe.scheduler.config, flow_shift=10.0
)
Run Text-to-Video Inference
With the pipeline ready, call it with a descriptive prompt and export the video tensor to an MP4 using export_to_video. The example below generates a 7.9-second clip at 720p and 24 FPS.
from diffusers.utils import export_to_video
result = pipe(
prompt="A mobile robot navigates a warehouse aisle and stops at a shelf.",
negative_prompt="",
image=None, # No image conditioning
num_frames=189, # ≈7.9 s at 24 FPS
height=720,
width=1280,
fps=24,
num_inference_steps=35,
guidance_scale=6.0,
enable_sound=False, # Set True to get synchronized audio
add_resolution_template=False,
add_duration_template=False,
generator=torch.Generator(device="cuda").manual_seed(1234),
)
export_to_video(result.video, "cosmos3_text2video.mp4", fps=24, macro_block_size=1)
print("Video saved to cosmos3_text2video.mp4")
What Happens During Diffusion
As implemented in the NVIDIA/cosmos source code, the pipeline invokes the diffusion branch of a shared Mixture-of-Transformers (MoT) backbone. This branch repeatedly denoises a latent video tensor across num_inference_steps while conditioning on the tokenized prompt. The architecture uses 3-D Rotary Position Embeddings (mRoPE) to encode spatial and temporal relationships, which helps maintain consistent motion and geometry across all 189 frames.
By default, the pipeline enforces safety guardrails that filter unsafe prompts and blur recognizable faces. You can disable these per-request via extra_params if your use case requires it.
Expected Performance
The total wall-clock time is dominated by the diffusion steps. On a single RTX 4090, a full 35-step inference run takes several minutes. The resulting output is a result.video tensor that export_to_video writes directly to disk.
Alternative Entry Point
If you prefer an interactive environment, a ready-to-run Jupyter notebook reproduces the same steps and includes inline visualization of intermediate frames. You can find it at:
cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb
This notebook is maintained in the NVIDIA/cosmos repository alongside the canonical README instructions.
Summary
- Install Diffusers from source along with
torch,accelerate,av, andimageioto obtain the latest Cosmos 3 pipeline code. - Load the
Cosmos3OmniPipelinefromnvidia/Cosmos3-Nanointorch.bfloat16withdevice_map="cuda". - Swap the default scheduler for
UniPCMultistepSchedulerand setflow_shift=10.0for smoother temporal dynamics. - Call the pipeline with a text prompt,
num_frames, resolution, andguidance_scale, then export the result withexport_to_video. - Safety guardrails run by default; disable them via
extra_paramsif required.
Frequently Asked Questions
What hardware is required to run Cosmos 3 text-to-video with Diffusers?
A CUDA-capable GPU is required. The Nano checkpoint downloads roughly 7 GB of weights, and the diffusion process is compute-intensive; expect a single RTX 4090 to take several minutes for a standard 35-step generation.
Can I generate synchronized audio along with the video?
Yes. Pass enable_sound=True in the pipeline call to request synchronized audio generation. The default example sets this to False to produce a silent video output.
Where is the official Diffusers example located in the repository?
The canonical snippet lives in README.md under the Generator with Diffusers section. A complete end-to-end Jupyter notebook is also provided at cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb.
How does the pipeline maintain consistent motion across frames?
The diffusion branch uses a shared Mixture-of-Transformers architecture and 3-D Rotary Position Embeddings (mRoPE) to encode spatial and temporal relationships across the latent video tensor during denoising, which preserves coherent motion and geometry throughout the clip.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →