# How to Integrate Cosmos 3 Generator with Diffusers for Research and Development

> Integrate NVIDIA Cosmos 3 Generator with Diffusers for research and development. Use Cosmos3OmniDiffusersPipeline for text to image, video generation without Docker or custom CUDA kernels.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-13

---

**You can integrate NVIDIA Cosmos 3 Generator with Diffusers by installing the cosmos3 package extras and instantiating the `Cosmos3OmniDiffusersPipeline` class, which provides a Python-first interface for text-to-image, text-to-video, and image-to-video generation without requiring Docker or custom CUDA kernels.**

NVIDIA Cosmos 3 is an omni-modal world model family designed for generative AI research. When you integrate Cosmos 3 Generator with Diffusers, you gain direct access to the model's **Unified Mixture-of-Transformers (MoT)** backbone through a standard Hugging Face interface, enabling rapid prototyping and closed-loop experimentation with multi-modal inputs.

## Setting Up the Diffusers Environment

Before loading the pipeline, you must configure a dedicated Python environment with the specific dependencies required by the Cosmos 3 framework.

### Creating the Virtual Environment

The repository documentation recommends creating an isolated virtual environment under `packages/cosmos3/.venv` to avoid conflicts with system packages. This location is referenced in the README's Diffusers setup section and ensures that the Cosmos-specific extensions are properly isolated.

Run the following commands from the repository root:

```bash

# Create the virtual environment

python -m venv packages/cosmos3/.venv

# Activate it

source packages/cosmos3/.venv/bin/activate  # On Windows: packages/cosmos3\.venv\Scripts\activate

```

### Installing Dependencies

Install the core Diffusers library along with the Cosmos-specific extras that expose the `Cosmos3OmniDiffusersPipeline` class. The `[diffusers]` extra pulls in the necessary tokenizer components and scheduler integrations.

```bash
pip install "diffusers[torch]"
pip install -e packages/cosmos3[diffusers]

```

This installation makes the `cosmos_framework` package available, including the pipeline source located at [`cosmos_framework/pipelines/cosmos3_omni_diffusers.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/pipelines/cosmos3_omni_diffusers.py).

## Loading and Running the Cosmos 3 Diffusers Pipeline

The `Cosmos3OmniDiffusersPipeline` class follows the standard Diffusers API pattern, accepting model identifiers and hardware configurations through `from_pretrained()` parameters.

### Instantiating the Pipeline

Load the checkpoint by specifying the model identifier (e.g., `nvidia/Cosmos3-Nano` or `nvidia/Cosmos3-Super`) and desired tensor precision. The pipeline automatically downloads the full checkpoint including the diffusion model, reasoner weights, and media tokenizers.

```python
from cosmos_framework import Cosmos3OmniDiffusersPipeline
import torch

pipeline = Cosmos3OmniDiffusersPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.float16,
    use_safetensors=True,
).to("cuda")

```

### Text-to-Image Generation

Generate photorealistic images by calling the pipeline with text prompts and dimension parameters. The method returns a Pillow Image object that you can save or process further.

```python
image = pipeline(
    prompt="A serene lake at sunrise, photorealistic",
    height=512,
    width=512,
    num_inference_steps=50,
    guidance_scale=7.5,
).images[0]

image.save("lake.png")

```

### Text-to-Video and Image-to-Video Generation

For video generation, enable the video mode by setting `video=True` and specify the number of frames. For image-to-video, pass an initial image path to the `image` parameter.

```python

# Text-to-video generation

video = pipeline(
    prompt="A futuristic city skyline at night, flying cars",
    video=True,
    num_frames=16,
    height=480,
    width=832,
    num_inference_steps=80,
).video  # Returns torch.Tensor[C, T, H, W, 3]

# Save video (requires torchvision)

import torchvision.io as io
io.write_video("city.mp4", video.permute(1, 0, 2, 3, 4), fps=24)

# Image-to-video generation

video = pipeline(
    image="seed.jpg",
    video=True,
    num_frames=24,
    height=480,
    width=832,
).video
io.write_video("output.mp4", video.permute(1, 0, 2, 3, 4), fps=24)

```

## Architecture and Key Components

Understanding the internal structure helps when modifying the pipeline for research purposes or integrating custom schedulers.

### Unified Mixture-of-Transformers (MoT)

Cosmos 3's backbone processes all modalities jointly through a **Unified Mixture-of-Transformers** architecture. The Diffusers pipeline maps this MoT's diffusion interface onto standard Diffusers scheduler APIs, allowing you to substitute schedulers like `DDIMScheduler` or `DPMSolverMultistepScheduler` without modifying the underlying model weights.

### Media Tokenizers and Scheduler Integration

The pipeline automatically instantiates integrated media tokenizers that translate raw video and audio inputs into the model's token space. These tokenizers handle text, image, video, and audio preprocessing, enabling true multi-modal conditioning. The diffusion scheduler and decoder then execute the denoising process, returning standard Python objects (Pillow images or NumPy/torch arrays) rather than raw tensors requiring manual post-processing.

## Advanced Research Workflows

Because the Diffusers integration loads the **full checkpoint**, you have direct access to the reasoner model inside the pipeline. This enables **closed-loop research** workflows where you generate content, feed it to the Cosmos 3 Reasoner for captioning or embedding extraction, and condition subsequent generations on those outputs—all within a single Python script.

The notebook `cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb` demonstrates this workflow, including kernel registration for the `Cosmos3 Diffusers (Python 3.13)` environment and verification steps to ensure the pipeline loads correctly.

## Performance Benchmarks

According to the [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) file in the repository, the Diffusers backend achieves competitive latency for research prototyping. For example, the Nano checkpoint generates 256p images in approximately **11 ms** on supported hardware. While the vLLM-Omni server provides higher throughput for production deployments, the Diffusers path offers comparable performance for low-resolution runs and eliminates the overhead of containerization.

## Summary

- **Install** the Cosmos 3 Diffusers extras using `pip install -e packages/cosmos3[diffusers]` after setting up the virtual environment at `packages/cosmos3/.venv`.
- **Instantiate** the `Cosmos3OmniDiffusersPipeline` class with `from_pretrained()` using model IDs like `nvidia/Cosmos3-Nano`.
- **Generate** images, videos, or image-to-video sequences by calling the pipeline with appropriate flags (`video=True` for video generation).
- **Access** the full checkpoint including reasoner weights for closed-loop research without building Docker images.
- **Benchmark** performance using the latencies documented in [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md), with 11 ms achievable for 256p image generation.

## Frequently Asked Questions

### What is the difference between using Cosmos 3 Generator with Diffusers versus the vLLM-Omni server?

The Diffusers integration provides a Python-first interface suitable for research and rapid prototyping, allowing you to inspect and modify the `Cosmos3OmniDiffusersPipeline` class directly. The vLLM-Omni server is optimized for high-throughput production inference and requires containerization. For development and experimentation, Diffusers offers greater flexibility without custom CUDA kernel development.

### Can I use custom schedulers with the Cosmos 3 Diffusers pipeline?

Yes. Because the pipeline follows the standard Diffusers API, you can replace the default scheduler with any compatible Diffusers scheduler such as `DDIMScheduler` or `DPMSolverMultistepScheduler`. The pipeline maps these schedulers onto the Cosmos 3 MoT backbone while maintaining the model's multi-modal tokenization capabilities.

### How do I access the reasoner model when using the Diffusers pipeline?

The Diffusers integration loads the full checkpoint including reasoner weights, giving you direct access to the reasoner component for conditioning generation on image captions, embeddings, or action trajectories. You can call the reasoner between generation steps in a single Python script to implement closed-loop research workflows that iterate between generation and reasoning.

### Where can I find executable examples for the Diffusers integration?

The repository includes the Jupyter notebook `cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb`, which provides step-by-step instructions for environment setup, kernel registration, and running text-to-image and text-to-video examples. This notebook also demonstrates how to verify that the `Cosmos3OmniDiffusersPipeline` loads correctly on your hardware.