How to Integrate Cosmos 3 Generator with Diffusers for Research and Development

You can integrate NVIDIA Cosmos 3 Generator with Diffusers by installing the cosmos3 package extras and instantiating the Cosmos3OmniDiffusersPipeline class, which provides a Python-first interface for text-to-image, text-to-video, and image-to-video generation without requiring Docker or custom CUDA kernels.

NVIDIA Cosmos 3 is an omni-modal world model family designed for generative AI research. When you integrate Cosmos 3 Generator with Diffusers, you gain direct access to the model's Unified Mixture-of-Transformers (MoT) backbone through a standard Hugging Face interface, enabling rapid prototyping and closed-loop experimentation with multi-modal inputs.

Setting Up the Diffusers Environment

Before loading the pipeline, you must configure a dedicated Python environment with the specific dependencies required by the Cosmos 3 framework.

Creating the Virtual Environment

The repository documentation recommends creating an isolated virtual environment under packages/cosmos3/.venv to avoid conflicts with system packages. This location is referenced in the README's Diffusers setup section and ensures that the Cosmos-specific extensions are properly isolated.

Run the following commands from the repository root:


# Create the virtual environment

python -m venv packages/cosmos3/.venv

# Activate it

source packages/cosmos3/.venv/bin/activate  # On Windows: packages/cosmos3\.venv\Scripts\activate

Installing Dependencies

Install the core Diffusers library along with the Cosmos-specific extras that expose the Cosmos3OmniDiffusersPipeline class. The [diffusers] extra pulls in the necessary tokenizer components and scheduler integrations.

pip install "diffusers[torch]"
pip install -e packages/cosmos3[diffusers]

This installation makes the cosmos_framework package available, including the pipeline source located at cosmos_framework/pipelines/cosmos3_omni_diffusers.py.

Loading and Running the Cosmos 3 Diffusers Pipeline

The Cosmos3OmniDiffusersPipeline class follows the standard Diffusers API pattern, accepting model identifiers and hardware configurations through from_pretrained() parameters.

Instantiating the Pipeline

Load the checkpoint by specifying the model identifier (e.g., nvidia/Cosmos3-Nano or nvidia/Cosmos3-Super) and desired tensor precision. The pipeline automatically downloads the full checkpoint including the diffusion model, reasoner weights, and media tokenizers.

from cosmos_framework import Cosmos3OmniDiffusersPipeline
import torch

pipeline = Cosmos3OmniDiffusersPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.float16,
    use_safetensors=True,
).to("cuda")

Text-to-Image Generation

Generate photorealistic images by calling the pipeline with text prompts and dimension parameters. The method returns a Pillow Image object that you can save or process further.

image = pipeline(
    prompt="A serene lake at sunrise, photorealistic",
    height=512,
    width=512,
    num_inference_steps=50,
    guidance_scale=7.5,
).images[0]

image.save("lake.png")

Text-to-Video and Image-to-Video Generation

For video generation, enable the video mode by setting video=True and specify the number of frames. For image-to-video, pass an initial image path to the image parameter.


# Text-to-video generation

video = pipeline(
    prompt="A futuristic city skyline at night, flying cars",
    video=True,
    num_frames=16,
    height=480,
    width=832,
    num_inference_steps=80,
).video  # Returns torch.Tensor[C, T, H, W, 3]

# Save video (requires torchvision)

import torchvision.io as io
io.write_video("city.mp4", video.permute(1, 0, 2, 3, 4), fps=24)

# Image-to-video generation

video = pipeline(
    image="seed.jpg",
    video=True,
    num_frames=24,
    height=480,
    width=832,
).video
io.write_video("output.mp4", video.permute(1, 0, 2, 3, 4), fps=24)

Architecture and Key Components

Understanding the internal structure helps when modifying the pipeline for research purposes or integrating custom schedulers.

Unified Mixture-of-Transformers (MoT)

Cosmos 3's backbone processes all modalities jointly through a Unified Mixture-of-Transformers architecture. The Diffusers pipeline maps this MoT's diffusion interface onto standard Diffusers scheduler APIs, allowing you to substitute schedulers like DDIMScheduler or DPMSolverMultistepScheduler without modifying the underlying model weights.

Media Tokenizers and Scheduler Integration

The pipeline automatically instantiates integrated media tokenizers that translate raw video and audio inputs into the model's token space. These tokenizers handle text, image, video, and audio preprocessing, enabling true multi-modal conditioning. The diffusion scheduler and decoder then execute the denoising process, returning standard Python objects (Pillow images or NumPy/torch arrays) rather than raw tensors requiring manual post-processing.

Advanced Research Workflows

Because the Diffusers integration loads the full checkpoint, you have direct access to the reasoner model inside the pipeline. This enables closed-loop research workflows where you generate content, feed it to the Cosmos 3 Reasoner for captioning or embedding extraction, and condition subsequent generations on those outputs—all within a single Python script.

The notebook cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb demonstrates this workflow, including kernel registration for the Cosmos3 Diffusers (Python 3.13) environment and verification steps to ensure the pipeline loads correctly.

Performance Benchmarks

According to the inference_benchmarks.md file in the repository, the Diffusers backend achieves competitive latency for research prototyping. For example, the Nano checkpoint generates 256p images in approximately 11 ms on supported hardware. While the vLLM-Omni server provides higher throughput for production deployments, the Diffusers path offers comparable performance for low-resolution runs and eliminates the overhead of containerization.

Summary

  • Install the Cosmos 3 Diffusers extras using pip install -e packages/cosmos3[diffusers] after setting up the virtual environment at packages/cosmos3/.venv.
  • Instantiate the Cosmos3OmniDiffusersPipeline class with from_pretrained() using model IDs like nvidia/Cosmos3-Nano.
  • Generate images, videos, or image-to-video sequences by calling the pipeline with appropriate flags (video=True for video generation).
  • Access the full checkpoint including reasoner weights for closed-loop research without building Docker images.
  • Benchmark performance using the latencies documented in inference_benchmarks.md, with 11 ms achievable for 256p image generation.

Frequently Asked Questions

What is the difference between using Cosmos 3 Generator with Diffusers versus the vLLM-Omni server?

The Diffusers integration provides a Python-first interface suitable for research and rapid prototyping, allowing you to inspect and modify the Cosmos3OmniDiffusersPipeline class directly. The vLLM-Omni server is optimized for high-throughput production inference and requires containerization. For development and experimentation, Diffusers offers greater flexibility without custom CUDA kernel development.

Can I use custom schedulers with the Cosmos 3 Diffusers pipeline?

Yes. Because the pipeline follows the standard Diffusers API, you can replace the default scheduler with any compatible Diffusers scheduler such as DDIMScheduler or DPMSolverMultistepScheduler. The pipeline maps these schedulers onto the Cosmos 3 MoT backbone while maintaining the model's multi-modal tokenization capabilities.

How do I access the reasoner model when using the Diffusers pipeline?

The Diffusers integration loads the full checkpoint including reasoner weights, giving you direct access to the reasoner component for conditioning generation on image captions, embeddings, or action trajectories. You can call the reasoner between generation steps in a single Python script to implement closed-loop research workflows that iterate between generation and reasoning.

Where can I find executable examples for the Diffusers integration?

The repository includes the Jupyter notebook cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb, which provides step-by-step instructions for environment setup, kernel registration, and running text-to-image and text-to-video examples. This notebook also demonstrates how to verify that the Cosmos3OmniDiffusersPipeline loads correctly on your hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →