# How to Use Cosmos 3 for Text-to-Video Generation with Synchronized Audio Output

> Learn to use Cosmos 3 for text-to-video generation with synchronized audio. Discover the Mixture-of-Transformers architecture and how to enable sound via Cosmos3OmniPipeline or vLLM-Omni API.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-13

---

**Cosmos 3 generates synchronized audio and video through a unified Mixture-of-Transformers architecture that processes visual and audio tokens in the same diffusion steps, accessible via the `Cosmos3OmniPipeline` with the `enable_sound=True` flag or the vLLM-Omni API with `generate_sound=true`.**

The NVIDIA Cosmos repository provides a state-of-the-art foundation for multimodal video generation. Cosmos 3's text-to-video generation with synchronized audio output leverages a novel **Mixture-of-Transformers (MoT)** architecture to produce temporally coherent visual frames and matching audio streams in a single inference pass. This guide demonstrates how to implement both local Diffusers-based inference and production-scale API deployment using the official cookbooks and source implementations.

## Understanding the Cosmos 3 Architecture

Cosmos 3 implements a unified **Mixture-of-Transformers** architecture that combines an autoregressive transformer for reasoning with a diffusion transformer for multimodal generation. As illustrated in `cookbooks/cosmos3/cosmos3-model-architecture.png`, the model treats video frames, audio samples, and optional conditioning modalities as a single sequence of tokens.

The diffusion transformer denoises this unified sequence, producing temporally coherent visual frames **and** a matching audio stream simultaneously. This architectural approach differs from traditional pipelines that generate video and audio separately, ensuring perfect synchronization between sound and motion without post-processing alignment steps.

## Prerequisites and Installation

Before running inference, install the required Python stack from the NVIDIA-maintained Diffusers fork. The official setup instructions in [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) specify the following environment configuration:

```bash

# Create a managed Python environment

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

# Install dependencies with automatic CUDA backend detection

uv pip install --torch-backend=auto \
  "diffusers @ git+https://github.com/huggingface/diffusers.git" \
  accelerate av torch torchvision transformers

```

The **Av** library is essential for audio codec handling, while `accelerate` enables efficient GPU memory management through `device_map="cuda"`.

## Method 1: Local Diffusers Pipeline

For research and development workflows, the `Cosmos3OmniPipeline` provides direct access to the full model checkpoint. The reference implementation in `cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb` demonstrates the complete end-to-end workflow.

### Loading the Pipeline

Initialize the pipeline with **bfloat16** precision and automatic device mapping to optimize memory utilization:

```python
import torch
from diffusers import Cosmos3OmniPipeline

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

```

### Configuring the Scheduler

Replace the default scheduler with **UniPCMultistepScheduler** to improve diffusion efficiency. The `flow_shift` parameter maintains smooth motion continuity across frames:

```python
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

```

### Generating Video with Audio

Set `enable_sound=True` to activate audio generation. The model returns a result object containing video frames and an embedded AAC audio stream:

```python
result = pipe(
    prompt="A small warehouse robot moves a blue box across a clean floor.",
    negative_prompt="blurry, low-quality",
    num_frames=189,            # Approximately 7.9 seconds at 24 FPS

    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    enable_sound=True,         # Enable synchronized audio generation

    generator=torch.Generator(device="cuda").manual_seed(1234),
)

# Export to MP4 with audio muxed

from diffusers.utils import export_to_video
export_to_video(
    result.video, 
    "cosmos3_text2video_with_sound.mp4", 
    fps=24, 
    macro_block_size=1
)

```

## Method 2: vLLM-Omni API for Production

For scalable deployment, the vLLM-Omni server exposes an OpenAI-compatible endpoint. As documented in `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb`, add `generate_sound=true` to the request JSON:

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=A bustling city street at night with neon signs" \
  --form-string "negative_prompt=blur, low quality" \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string "flow_shift=10.0" \
  --form-string "seed=0" \
  --form-string 'extra_params={"generate_sound":true,"use_resolution_template":false,"use_duration_template":false}' \
  -o cosmos3_t2v_sound.mp4

```

The server returns an MP4 container with synchronized audio and video ready for immediate playback.

## Key Parameters for Audio-Video Synchronization

Several configuration parameters control the quality and characteristics of the generated output:

- **`enable_sound`** (Pipeline) / **`generate_sound`** (API): Boolean flags that activate the audio generation pathway. When enabled, the model allocates tokens for audio samples within the unified diffusion sequence.
- **`flow_shift`**: Configures the scheduler's temporal bias. Values between **8.0** and **12.0** typically produce optimal motion smoothness for 24 FPS output.
- **`num_frames`**: Must align with your target duration and FPS. For 24 FPS video, 189 frames equals approximately 7.9 seconds of content.
- **`macro_block_size`**: Set to **1** in `export_to_video` to ensure compatibility with the AAC audio muxing process.

## Summary

- **Cosmos 3** uses a unified **Mixture-of-Transformers** architecture to generate video and audio tokens in the same diffusion steps, ensuring perfect synchronization.
- The **`Cosmos3OmniPipeline`** from the NVIDIA Diffusers fork provides local inference with the `enable_sound=True` flag.
- The **vLLM-Omni API** offers production-scale deployment via the `/v1/videos/sync` endpoint with `generate_sound=true`.
- Reference implementations are available in `cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb` and `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb`.
- Output files use **AAC** audio codec muxed into standard **MP4** containers.

## Frequently Asked Questions

### Does Cosmos 3 generate audio in the same diffusion process as video?

Yes. According to the architecture implementation in `cookbooks/cosmos3/cosmos3-model-architecture.png`, Cosmos 3 processes video frames and audio samples as a single token sequence within the diffusion transformer. The denoising steps simultaneously refine both modalities, eliminating the need for separate audio generation and post-hoc synchronization.

### What audio format does Cosmos 3 output?

Cosmos 3 generates **AAC** (Advanced Audio Coding) audio tracks. When using the Diffusers pipeline with `enable_sound=True`, the `export_to_video` utility automatically muxes the audio into an MP4 container. The vLLM-Omni API returns pre-muxed MP4 files containing the same AAC codec.

### How do I adjust the audio quality or characteristics?

Audio characteristics are primarily controlled through the text prompt and the **`guidance_scale`** parameter. Higher guidance scales (6.0–7.0) strengthen adherence to sonic descriptions in your prompt. The `flow_shift` scheduler parameter also affects temporal coherence between audio events and visual motion. Currently, there is no separate audio-specific quality parameter exposed in the API.

### Can I use Cosmos 3 with other conditioning modalities besides text?

Yes. The `Cosmos3OmniPipeline` accepts optional conditioning images and action sequences alongside text prompts. Set the `image` parameter to provide visual conditioning, or use the vLLM-Omni API's `extra_params` to include action conditioning tokens. The MoT architecture handles these additional modalities through the same unified token sequence that processes audio and video.