How to Use Cosmos 3 for Text-to-Video Generation with Synchronized Audio Output

Cosmos 3 generates synchronized audio and video through a unified Mixture-of-Transformers architecture that processes visual and audio tokens in the same diffusion steps, accessible via the Cosmos3OmniPipeline with the enable_sound=True flag or the vLLM-Omni API with generate_sound=true.

The NVIDIA Cosmos repository provides a state-of-the-art foundation for multimodal video generation. Cosmos 3's text-to-video generation with synchronized audio output leverages a novel Mixture-of-Transformers (MoT) architecture to produce temporally coherent visual frames and matching audio streams in a single inference pass. This guide demonstrates how to implement both local Diffusers-based inference and production-scale API deployment using the official cookbooks and source implementations.

Understanding the Cosmos 3 Architecture

Cosmos 3 implements a unified Mixture-of-Transformers architecture that combines an autoregressive transformer for reasoning with a diffusion transformer for multimodal generation. As illustrated in cookbooks/cosmos3/cosmos3-model-architecture.png, the model treats video frames, audio samples, and optional conditioning modalities as a single sequence of tokens.

The diffusion transformer denoises this unified sequence, producing temporally coherent visual frames and a matching audio stream simultaneously. This architectural approach differs from traditional pipelines that generate video and audio separately, ensuring perfect synchronization between sound and motion without post-processing alignment steps.

Prerequisites and Installation

Before running inference, install the required Python stack from the NVIDIA-maintained Diffusers fork. The official setup instructions in cookbooks/cosmos3/README.md specify the following environment configuration:


# Create a managed Python environment

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

# Install dependencies with automatic CUDA backend detection

uv pip install --torch-backend=auto \
  "diffusers @ git+https://github.com/huggingface/diffusers.git" \
  accelerate av torch torchvision transformers

The Av library is essential for audio codec handling, while accelerate enables efficient GPU memory management through device_map="cuda".

Method 1: Local Diffusers Pipeline

For research and development workflows, the Cosmos3OmniPipeline provides direct access to the full model checkpoint. The reference implementation in cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb demonstrates the complete end-to-end workflow.

Loading the Pipeline

Initialize the pipeline with bfloat16 precision and automatic device mapping to optimize memory utilization:

import torch
from diffusers import Cosmos3OmniPipeline

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

Configuring the Scheduler

Replace the default scheduler with UniPCMultistepScheduler to improve diffusion efficiency. The flow_shift parameter maintains smooth motion continuity across frames:

from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

Generating Video with Audio

Set enable_sound=True to activate audio generation. The model returns a result object containing video frames and an embedded AAC audio stream:

result = pipe(
    prompt="A small warehouse robot moves a blue box across a clean floor.",
    negative_prompt="blurry, low-quality",
    num_frames=189,            # Approximately 7.9 seconds at 24 FPS

    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    enable_sound=True,         # Enable synchronized audio generation

    generator=torch.Generator(device="cuda").manual_seed(1234),
)

# Export to MP4 with audio muxed

from diffusers.utils import export_to_video
export_to_video(
    result.video, 
    "cosmos3_text2video_with_sound.mp4", 
    fps=24, 
    macro_block_size=1
)

Method 2: vLLM-Omni API for Production

For scalable deployment, the vLLM-Omni server exposes an OpenAI-compatible endpoint. As documented in cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb, add generate_sound=true to the request JSON:

curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=A bustling city street at night with neon signs" \
  --form-string "negative_prompt=blur, low quality" \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string "flow_shift=10.0" \
  --form-string "seed=0" \
  --form-string 'extra_params={"generate_sound":true,"use_resolution_template":false,"use_duration_template":false}' \
  -o cosmos3_t2v_sound.mp4

The server returns an MP4 container with synchronized audio and video ready for immediate playback.

Key Parameters for Audio-Video Synchronization

Several configuration parameters control the quality and characteristics of the generated output:

  • enable_sound (Pipeline) / generate_sound (API): Boolean flags that activate the audio generation pathway. When enabled, the model allocates tokens for audio samples within the unified diffusion sequence.
  • flow_shift: Configures the scheduler's temporal bias. Values between 8.0 and 12.0 typically produce optimal motion smoothness for 24 FPS output.
  • num_frames: Must align with your target duration and FPS. For 24 FPS video, 189 frames equals approximately 7.9 seconds of content.
  • macro_block_size: Set to 1 in export_to_video to ensure compatibility with the AAC audio muxing process.

Summary

  • Cosmos 3 uses a unified Mixture-of-Transformers architecture to generate video and audio tokens in the same diffusion steps, ensuring perfect synchronization.
  • The Cosmos3OmniPipeline from the NVIDIA Diffusers fork provides local inference with the enable_sound=True flag.
  • The vLLM-Omni API offers production-scale deployment via the /v1/videos/sync endpoint with generate_sound=true.
  • Reference implementations are available in cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb and cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb.
  • Output files use AAC audio codec muxed into standard MP4 containers.

Frequently Asked Questions

Does Cosmos 3 generate audio in the same diffusion process as video?

Yes. According to the architecture implementation in cookbooks/cosmos3/cosmos3-model-architecture.png, Cosmos 3 processes video frames and audio samples as a single token sequence within the diffusion transformer. The denoising steps simultaneously refine both modalities, eliminating the need for separate audio generation and post-hoc synchronization.

What audio format does Cosmos 3 output?

Cosmos 3 generates AAC (Advanced Audio Coding) audio tracks. When using the Diffusers pipeline with enable_sound=True, the export_to_video utility automatically muxes the audio into an MP4 container. The vLLM-Omni API returns pre-muxed MP4 files containing the same AAC codec.

How do I adjust the audio quality or characteristics?

Audio characteristics are primarily controlled through the text prompt and the guidance_scale parameter. Higher guidance scales (6.0–7.0) strengthen adherence to sonic descriptions in your prompt. The flow_shift scheduler parameter also affects temporal coherence between audio events and visual motion. Currently, there is no separate audio-specific quality parameter exposed in the API.

Can I use Cosmos 3 with other conditioning modalities besides text?

Yes. The Cosmos3OmniPipeline accepts optional conditioning images and action sequences alongside text prompts. Set the image parameter to provide visual conditioning, or use the vLLM-Omni API's extra_params to include action conditioning tokens. The MoT architecture handles these additional modalities through the same unified token sequence that processes audio and video.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →