How to Enable Sound Generation with Video Output in NVIDIA Cosmos
To enable sound generation in NVIDIA Cosmos, set enable_sound=True in the Diffusers pipeline or generate_sound=true in vLLM-Omni API requests, ensuring you use a Cosmos 3 Nano or Super checkpoint that includes the AVAE audio tokenizer.
NVIDIA Cosmos 3 generates synchronized audio-visual content by integrating an audio tokenizer (AVAE) into its diffusion transformer pipeline. This guide explains how to enable sound generation with video output in Cosmos using three different interfaces: the Diffusers Python library, the vLLM-Omni API, and the native Cosmos Framework CLI, based on the actual implementation in the nvidia/cosmos repository.
Architecture of Audio Generation in Cosmos
Cosmos 3 generates synchronized audio through a multimodal latent diffusion process that jointly denoises video and audio tokens.
The AVAE Audio Tokenizer
The AVAE (Audio Variational Autoencoder) tokenizer converts raw audio into compact latent representations and back again. According to the source code, the audio tokenizer uses the checkpoint at pretrained/tokenizers/audio/avae/avae_48k_noncausal_25hz_64ch.ckpt with default parameters of sound_dim = 64 and sound_latent_fps = 25. This produces a compressed audio latent sequence that the diffusion model processes alongside video frames.
Mixture-of-Transformers Integration
The Mixture-of-Transformers (MoT) diffusion transformer receives a concatenated token stream containing both video latents (from the vision tokenizer) and audio latents (from the AVAE interface). When the enable_sound flag is active, the pipeline decodes the audio latent sequence after denoising and muxes the resulting 48 kHz stereo AAC stream into the final MP4 output.
How to Enable Sound Generation in Diffusers
The Cosmos3OmniPipeline class in the Diffusers library provides the most direct way to generate video with synchronized audio. The pipeline supports the enable_sound parameter, which defaults to False in the notebook configuration at line 560 of cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb.
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers import UniPCMultistepScheduler
from diffusers.utils import export_to_video
# Load a checkpoint that supports audio (Cosmos3-Nano or Cosmos3-Super)
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
# Configure the scheduler (recommended)
pipe.scheduler = UniPCMultistepScheduler.from_config(
pipe.scheduler.config, flow_shift=10.0
)
# Enable sound generation
result = pipe(
prompt="A small drone flies over a forest and the wind whistles past the rotors.",
num_frames=189,
height=720,
width=1280,
fps=24,
num_inference_steps=35,
guidance_scale=6.0,
enable_sound=True, # Activates audio generation
generator=torch.Generator(device="cuda").manual_seed(42),
)
# Export MP4 with embedded AAC sound
export_to_video(
result.video,
"drone_flight_with_sound.mp4",
fps=24,
macro_block_size=1,
)
When enable_sound=True, the pipeline returns a result.sound object containing a tensor of shape [channels, timesteps], which export_to_video automatically muxes into the MP4 container.
How to Enable Sound Generation via vLLM-Omni API
For production deployments using the vLLM-Omni server, you can enable audio generation by including the generate_sound parameter in your request JSON. The endpoint description at lines 446-447 of cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb documents this flag.
curl -sS -X POST http://localhost:8000/v1/videos/sync \
--form-string "prompt=An underwater robot explores a coral reef while bubbly sounds echo." \
--form-string "size=1280x720" \
--form-string "num_frames=189" \
--form-string "fps=24" \
--form-string "num_inference_steps=35" \
--form-string "guidance_scale=6.0" \
--form-string "generate_sound=true" \
-o reef_exploration.mp4
The vLLM-Omni server processes the generate_sound=true flag identically to the Diffusers pipeline, invoking the AVAE decoder to produce the synchronized audio track before returning the complete MP4 file.
How to Enable Sound Generation in Cosmos Framework CLI
If you are using the native Cosmos Framework entry point, pass the --enable-sound flag to the inference script. This flag propagates to the underlying Diffusers pipeline internally.
cosmos_framework run \
--model nvidia/Cosmos3-Nano \
--prompt "A bustling city street at night with distant traffic noise." \
--output video.mp4 \
--enable-sound
This approach is functionally equivalent to setting enable_sound=True in Python, as both use the same underlying Cosmos3OmniPipeline implementation.
Audio Configuration Parameters
While the defaults produce broadcast-quality 48 kHz stereo audio, you can optionally adjust these parameters when configuring the pipeline:
- sound_dim: Latent dimension for audio tokens (default: 64)
- sound_latent_fps: Frame rate for audio latents (default: 25 Hz)
- Audio codec: AAC at 48 kHz stereo (fixed in the current implementation)
These settings are defined in the AVAE checkpoint configuration and referenced in the repository's README.md under the "Sound output" specification.
Summary
- Use an audio-capable checkpoint: Only Cosmos 3 Nano and Super checkpoints include the AVAE tokenizer required for sound generation.
- Activate the flag: Set
enable_sound=Truein Python,generate_sound=truein HTTP requests, or--enable-soundin CLI commands. - Output format: The pipeline automatically muxes 48 kHz stereo AAC audio into the MP4 container alongside the video frames.
- Source verification: Implementation details are found in
cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynband the AVAE checkpoint atpretrained/tokenizers/audio/avae/avae_48k_noncausal_25hz_64ch.ckpt.
Frequently Asked Questions
What model checkpoints support sound generation in Cosmos?
Only the Cosmos 3 Nano and Cosmos 3 Super checkpoints include the AVAE audio tokenizer module required for sound generation. These checkpoints reference the specific tokenizer weights at pretrained/tokenizers/audio/avae/avae_48k_noncausal_25hz_64ch.ckpt. Standard Cosmos 3 checkpoints without the Omni designation do not support audio output.
What audio format does Cosmos output when sound is enabled?
Cosmos outputs 48 kHz stereo AAC audio muxed into the MP4 container alongside the video stream. This is documented in the "Sound output" table in the repository's README.md. The audio tokenizer produces this format through the AVAE decoder, which converts the denoised audio latents back to raw waveforms at the specified sample rate.
Can I adjust the audio quality or sampling rate?
While the output format is fixed at 48 kHz AAC, you can modify the intermediate latent parameters sound_dim (default 64) and sound_latent_fps (default 25) if using a custom AVAE configuration. However, the standard checkpoints are optimized for these specific values, and altering them requires training or fine-tuning the audio tokenizer component.
Does enabling sound generation affect video generation speed?
Yes, enabling sound generation increases inference time because the Mixture-of-Transformers must jointly denoise both video and audio token streams, and the pipeline must additionally run the AVAE decoder to convert audio latents back to waveforms. The overhead is proportional to the audio sequence length, which scales with video duration at 25 latent frames per second.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →