# How to Enable Sound Generation with Video Output in NVIDIA Cosmos

> Discover how to enable sound generation with video output in NVIDIA Cosmos. Learn to set enable_sound=True or generate_sound=true for dynamic multimedia experiences.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-12

---

**To enable sound generation in NVIDIA Cosmos, set `enable_sound=True` in the Diffusers pipeline or `generate_sound=true` in vLLM-Omni API requests, ensuring you use a Cosmos 3 Nano or Super checkpoint that includes the AVAE audio tokenizer.**

NVIDIA Cosmos 3 generates synchronized audio-visual content by integrating an **audio tokenizer (AVAE)** into its diffusion transformer pipeline. This guide explains how to enable sound generation with video output in Cosmos using three different interfaces: the Diffusers Python library, the vLLM-Omni API, and the native Cosmos Framework CLI, based on the actual implementation in the `nvidia/cosmos` repository.

## Architecture of Audio Generation in Cosmos

Cosmos 3 generates synchronized audio through a multimodal latent diffusion process that jointly denoises video and audio tokens.

### The AVAE Audio Tokenizer

The **AVAE (Audio Variational Autoencoder)** tokenizer converts raw audio into compact latent representations and back again. According to the source code, the audio tokenizer uses the checkpoint at `pretrained/tokenizers/audio/avae/avae_48k_noncausal_25hz_64ch.ckpt` with default parameters of `sound_dim = 64` and `sound_latent_fps = 25`. This produces a compressed audio latent sequence that the diffusion model processes alongside video frames.

### Mixture-of-Transformers Integration

The **Mixture-of-Transformers (MoT)** diffusion transformer receives a concatenated token stream containing both video latents (from the vision tokenizer) and audio latents (from the AVAE interface). When the `enable_sound` flag is active, the pipeline decodes the audio latent sequence after denoising and muxes the resulting **48 kHz stereo AAC** stream into the final MP4 output.

## How to Enable Sound Generation in Diffusers

The `Cosmos3OmniPipeline` class in the Diffusers library provides the most direct way to generate video with synchronized audio. The pipeline supports the `enable_sound` parameter, which defaults to `False` in the notebook configuration at line 560 of `cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb`.

```python
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers import UniPCMultistepScheduler
from diffusers.utils import export_to_video

# Load a checkpoint that supports audio (Cosmos3-Nano or Cosmos3-Super)

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

# Configure the scheduler (recommended)

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

# Enable sound generation

result = pipe(
    prompt="A small drone flies over a forest and the wind whistles past the rotors.",
    num_frames=189,
    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    enable_sound=True,          # Activates audio generation

    generator=torch.Generator(device="cuda").manual_seed(42),
)

# Export MP4 with embedded AAC sound

export_to_video(
    result.video,
    "drone_flight_with_sound.mp4",
    fps=24,
    macro_block_size=1,
)

```

When `enable_sound=True`, the pipeline returns a `result.sound` object containing a tensor of shape `[channels, timesteps]`, which `export_to_video` automatically muxes into the MP4 container.

## How to Enable Sound Generation via vLLM-Omni API

For production deployments using the vLLM-Omni server, you can enable audio generation by including the `generate_sound` parameter in your request JSON. The endpoint description at lines 446-447 of `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb` documents this flag.

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=An underwater robot explores a coral reef while bubbly sounds echo." \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string "generate_sound=true" \
  -o reef_exploration.mp4

```

The vLLM-Omni server processes the `generate_sound=true` flag identically to the Diffusers pipeline, invoking the AVAE decoder to produce the synchronized audio track before returning the complete MP4 file.

## How to Enable Sound Generation in Cosmos Framework CLI

If you are using the native Cosmos Framework entry point, pass the `--enable-sound` flag to the inference script. This flag propagates to the underlying Diffusers pipeline internally.

```bash
cosmos_framework run \
  --model nvidia/Cosmos3-Nano \
  --prompt "A bustling city street at night with distant traffic noise." \
  --output video.mp4 \
  --enable-sound

```

This approach is functionally equivalent to setting `enable_sound=True` in Python, as both use the same underlying `Cosmos3OmniPipeline` implementation.

## Audio Configuration Parameters

While the defaults produce broadcast-quality 48 kHz stereo audio, you can optionally adjust these parameters when configuring the pipeline:

- **sound_dim**: Latent dimension for audio tokens (default: 64)
- **sound_latent_fps**: Frame rate for audio latents (default: 25 Hz)
- **Audio codec**: AAC at 48 kHz stereo (fixed in the current implementation)

These settings are defined in the AVAE checkpoint configuration and referenced in the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) under the "Sound output" specification.

## Summary

- **Use an audio-capable checkpoint**: Only Cosmos 3 Nano and Super checkpoints include the AVAE tokenizer required for sound generation.
- **Activate the flag**: Set `enable_sound=True` in Python, `generate_sound=true` in HTTP requests, or `--enable-sound` in CLI commands.
- **Output format**: The pipeline automatically muxes 48 kHz stereo AAC audio into the MP4 container alongside the video frames.
- **Source verification**: Implementation details are found in `cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb` and the AVAE checkpoint at `pretrained/tokenizers/audio/avae/avae_48k_noncausal_25hz_64ch.ckpt`.

## Frequently Asked Questions

### What model checkpoints support sound generation in Cosmos?

Only the **Cosmos 3 Nano** and **Cosmos 3 Super** checkpoints include the AVAE audio tokenizer module required for sound generation. These checkpoints reference the specific tokenizer weights at `pretrained/tokenizers/audio/avae/avae_48k_noncausal_25hz_64ch.ckpt`. Standard Cosmos 3 checkpoints without the Omni designation do not support audio output.

### What audio format does Cosmos output when sound is enabled?

Cosmos outputs **48 kHz stereo AAC** audio muxed into the MP4 container alongside the video stream. This is documented in the "Sound output" table in the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md). The audio tokenizer produces this format through the AVAE decoder, which converts the denoised audio latents back to raw waveforms at the specified sample rate.

### Can I adjust the audio quality or sampling rate?

While the output format is fixed at 48 kHz AAC, you can modify the intermediate latent parameters `sound_dim` (default 64) and `sound_latent_fps` (default 25) if using a custom AVAE configuration. However, the standard checkpoints are optimized for these specific values, and altering them requires training or fine-tuning the audio tokenizer component.

### Does enabling sound generation affect video generation speed?

Yes, enabling sound generation increases inference time because the **Mixture-of-Transformers** must jointly denoise both video and audio token streams, and the pipeline must additionally run the AVAE decoder to convert audio latents back to waveforms. The overhead is proportional to the audio sequence length, which scales with video duration at 25 latent frames per second.