# How to Set Up vLLM-Omni for Video Generation with Cosmos 3

> Learn to set up vLLM-Omni for video generation with Cosmos 3. Utilize the official Docker container and its OpenAI-compatible API for seamless multimedia creation.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-05

---

**You can generate video, audio, and images with Cosmos 3 by running the official `vllm/vllm-omni:cosmos3` Docker container and submitting JSON payloads to its OpenAI-compatible API.**

The NVIDIA `cosmos` repository provides a complete vLLM-Omni backend that exposes the Cosmos 3 model through a standard HTTP interface. By following the official cookbooks in `cookbooks/cosmos3/generator/audiovisual`, you can launch a local server, construct request payloads, and retrieve synthesized MP4 videos from either text or image inputs.

## Install and Launch the vLLM-Omni Server

The recommended deployment path uses the pre-built Docker image. You must set three environment variables before launching the container: `HF_HOME` for Hugging Face cache, `COSMOS3_WORKDIR` for your mounted workspace, and `COSMOS3_HOST_PORT` for the exposed host port.

### Run Cosmos 3 Nano on a Single GPU

For single-GPU inference with the Cosmos 3 Nano model, pull the image and start the container with the NVIDIA runtime:

```bash
docker pull vllm/vllm-omni:cosmos3

export HF_HOME="${HOME}/.cache/huggingface"
export COSMOS3_WORKDIR="$(pwd)"
export COSMOS3_HOST_PORT=8000

docker run --runtime nvidia --gpus '"device=0"' \
  -e CUDA_DEVICE_ORDER=PCI_BUS_ID \
  -v "${HF_HOME}:/root/.cache/huggingface" \
  -v "${COSMOS3_WORKDIR}:/workspace" \
  -p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

```

The container logs `Application startup complete.` once the `/v1/models` endpoint is ready. Verify the health of your vLLM-Omni server with:

```bash
curl http://localhost:8000/v1/models

```

### Run Cosmos 3 Super on Multiple GPUs

For the larger Cosmos 3 Super model, add tensor parallelism and layerwise offloading across all available GPUs:

```bash
docker run --runtime nvidia --gpus all \
  -e HF_HOME="${HF_HOME}" \
  -v "${COSMOS3_WORKDIR}:/workspace" \
  -p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Super \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --port 8000 \
    --init-timeout 1800

```

Both variants require the `--omni` flag and the exact `--model-class-name Cosmos3OmniDiffusersPipeline` argument to activate the multimodal pipeline inside vLLM-Omni.

## Build the Request Payload

With the server running, the next step is to construct the JSON body that vLLM-Omni expects. The notebook `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb` implements a helper function named `create_payload()` that automates this process.

The helper performs the following actions:

- Loads a human-readable prompt JSON from `cookbooks/cosmos3/generator/audiovisual/assets/prompts/...`
- Injects negative prompts and fixed sampling parameters such as steps, guidance scale, FPS, and resolution
- Sets `vision_path` for image-to-video workflows
- Writes the finalized JSON to `outputs/notebooks/vllm/payloads/<use-case>.json`

You can generate a payload for the text-to-video-with-audio use case (`t2vs`) with:

```python
from run_with_vllm_omni import create_payload

payload_path, output_dir, model_name = create_payload(
    "t2vs",
    backend="vllm"
)

print("Payload JSON:", payload_path)
print("Output folder:", output_dir)
print("Model:", model_name)

```

The notebook resolves server URLs using the environment variables `COSMOS3_VLLM_BASE_URL`, `COSMOS3_VLLM_NANO_BASE_URL`, and `COSMOS3_VLLM_SUPER_BASE_URL`, falling back to `http://localhost:8000` when these are unset. This mapping is defined in lines 31–40 of `run_with_vllm_omni.ipynb`.

## Submit the Request and Retrieve Video

The `run_with_vllm_omni.ipynb` notebook wraps the actual HTTP call in a `post_video()` routine that streams an MP4 response back to disk. Internally, it targets the synchronous video endpoint at `http://localhost:8000/v1/videos/sync`.

A simplified version of the submission logic looks like this:

```python
from pathlib import Path
import subprocess

def post_video(*, payload_path: Path, payload: dict, output_path: Path, model: str):
    url = "http://localhost:8000/v1/videos/sync"
    cmd = [
        "curl", "-sS", "--fail-with-body", "-X", "POST", url,
        "-H", "Accept: video/mp4"
    ]
    # Attach payload fields as form strings

    for k, v in build_vllm_form(payload).items():
        cmd += ["--form-string", f"{k}={v}"]
    # Attach reference image for image-to-video generation

    if payload["model_mode"] == "image2video":
        img_path = resolve_payload_path(payload_path, payload["vision_path"])
        cmd += ["-F", f"input_reference=@{img_path}"]
    # Stream to a temporary file and atomically rename

    tmp_path = output_path.with_suffix(".tmp")
    subprocess.run(cmd + ["-o", str(tmp_path)], check=True)
    tmp_path.replace(output_path)

```

After the file is written, the notebook displays it inline with `view_run(output_dir)`:

```python
view_run(output_dir)

```

## Full End-to-End Example

The following complete workflow demonstrates text-to-video generation with audio using Cosmos 3 Nano and the vLLM-Omni backend:

```python
import os
os.environ["COSMOS3_VLLM_NANO_BASE_URL"] = "http://localhost:8000"

# 1. Build the payload

payload_path, out_dir, model = create_payload("t2vs", backend="vllm")

# 2. Run generation

run_vllm_payload(payload_path, out_dir, model="Cosmos3-Nano")

# 3. Visualize the result

view_run(out_dir)

```

Executing these cells produces an MP4 file—for example, a robot pouring water with synchronized audio. You can replace `"t2vs"` with any identifier defined in the notebook’s `ASSET_SETS`, such as `"i2v_nano_noaudio"` for image-to-video without sound.

Prompt assets are stored under `cookbooks/cosmos3/generator/audiovisual/assets/prompts/`. The file [`text2video/robot_pouring_water_audio.json`](https://github.com/NVIDIA/cosmos/blob/main/text2video/robot_pouring_water_audio.json) supplies the default `t2vs` prompt, while [`image2video/car_driving.json`](https://github.com/NVIDIA/cosmos/blob/main/image2video/car_driving.json) shows how a `vision_path` is injected for image-conditioned generation.

## Summary

- Pull the official `vllm/vllm-omni:cosmos3` image to obtain the vLLM-Omni backend bundled with Cosmos 3 checkpoints.
- Launch the container with `--omni`, `--model-class-name Cosmos3OmniDiffusersPipeline`, and `--init-timeout 1800`; use `--tensor-parallel-size` and `--enable-layerwise-offload` for the Super variant.
- Verify the server is ready by querying `http://localhost:8000/v1/models`.
- Use `create_payload()` from `run_with_vllm_omni.ipynb` to convert human-readable prompts into the JSON body expected by the API.
- Post the payload to `/v1/videos/sync` and stream the returned MP4 to disk with `post_video()`.
- Display completed generations inside the notebook via `view_run()`.

## Frequently Asked Questions

### What Docker image do I need for vLLM-Omni with Cosmos 3?

Use the official `vllm/vllm-omni:cosmos3` image as implemented in the NVIDIA `cosmos` repository. It ships with vLLM-Omni and the Cosmos 3 model weights pre-configured, so you do not need to build a native virtual environment.

### How do I switch between Cosmos 3 Nano and Cosmos 3 Super?

Set the model name in the `vllm serve` command to `nvidia/Cosmos3-Nano` for single-GPU inference or `nvidia/Cosmos3-Super` for multi-GPU inference. For Super, append `--tensor-parallel-size 4` and `--enable-layerwise-offload` to distribute layers across all visible GPUs.

### Can I generate video from an image instead of text?

Yes. Select an image-to-video identifier from `ASSET_SETS` such as `"i2v_nano_noaudio"`, ensure the prompt JSON contains a `vision_path`, and let `post_video()` attach the reference image via the `input_reference` form field. The server then performs image-conditioned generation through the same `/v1/videos/sync` endpoint.

### Why does my container need `--init-timeout 1800`?

Cosmos 3 checkpoints are large, and vLLM-Omni requires substantial time to load weights, compile CUDA graphs, and initialize the diffusers pipeline. The `--init-timeout 1800` flag gives the server a 30-minute window to complete startup before the orchestrator marks it unhealthy.