How to Set Up vLLM-Omni for Video Generation with Cosmos 3

You can generate video, audio, and images with Cosmos 3 by running the official vllm/vllm-omni:cosmos3 Docker container and submitting JSON payloads to its OpenAI-compatible API.

The NVIDIA cosmos repository provides a complete vLLM-Omni backend that exposes the Cosmos 3 model through a standard HTTP interface. By following the official cookbooks in cookbooks/cosmos3/generator/audiovisual, you can launch a local server, construct request payloads, and retrieve synthesized MP4 videos from either text or image inputs.

Install and Launch the vLLM-Omni Server

The recommended deployment path uses the pre-built Docker image. You must set three environment variables before launching the container: HF_HOME for Hugging Face cache, COSMOS3_WORKDIR for your mounted workspace, and COSMOS3_HOST_PORT for the exposed host port.

Run Cosmos 3 Nano on a Single GPU

For single-GPU inference with the Cosmos 3 Nano model, pull the image and start the container with the NVIDIA runtime:

docker pull vllm/vllm-omni:cosmos3

export HF_HOME="${HOME}/.cache/huggingface"
export COSMOS3_WORKDIR="$(pwd)"
export COSMOS3_HOST_PORT=8000

docker run --runtime nvidia --gpus '"device=0"' \
  -e CUDA_DEVICE_ORDER=PCI_BUS_ID \
  -v "${HF_HOME}:/root/.cache/huggingface" \
  -v "${COSMOS3_WORKDIR}:/workspace" \
  -p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

The container logs Application startup complete. once the /v1/models endpoint is ready. Verify the health of your vLLM-Omni server with:

curl http://localhost:8000/v1/models

Run Cosmos 3 Super on Multiple GPUs

For the larger Cosmos 3 Super model, add tensor parallelism and layerwise offloading across all available GPUs:

docker run --runtime nvidia --gpus all \
  -e HF_HOME="${HF_HOME}" \
  -v "${COSMOS3_WORKDIR}:/workspace" \
  -p "${COSMOS3_HOST_PORT}:8000" --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Super \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --port 8000 \
    --init-timeout 1800

Both variants require the --omni flag and the exact --model-class-name Cosmos3OmniDiffusersPipeline argument to activate the multimodal pipeline inside vLLM-Omni.

Build the Request Payload

With the server running, the next step is to construct the JSON body that vLLM-Omni expects. The notebook cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb implements a helper function named create_payload() that automates this process.

The helper performs the following actions:

  • Loads a human-readable prompt JSON from cookbooks/cosmos3/generator/audiovisual/assets/prompts/...
  • Injects negative prompts and fixed sampling parameters such as steps, guidance scale, FPS, and resolution
  • Sets vision_path for image-to-video workflows
  • Writes the finalized JSON to outputs/notebooks/vllm/payloads/<use-case>.json

You can generate a payload for the text-to-video-with-audio use case (t2vs) with:

from run_with_vllm_omni import create_payload

payload_path, output_dir, model_name = create_payload(
    "t2vs",
    backend="vllm"
)

print("Payload JSON:", payload_path)
print("Output folder:", output_dir)
print("Model:", model_name)

The notebook resolves server URLs using the environment variables COSMOS3_VLLM_BASE_URL, COSMOS3_VLLM_NANO_BASE_URL, and COSMOS3_VLLM_SUPER_BASE_URL, falling back to http://localhost:8000 when these are unset. This mapping is defined in lines 31–40 of run_with_vllm_omni.ipynb.

Submit the Request and Retrieve Video

The run_with_vllm_omni.ipynb notebook wraps the actual HTTP call in a post_video() routine that streams an MP4 response back to disk. Internally, it targets the synchronous video endpoint at http://localhost:8000/v1/videos/sync.

A simplified version of the submission logic looks like this:

from pathlib import Path
import subprocess

def post_video(*, payload_path: Path, payload: dict, output_path: Path, model: str):
    url = "http://localhost:8000/v1/videos/sync"
    cmd = [
        "curl", "-sS", "--fail-with-body", "-X", "POST", url,
        "-H", "Accept: video/mp4"
    ]
    # Attach payload fields as form strings

    for k, v in build_vllm_form(payload).items():
        cmd += ["--form-string", f"{k}={v}"]
    # Attach reference image for image-to-video generation

    if payload["model_mode"] == "image2video":
        img_path = resolve_payload_path(payload_path, payload["vision_path"])
        cmd += ["-F", f"input_reference=@{img_path}"]
    # Stream to a temporary file and atomically rename

    tmp_path = output_path.with_suffix(".tmp")
    subprocess.run(cmd + ["-o", str(tmp_path)], check=True)
    tmp_path.replace(output_path)

After the file is written, the notebook displays it inline with view_run(output_dir):

view_run(output_dir)

Full End-to-End Example

The following complete workflow demonstrates text-to-video generation with audio using Cosmos 3 Nano and the vLLM-Omni backend:

import os
os.environ["COSMOS3_VLLM_NANO_BASE_URL"] = "http://localhost:8000"

# 1. Build the payload

payload_path, out_dir, model = create_payload("t2vs", backend="vllm")

# 2. Run generation

run_vllm_payload(payload_path, out_dir, model="Cosmos3-Nano")

# 3. Visualize the result

view_run(out_dir)

Executing these cells produces an MP4 file—for example, a robot pouring water with synchronized audio. You can replace "t2vs" with any identifier defined in the notebook’s ASSET_SETS, such as "i2v_nano_noaudio" for image-to-video without sound.

Prompt assets are stored under cookbooks/cosmos3/generator/audiovisual/assets/prompts/. The file text2video/robot_pouring_water_audio.json supplies the default t2vs prompt, while image2video/car_driving.json shows how a vision_path is injected for image-conditioned generation.

Summary

  • Pull the official vllm/vllm-omni:cosmos3 image to obtain the vLLM-Omni backend bundled with Cosmos 3 checkpoints.
  • Launch the container with --omni, --model-class-name Cosmos3OmniDiffusersPipeline, and --init-timeout 1800; use --tensor-parallel-size and --enable-layerwise-offload for the Super variant.
  • Verify the server is ready by querying http://localhost:8000/v1/models.
  • Use create_payload() from run_with_vllm_omni.ipynb to convert human-readable prompts into the JSON body expected by the API.
  • Post the payload to /v1/videos/sync and stream the returned MP4 to disk with post_video().
  • Display completed generations inside the notebook via view_run().

Frequently Asked Questions

What Docker image do I need for vLLM-Omni with Cosmos 3?

Use the official vllm/vllm-omni:cosmos3 image as implemented in the NVIDIA cosmos repository. It ships with vLLM-Omni and the Cosmos 3 model weights pre-configured, so you do not need to build a native virtual environment.

How do I switch between Cosmos 3 Nano and Cosmos 3 Super?

Set the model name in the vllm serve command to nvidia/Cosmos3-Nano for single-GPU inference or nvidia/Cosmos3-Super for multi-GPU inference. For Super, append --tensor-parallel-size 4 and --enable-layerwise-offload to distribute layers across all visible GPUs.

Can I generate video from an image instead of text?

Yes. Select an image-to-video identifier from ASSET_SETS such as "i2v_nano_noaudio", ensure the prompt JSON contains a vision_path, and let post_video() attach the reference image via the input_reference form field. The server then performs image-conditioned generation through the same /v1/videos/sync endpoint.

Why does my container need --init-timeout 1800?

Cosmos 3 checkpoints are large, and vLLM-Omni requires substantial time to load weights, compile CUDA graphs, and initialize the diffusers pipeline. The --init-timeout 1800 flag gives the server a 30-minute window to complete startup before the orchestrator marks it unhealthy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →