# How to Run the Cosmos 3 Reasoner with vLLM for Video Understanding

> Learn how to run the Cosmos 3 Reasoner with vLLM for efficient video understanding. Get an OpenAI-compatible API for multimodal inference with video URLs.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-05

---

**The NVIDIA Cosmos 3 Reasoner can be served efficiently through vLLM, exposing an OpenAI-compatible REST API that accepts video URLs and frame-sampling parameters for multimodal inference.**

The NVIDIA `cosmos` repository provides a production-ready pipeline for video understanding using the Cosmos 3 Reasoner family of models. By leveraging vLLM as the inference backend, you can run both the Nano and Super checkpoints with standard HTTP requests. This guide walks through the exact steps to launch the server and issue video understanding queries from a Python client.

## Architecture Overview

The deployment stack consists of three layers:

- **vLLM server** — Loads the Cosmos 3 checkpoint (e.g., `nvidia/cosmos3-super`) and listens on a local port. `8000` is used for Nano, while `8001` is commonly used for Super.
- **OpenAI-compatible client** — A thin Python wrapper built on the `openai` library that constructs chat requests containing a `video_url`, text prompt, and optional `media_io_kwargs`.
- **Media handling** — vLLM downloads or decodes the video, extracts frames according to the supplied sampling configuration, runs the visual encoder, and streams back a text response.

The request flows from the user client to the `vllm.entrypoint.api_server`, through the multimodal encoder, and finally to the LLM before returning captions or temporal reasoning results.

## Environment Setup

Before launching the server, install vLLM and any remaining Cosmos 3 dependencies as described in the environment guide. The official cookbook at [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) contains the full walkthrough for preparing your Python environment.

You will also need the `openai` package in the client environment:

```bash
pip install openai

```

## Launching the vLLM Server

### Single-GPU Nano Setup

For smaller workloads or development, launch the Nano checkpoint on a single GPU:

```bash
CUDA_VISIBLE_DEVICES=0 \
python -m vllm.entrypoint.api_server \
    --model nvidia/cosmos3-nano \
    --port 8000

```

### Multi-GPU Super Setup

For maximum reasoning quality, serve the Super checkpoint with tensor parallelism. As implemented in [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md), allocate four GPUs and bind to port `8001`:

```bash
CUDA_VISIBLE_DEVICES=0,1,2,3 \
python -m vllm.entrypoint.api_server \
    --model nvidia/cosmos3-super \
    --port 8001 \
    --tensor-parallel-size 4

```

Wait until the server reports that the model weights are loaded and the HTTP endpoint is ready.

## Running Video Understanding Requests via vLLM

### Minimal Python Client Example

The repository notebook at `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` demonstrates the canonical client pattern. Below is a self-contained script that queries the local vLLM server:

```python
import openai

video_url = (
    "https://github.com/nvidia-cosmos/cosmos-dependencies/raw/refs/heads/assets/"
    "cosmos3/inputs/video/temporal_localization_1.mp4"
)

client = openai.OpenAI(
    api_key="EMPTY",
    base_url="http://localhost:8001/v1"
)

response = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": video_url}},
                {"type": "text", "text": "Describe what is happening in this video."},
            ],
        }
    ],
    max_tokens=1024,
    temperature=0.0,
)

print(response.choices[0].message.content)

```

The `api_key` is set to `"EMPTY"` because vLLM does not enforce authentication by default. The client automatically discovers the loaded checkpoint via `client.models.list().data[0].id`.

### Controlling Frame Sampling with media_io_kwargs

You can adjust how many frames are extracted by passing `media_io_kwargs` inside the request's `extra_body`. This is useful for trading off inference speed against temporal granularity:

```python
response = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": video_url}},
                {"type": "text", "text": "Describe what is happening in this video."},
            ],
        }
    ],
    max_tokens=1024,
    temperature=0.0,
    extra_body={
        "media_io_kwargs": {
            "video": {
                "fps": 4.0,
                "max_frames": 32,
                "start_time": 0.0,
                "end_time": 10.0,
            }
        }
    },
)

```

vLLM uses these parameters during multimodal preprocessing to subsample the video before feeding frames to the Cosmos 3 visual encoder.

## Advanced Configuration

### Switching Checkpoints

Replace `--model nvidia/cosmos3-super` with `nvidia/cosmos3-nano` or a local path to a fine-tuned checkpoint. Update `base_url` in the client to match the new server port.

### Batch Video Processing

Wrap the `client.chat.completions.create` call in a loop over multiple `video_url` values. Each request is handled independently by the vLLM asynchronous engine, so throughput scales with server concurrency limits.

## Summary

- **vLLM** exposes the Cosmos 3 Reasoner as an OpenAI-compatible REST API on ports `8000` (Nano) or `8001` (Super).
- **Python clients** issue standard `chat.completions.create` calls with `video_url` payloads and text prompts.
- **Frame sampling** is controlled through `extra_body={"media_io_kwargs": {"video": {...}}}`.
- **Tensor parallelism** via `--tensor-parallel-size` enables serving the Super model across multiple GPUs.

## Frequently Asked Questions

### What port does the Cosmos 3 Reasoner vLLM server use by default?

Port `8000` is typical for the Nano checkpoint, while port `8001` is commonly assigned to the larger Super checkpoint in the official cookbook. You can override either with the `--port` flag when launching `vllm.entrypoint.api_server`.

### How do I control video frame sampling when using vLLM?

Pass a `media_io_kwargs` dictionary inside the request's `extra_body` field. Keys such as `fps`, `max_frames`, `start_time`, and `end_time` inside the `"video"` nested dict let you subsample or trim the input before the visual encoder processes it.

### Can I use the same client code for both image and video inputs?

Yes. The request structure is identical; simply change the content type from `image_url` to `video_url` in the message payload. The server handles modality-specific preprocessing automatically.

### What hardware is required to run the Cosmos 3 Super model?

The Super checkpoint is designed for multi-GPU inference. The reference command in [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md) uses `--tensor-parallel-size 4` across four GPUs. The Nano checkpoint can run on a single GPU.