# Setting Up Cosmos 3 Reasoner with vLLM for Production Inference

> Deploy Cosmos 3 Reasoner with vLLM for production inference. Install the vllm-cosmos3 plugin, override the vLLM architecture, and configure I/O paths for seamless deployment.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**You can deploy Cosmos 3 Reasoner as an OpenAI-compatible chat-completion endpoint by installing the `vllm-cosmos3` plugin, launching vLLM with the `Cosmos3ReasonerForConditionalGeneration` architecture override, and configuring multimodal I/O paths for production traffic.**

The NVIDIA Cosmos repository provides a production-ready pathway for serving Cosmos 3 Reasoner models using vLLM. By setting up Cosmos 3 Reasoner with vLLM for production inference, you unlock an OpenAI-compatible API capable of processing images and video streams with structured text generation. This implementation leverages the `Cosmos3ReasonerForConditionalGeneration` architecture registered through the `vllm-cosmos3` plugin to handle multimodal workloads efficiently.

## Install the Correct vLLM Build

Production deployment begins with an isolated Python environment and a CUDA-matched vLLM installation. The `vllm-cosmos3` plugin must be installed alongside vLLM to register the Cosmos 3 Reasoner architecture.

Create a clean virtual environment and install the dependencies:

```bash

# Create and activate a Python 3.13 environment

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

# Install CUDA-matched vLLM and the Cosmos-3 plugin

# CUDA 13 → cu130 + vllm==0.21.0

# CUDA 12.8 → cu128 + vllm==0.19.1

uv pip install --torch-backend=cu130 "vllm==0.21.0" \
  "vllm-cosmos3 @ git+https://github.com/NVIDIA/cosmos-framework.git#subdirectory=packages/vllm-cosmos3"

```

**Critical:** vLLM wheels are compiled for specific CUDA versions. Mismatching the driver results in `torch.cuda.is_available() == False`. The version pairings above are reproduced from the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 38-40). If your driver is older than CUDA 12.8, substitute `--torch-backend=cu128` with the appropriate legacy tag and use the matching vLLM version.

## Launch the vLLM Reasoner Server

With dependencies installed, launch the server using the `vllm serve` command with architecture overrides and multimodal configuration flags.

First, clone the Cosmos framework repository to access local media paths:

```bash
COSMOS3_REPO=$HOME/cosmos-framework
git clone https://github.com/NVIDIA/cosmos-framework.git $COSMOS3_REPO

```

Start the server for the **Nano** model on a single GPU:

```bash
VLLM_PORT=8000
COSMOS3_MEDIA_ROOT=$(pwd)/cookbooks/cosmos3   # Root for local file:// media paths

setsid .venv/bin/vllm serve nvidia/Cosmos3-Nano \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --media-io-kwargs '{"video": {"num_frames": -1}}' \
  --port $VLLM_PORT > vllm_server.log 2>&1 &
disown
echo "vLLM server started on port $VLLM_PORT (log → vllm_server.log)"

```

**Key flags explained:**
- **`--hf-overrides`**: Exposes the `Cosmos3ReasonerForConditionalGeneration` architecture required by the model checkpoint.
- **`--async-scheduling`**: Enables asynchronous request processing for higher throughput.
- **`--allowed-local-media-path`**: Restricts file system access to the specified directory for security.
- **`--media-io-kwargs`**: Configures video processing; `num_frames: -1` processes full-length videos.

For the **Super** model on multiple GPUs, add tensor parallelism:

```bash
CUDA_VISIBLE_DEVICES=0,1,2,3 .venv/bin/vllm serve nvidia/Cosmos3-Super \
  --tensor-parallel-size 4 \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --port 8000

```

**Guardrails configuration:** To disable optional guardrail models, export a deploy config as shown in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 403-417) and pass `--deploy-config no_guardrails.yaml` to the serve command.

## Query the Reasoner with Production Clients

The server exposes an OpenAI-compatible `POST /v1/chat/completions` endpoint. Use standard HTTP clients or the OpenAI Python SDK to interact with the model.

### Using curl for Image Inference

Send base64-encoded images or HTTPS URLs:

```bash
IMAGE_URI="data:image/jpeg;base64,$(base64 -w 0 cookbooks/cosmos3/reasoner/assets/robot_153.jpg)"

curl -X POST http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "nvidia/Cosmos3-Nano",
        "messages": [
          {"role":"system","content":"You are a helpful assistant."},
          {"role":"user","content":[
            {"type":"image_url","image_url":{"url":"'"$IMAGE_URI"'"}},
            {"type":"text","text":"Describe what is happening in this image in one sentence."}
          ]}
        ],
        "max_tokens":256,
        "stream":false
      }'

```

### Using the OpenAI Python Client for Video

Process video URLs with custom frame sampling rates:

```python
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")

response = client.chat.completions.create(
    model="nvidia/Cosmos3-Nano",
    messages=[
        {"role":"system","content":"You are a helpful assistant."},
        {"role":"user","content":[
            {"type":"video_url","video_url":{"url":"https://download.samplelib.com/mp4/sample-5s.mp4"}},
            {"type":"text","text":"List the notable events with approximate timestamps."}
        ]}
    ],
    max_tokens=256,
    stream=False,
    extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}}
)

print(response.choices[0].message.content)

```

**Production scaling:** Wrap the HTTP endpoint behind a load balancer (NGINX, Envoy, or cloud-native Ingress) and enable health checks via `GET /v1/health`. Deploy multiple vLLM pods on distinct ports and aggregate traffic using a reverse proxy to handle high concurrency loads.

## Key Repository References

The implementation details are derived from the following source files in the NVIDIA Cosmos repository:

- **`cookbooks/cosmos3/reasoner/run_with_vllm.ipynb`**: Complete notebook automating environment setup, server launch, and client calls (see lines 166-172 for multi-GPU configuration).
- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (Reasoner with vLLM section)**: Contains the version matrix, guardrail configuration templates (lines 403-417), and example requests (lines 521-638).
- **`packages/vllm-cosmos3`**: Plugin directory in the `cosmos-framework` submodule that registers the `Cosmos3ReasonerForConditionalGeneration` architecture.
- **`cookbooks/cosmos3/reasoner/assets/robot_153.jpg`**: Sample image used for testing local file path resolutions.

## Summary

- **Match CUDA and vLLM versions** exactly: use `--torch-backend=cu130` with `vllm==0.21.0` for CUDA 13, or `--torch-backend=cu128` with `vllm==0.19.1` for CUDA 12.8.
- **Install the `vllm-cosmos3` plugin** from the Cosmos-framework submodule to register the Reasoner architecture.
- **Launch with architecture overrides** using `--hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}'` and enable `--async-scheduling` for throughput.
- **Configure media paths** with `--allowed-local-media-path` and tune video processing via `--media-io-kwargs`.
- **Deploy behind a reverse proxy** with health checks and load balancing for production resilience.

## Frequently Asked Questions

### Which CUDA version should I use with vLLM for Cosmos 3 Reasoner?

You must pair your CUDA driver version with a specific vLLM wheel. For CUDA 13, install with `--torch-backend=cu130` and `vllm==0.21.0`. For CUDA 12.8, use `--torch-backend=cu128` and `vllm==0.19.1`. Mismatching these versions causes `torch.cuda.is_available()` to return `False`, preventing GPU initialization.

### How do I enable multi-GPU inference for Cosmos 3 Super?

Add the `--tensor-parallel-size 4` flag to the `vllm serve` command and set `CUDA_VISIBLE_DEVICES=0,1,2,3` to specify which GPUs to use. This configuration is documented in `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` (lines 166-172) and distributes the model across four GPUs for higher throughput.

### Can I disable guardrails in production deployments?

Yes. Export a deployment configuration file that disables guardrail models as shown in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 403-417), then pass `--deploy-config no_guardrails.yaml` when starting the server. This reduces latency by skipping the safety classification steps.

### What media formats does the OpenAI-compatible endpoint support?

The endpoint supports images via `image_url` content types (base64 or HTTPS URLs) and videos via `video_url` content types with configurable preprocessing. Use `--media-io-kwargs` to set `video.num_frames` (use `-1` for full-length) or pass `extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}}` in client requests to control frame sampling rates.