# How to Deploy Cosmos 3 Reasoner Using vLLM for Production Inference

> Deploy Cosmos 3 Reasoner with vLLM for production inference. Install the plugin, launch the server, and expose an OpenAI-compatible endpoint for efficient multimodal processing.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-13

---

**Deploy Cosmos 3 Reasoner using vLLM by installing the CUDA‑matched vLLM wheel alongside the `vllm‑cosmos3` plugin, then launching the server with architecture overrides and multimodal I/O flags to expose an OpenAI‑compatible chat completions endpoint.**

The NVIDIA Cosmos repository provides a production‑ready inference stack for the Cosmos 3 Reasoner family of multimodal language models. By leveraging the `vllm‑cosmos3` plugin, you can serve these models through a standard OpenAI‑compatible API, enabling seamless integration with existing client libraries and production infrastructure.

## Prerequisites and Environment Setup

Proper deployment begins with a CUDA‑matched Python environment and the correct plugin installation.

### CUDA Version Compatibility

vLLM wheels are compiled against specific CUDA versions; mismatched drivers will cause `torch.cuda.is_available()` to return `False`. The NVIDIA Cosmos documentation specifies the following pairings:

- **CUDA 13**: Use `--torch-backend=cu130` with `vllm==0.21.0`
- **CUDA 12.8**: Use `--torch-backend=cu128` with `vllm==0.19.1`

Verify your driver version before proceeding, as this determines which vLLM build you must install.

### Installing vLLM and the Cosmos3 Plugin

Create an isolated environment and install the dependencies from the `cosmos-framework` submodule:

```bash

# Create virtual environment with Python 3.13

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

# Install CUDA-matched vLLM and the Cosmos3 plugin

# Replace cu130/cu128 based on your driver

uv pip install --torch-backend=cu130 "vllm==0.21.0" \
  "vllm-cosmos3 @ git+https://github.com/NVIDIA/cosmos-framework.git#subdirectory=packages/vllm-cosmos3"

```

The `vllm‑cosmos3` package registers the `Cosmos3ReasonerForConditionalGeneration` architecture, which vLLM requires to load the model checkpoints correctly.

## Launching the vLLM Server

The server initialization differs based on your hardware configuration and whether you are running the Nano or Super variant.

### Single GPU Configuration (Nano)

For the `nvidia/Cosmos3-Nano` model on a single GPU, execute the following command from your repository root:

```bash
COSMOS3_MEDIA_ROOT=$(pwd)/cookbooks/cosmos3
VLLM_PORT=8000

setsid .venv/bin/vllm serve nvidia/Cosmos3-Nano \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --media-io-kwargs '{"video": {"num_frames": -1}}' \
  --port $VLLM_PORT > vllm_server.log 2>&1 &
disown

```

Key flags explained:
- `--hf-overrides`: Explicitly sets the model architecture class to `Cosmos3ReasonerForConditionalGeneration`
- `--allowed-local-media-path`: Permits the server to access media files via `file://` URLs within the specified directory
- `--media-io-kwargs`: Configures video processing to load all frames (`num_frames: -1`)

### Multi-GPU Configuration (Super)

For the larger `nvidia/Cosmos3-Super` model, enable tensor parallelism across multiple GPUs:

```bash
export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --tensor-parallel-size 4 \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --port 8000

```

The `--tensor-parallel-size 4` flag distributes the model across four GPUs, as demonstrated in the notebook `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` (lines 166‑172).

### Disabling Guardrails (Optional)

To run without safety guardrail models, export a deployment configuration file and reference it at launch:

```bash

# Create config file (referencing README.md lines 403-417)

cat > no_guardrails.yaml << 'EOF'
deploy:
  guardrails: false
EOF

vllm serve nvidia/Cosmos3-Nano \
  --deploy-config no_guardrails.yaml \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --async-scheduling \
  --port 8000

```

## Sending Inference Requests

Once the server is running on port `8000`, you can interact with it using standard OpenAI‑compatible clients.

### Using curl for Image Inference

Send a base64‑encoded image with a text prompt:

```bash
IMAGE_URI="data:image/jpeg;base64,$(base64 -w 0 cookbooks/cosmos3/reasoner/assets/robot_153.jpg)"

curl -X POST http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "nvidia/Cosmos3-Nano",
        "messages": [
          {"role":"system","content":"You are a helpful assistant."},
          {"role":"user","content":[
            {"type":"image_url","image_url":{"url":"'"$IMAGE_URI"'"}},
            {"type":"text","text":"Describe what is happening in this image in one sentence."}
          ]}
        ],
        "max_tokens":256,
        "stream":false
      }'

```

### Using the OpenAI Python Client for Video Analysis

For video understanding, pass a video URL and specify frame sampling rates via `extra_body`:

```python
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")

response = client.chat.completions.create(
    model="nvidia/Cosmos3-Nano",
    messages=[
        {"role":"system","content":"You are a helpful assistant."},
        {"role":"user","content":[
            {"type":"video_url","video_url":{"url":"https://download.samplelib.com/mp4/sample-5s.mp4"}},
            {"type":"text","text":"List the notable events with approximate timestamps."}
        ]}
    ],
    max_tokens=256,
    stream=False,
    extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}}
)

print(response.choices[0].message.content)

```

## Production Deployment Considerations

For production workloads, expose the vLLM server behind a reverse proxy (NGINX, Envoy, or cloud‑native Ingress) to handle SSL termination and load balancing. Enable health checks using the `GET /v1/health` endpoint to ensure uptime.

When scaling horizontally, run multiple vLLM instances on distinct ports and aggregate them behind a load balancer. Ensure all instances share the same `--allowed-local-media-path` or use object storage URLs to maintain consistency across replicas.

## Summary

- **Match CUDA and vLLM versions**: Use `cu130` with vLLM 0.21.0 for CUDA 13, or `cu128` with vLLM 0.19.1 for CUDA 12.8.
- **Install the plugin**: The `vllm‑cosmos3` package from the `cosmos-framework` repository is required to register the `Cosmos3ReasonerForConditionalGeneration` architecture.
- **Launch with correct flags**: Always include `--hf-overrides`, `--async-scheduling`, and `--allowed-local-media-path` for multimodal support.
- **Scale with tensor parallelism**: Use `--tensor-parallel-size` for multi‑GPU deployment of the Super model.
- **Consume via OpenAI API**: Use standard clients pointing to `http://localhost:8000/v1` with image or video payloads.

## Frequently Asked Questions

### What CUDA version should I use with Cosmos 3 Reasoner and vLLM?

You must match your CUDA driver version to the vLLM wheel. According to the NVIDIA Cosmos repository README, use `--torch-backend=cu130` with `vllm==0.21.0` for CUDA 13 systems, or `--torch-backend=cu128` with `vllm==0.19.1` for CUDA 12.8 systems. Mismatched versions will result in CUDA initialization failures where PyTorch cannot detect the GPU.

### How do I enable tensor parallelism for the Cosmos 3 Super model?

Add the `--tensor-parallel-size 4` flag to your `vllm serve` command and set `CUDA_VISIBLE_DEVICES=0,1,2,3` to specify which GPUs to use. This configuration is documented in the `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` notebook and distributes the model weights across four GPUs for efficient inference of the larger Super checkpoint.

### Can I use local media files with the vLLM Cosmos 3 Reasoner server?

Yes, but you must explicitly allow local file access using the `--allowed-local-media-path` flag when starting the server. Set this to the parent directory containing your media assets (e.g., `cookbooks/cosmos3`), then reference files using `file://` URLs in your API requests. Without this flag, the server will reject local file paths for security reasons.

### How do I disable the guardrail models in production?

Create a YAML deployment configuration file setting `deploy.guardrails` to `false`, then pass this file to the server using the `--deploy-config` flag. This bypasses the safety models entirely, reducing latency and memory overhead. Refer to the README.md lines 403‑417 in the NVIDIA Cosmos repository for the exact configuration syntax.