How to Deploy Cosmos 3 Reasoner Using vLLM for Production Inference

Deploy Cosmos 3 Reasoner using vLLM by installing the CUDA‑matched vLLM wheel alongside the vllm‑cosmos3 plugin, then launching the server with architecture overrides and multimodal I/O flags to expose an OpenAI‑compatible chat completions endpoint.

The NVIDIA Cosmos repository provides a production‑ready inference stack for the Cosmos 3 Reasoner family of multimodal language models. By leveraging the vllm‑cosmos3 plugin, you can serve these models through a standard OpenAI‑compatible API, enabling seamless integration with existing client libraries and production infrastructure.

Prerequisites and Environment Setup

Proper deployment begins with a CUDA‑matched Python environment and the correct plugin installation.

CUDA Version Compatibility

vLLM wheels are compiled against specific CUDA versions; mismatched drivers will cause torch.cuda.is_available() to return False. The NVIDIA Cosmos documentation specifies the following pairings:

  • CUDA 13: Use --torch-backend=cu130 with vllm==0.21.0
  • CUDA 12.8: Use --torch-backend=cu128 with vllm==0.19.1

Verify your driver version before proceeding, as this determines which vLLM build you must install.

Installing vLLM and the Cosmos3 Plugin

Create an isolated environment and install the dependencies from the cosmos-framework submodule:


# Create virtual environment with Python 3.13

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

# Install CUDA-matched vLLM and the Cosmos3 plugin

# Replace cu130/cu128 based on your driver

uv pip install --torch-backend=cu130 "vllm==0.21.0" \
  "vllm-cosmos3 @ git+https://github.com/NVIDIA/cosmos-framework.git#subdirectory=packages/vllm-cosmos3"

The vllm‑cosmos3 package registers the Cosmos3ReasonerForConditionalGeneration architecture, which vLLM requires to load the model checkpoints correctly.

Launching the vLLM Server

The server initialization differs based on your hardware configuration and whether you are running the Nano or Super variant.

Single GPU Configuration (Nano)

For the nvidia/Cosmos3-Nano model on a single GPU, execute the following command from your repository root:

COSMOS3_MEDIA_ROOT=$(pwd)/cookbooks/cosmos3
VLLM_PORT=8000

setsid .venv/bin/vllm serve nvidia/Cosmos3-Nano \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --media-io-kwargs '{"video": {"num_frames": -1}}' \
  --port $VLLM_PORT > vllm_server.log 2>&1 &
disown

Key flags explained:

  • --hf-overrides: Explicitly sets the model architecture class to Cosmos3ReasonerForConditionalGeneration
  • --allowed-local-media-path: Permits the server to access media files via file:// URLs within the specified directory
  • --media-io-kwargs: Configures video processing to load all frames (num_frames: -1)

Multi-GPU Configuration (Super)

For the larger nvidia/Cosmos3-Super model, enable tensor parallelism across multiple GPUs:

export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --tensor-parallel-size 4 \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --port 8000

The --tensor-parallel-size 4 flag distributes the model across four GPUs, as demonstrated in the notebook cookbooks/cosmos3/reasoner/run_with_vllm.ipynb (lines 166‑172).

Disabling Guardrails (Optional)

To run without safety guardrail models, export a deployment configuration file and reference it at launch:


# Create config file (referencing README.md lines 403-417)

cat > no_guardrails.yaml << 'EOF'
deploy:
  guardrails: false
EOF

vllm serve nvidia/Cosmos3-Nano \
  --deploy-config no_guardrails.yaml \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --async-scheduling \
  --port 8000

Sending Inference Requests

Once the server is running on port 8000, you can interact with it using standard OpenAI‑compatible clients.

Using curl for Image Inference

Send a base64‑encoded image with a text prompt:

IMAGE_URI="data:image/jpeg;base64,$(base64 -w 0 cookbooks/cosmos3/reasoner/assets/robot_153.jpg)"

curl -X POST http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "nvidia/Cosmos3-Nano",
        "messages": [
          {"role":"system","content":"You are a helpful assistant."},
          {"role":"user","content":[
            {"type":"image_url","image_url":{"url":"'"$IMAGE_URI"'"}},
            {"type":"text","text":"Describe what is happening in this image in one sentence."}
          ]}
        ],
        "max_tokens":256,
        "stream":false
      }'

Using the OpenAI Python Client for Video Analysis

For video understanding, pass a video URL and specify frame sampling rates via extra_body:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")

response = client.chat.completions.create(
    model="nvidia/Cosmos3-Nano",
    messages=[
        {"role":"system","content":"You are a helpful assistant."},
        {"role":"user","content":[
            {"type":"video_url","video_url":{"url":"https://download.samplelib.com/mp4/sample-5s.mp4"}},
            {"type":"text","text":"List the notable events with approximate timestamps."}
        ]}
    ],
    max_tokens=256,
    stream=False,
    extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}}
)

print(response.choices[0].message.content)

Production Deployment Considerations

For production workloads, expose the vLLM server behind a reverse proxy (NGINX, Envoy, or cloud‑native Ingress) to handle SSL termination and load balancing. Enable health checks using the GET /v1/health endpoint to ensure uptime.

When scaling horizontally, run multiple vLLM instances on distinct ports and aggregate them behind a load balancer. Ensure all instances share the same --allowed-local-media-path or use object storage URLs to maintain consistency across replicas.

Summary

  • Match CUDA and vLLM versions: Use cu130 with vLLM 0.21.0 for CUDA 13, or cu128 with vLLM 0.19.1 for CUDA 12.8.
  • Install the plugin: The vllm‑cosmos3 package from the cosmos-framework repository is required to register the Cosmos3ReasonerForConditionalGeneration architecture.
  • Launch with correct flags: Always include --hf-overrides, --async-scheduling, and --allowed-local-media-path for multimodal support.
  • Scale with tensor parallelism: Use --tensor-parallel-size for multi‑GPU deployment of the Super model.
  • Consume via OpenAI API: Use standard clients pointing to http://localhost:8000/v1 with image or video payloads.

Frequently Asked Questions

What CUDA version should I use with Cosmos 3 Reasoner and vLLM?

You must match your CUDA driver version to the vLLM wheel. According to the NVIDIA Cosmos repository README, use --torch-backend=cu130 with vllm==0.21.0 for CUDA 13 systems, or --torch-backend=cu128 with vllm==0.19.1 for CUDA 12.8 systems. Mismatched versions will result in CUDA initialization failures where PyTorch cannot detect the GPU.

How do I enable tensor parallelism for the Cosmos 3 Super model?

Add the --tensor-parallel-size 4 flag to your vllm serve command and set CUDA_VISIBLE_DEVICES=0,1,2,3 to specify which GPUs to use. This configuration is documented in the cookbooks/cosmos3/reasoner/run_with_vllm.ipynb notebook and distributes the model weights across four GPUs for efficient inference of the larger Super checkpoint.

Can I use local media files with the vLLM Cosmos 3 Reasoner server?

Yes, but you must explicitly allow local file access using the --allowed-local-media-path flag when starting the server. Set this to the parent directory containing your media assets (e.g., cookbooks/cosmos3), then reference files using file:// URLs in your API requests. Without this flag, the server will reject local file paths for security reasons.

How do I disable the guardrail models in production?

Create a YAML deployment configuration file setting deploy.guardrails to false, then pass this file to the server using the --deploy-config flag. This bypasses the safety models entirely, reducing latency and memory overhead. Refer to the README.md lines 403‑417 in the NVIDIA Cosmos repository for the exact configuration syntax.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →