Setting Up Cosmos 3 Reasoner with vLLM for Production Inference
You can deploy Cosmos 3 Reasoner as an OpenAI-compatible chat-completion endpoint by installing the vllm-cosmos3 plugin, launching vLLM with the Cosmos3ReasonerForConditionalGeneration architecture override, and configuring multimodal I/O paths for production traffic.
The NVIDIA Cosmos repository provides a production-ready pathway for serving Cosmos 3 Reasoner models using vLLM. By setting up Cosmos 3 Reasoner with vLLM for production inference, you unlock an OpenAI-compatible API capable of processing images and video streams with structured text generation. This implementation leverages the Cosmos3ReasonerForConditionalGeneration architecture registered through the vllm-cosmos3 plugin to handle multimodal workloads efficiently.
Install the Correct vLLM Build
Production deployment begins with an isolated Python environment and a CUDA-matched vLLM installation. The vllm-cosmos3 plugin must be installed alongside vLLM to register the Cosmos 3 Reasoner architecture.
Create a clean virtual environment and install the dependencies:
# Create and activate a Python 3.13 environment
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
# Install CUDA-matched vLLM and the Cosmos-3 plugin
# CUDA 13 → cu130 + vllm==0.21.0
# CUDA 12.8 → cu128 + vllm==0.19.1
uv pip install --torch-backend=cu130 "vllm==0.21.0" \
"vllm-cosmos3 @ git+https://github.com/NVIDIA/cosmos-framework.git#subdirectory=packages/vllm-cosmos3"
Critical: vLLM wheels are compiled for specific CUDA versions. Mismatching the driver results in torch.cuda.is_available() == False. The version pairings above are reproduced from the repository's README.md (lines 38-40). If your driver is older than CUDA 12.8, substitute --torch-backend=cu128 with the appropriate legacy tag and use the matching vLLM version.
Launch the vLLM Reasoner Server
With dependencies installed, launch the server using the vllm serve command with architecture overrides and multimodal configuration flags.
First, clone the Cosmos framework repository to access local media paths:
COSMOS3_REPO=$HOME/cosmos-framework
git clone https://github.com/NVIDIA/cosmos-framework.git $COSMOS3_REPO
Start the server for the Nano model on a single GPU:
VLLM_PORT=8000
COSMOS3_MEDIA_ROOT=$(pwd)/cookbooks/cosmos3 # Root for local file:// media paths
setsid .venv/bin/vllm serve nvidia/Cosmos3-Nano \
--hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
--async-scheduling \
--allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
--media-io-kwargs '{"video": {"num_frames": -1}}' \
--port $VLLM_PORT > vllm_server.log 2>&1 &
disown
echo "vLLM server started on port $VLLM_PORT (log → vllm_server.log)"
Key flags explained:
--hf-overrides: Exposes theCosmos3ReasonerForConditionalGenerationarchitecture required by the model checkpoint.--async-scheduling: Enables asynchronous request processing for higher throughput.--allowed-local-media-path: Restricts file system access to the specified directory for security.--media-io-kwargs: Configures video processing;num_frames: -1processes full-length videos.
For the Super model on multiple GPUs, add tensor parallelism:
CUDA_VISIBLE_DEVICES=0,1,2,3 .venv/bin/vllm serve nvidia/Cosmos3-Super \
--tensor-parallel-size 4 \
--hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
--async-scheduling \
--allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
--port 8000
Guardrails configuration: To disable optional guardrail models, export a deploy config as shown in README.md (lines 403-417) and pass --deploy-config no_guardrails.yaml to the serve command.
Query the Reasoner with Production Clients
The server exposes an OpenAI-compatible POST /v1/chat/completions endpoint. Use standard HTTP clients or the OpenAI Python SDK to interact with the model.
Using curl for Image Inference
Send base64-encoded images or HTTPS URLs:
IMAGE_URI="data:image/jpeg;base64,$(base64 -w 0 cookbooks/cosmos3/reasoner/assets/robot_153.jpg)"
curl -X POST http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/Cosmos3-Nano",
"messages": [
{"role":"system","content":"You are a helpful assistant."},
{"role":"user","content":[
{"type":"image_url","image_url":{"url":"'"$IMAGE_URI"'"}},
{"type":"text","text":"Describe what is happening in this image in one sentence."}
]}
],
"max_tokens":256,
"stream":false
}'
Using the OpenAI Python Client for Video
Process video URLs with custom frame sampling rates:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
response = client.chat.completions.create(
model="nvidia/Cosmos3-Nano",
messages=[
{"role":"system","content":"You are a helpful assistant."},
{"role":"user","content":[
{"type":"video_url","video_url":{"url":"https://download.samplelib.com/mp4/sample-5s.mp4"}},
{"type":"text","text":"List the notable events with approximate timestamps."}
]}
],
max_tokens=256,
stream=False,
extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}}
)
print(response.choices[0].message.content)
Production scaling: Wrap the HTTP endpoint behind a load balancer (NGINX, Envoy, or cloud-native Ingress) and enable health checks via GET /v1/health. Deploy multiple vLLM pods on distinct ports and aggregate traffic using a reverse proxy to handle high concurrency loads.
Key Repository References
The implementation details are derived from the following source files in the NVIDIA Cosmos repository:
cookbooks/cosmos3/reasoner/run_with_vllm.ipynb: Complete notebook automating environment setup, server launch, and client calls (see lines 166-172 for multi-GPU configuration).README.md(Reasoner with vLLM section): Contains the version matrix, guardrail configuration templates (lines 403-417), and example requests (lines 521-638).packages/vllm-cosmos3: Plugin directory in thecosmos-frameworksubmodule that registers theCosmos3ReasonerForConditionalGenerationarchitecture.cookbooks/cosmos3/reasoner/assets/robot_153.jpg: Sample image used for testing local file path resolutions.
Summary
- Match CUDA and vLLM versions exactly: use
--torch-backend=cu130withvllm==0.21.0for CUDA 13, or--torch-backend=cu128withvllm==0.19.1for CUDA 12.8. - Install the
vllm-cosmos3plugin from the Cosmos-framework submodule to register the Reasoner architecture. - Launch with architecture overrides using
--hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}'and enable--async-schedulingfor throughput. - Configure media paths with
--allowed-local-media-pathand tune video processing via--media-io-kwargs. - Deploy behind a reverse proxy with health checks and load balancing for production resilience.
Frequently Asked Questions
Which CUDA version should I use with vLLM for Cosmos 3 Reasoner?
You must pair your CUDA driver version with a specific vLLM wheel. For CUDA 13, install with --torch-backend=cu130 and vllm==0.21.0. For CUDA 12.8, use --torch-backend=cu128 and vllm==0.19.1. Mismatching these versions causes torch.cuda.is_available() to return False, preventing GPU initialization.
How do I enable multi-GPU inference for Cosmos 3 Super?
Add the --tensor-parallel-size 4 flag to the vllm serve command and set CUDA_VISIBLE_DEVICES=0,1,2,3 to specify which GPUs to use. This configuration is documented in cookbooks/cosmos3/reasoner/run_with_vllm.ipynb (lines 166-172) and distributes the model across four GPUs for higher throughput.
Can I disable guardrails in production deployments?
Yes. Export a deployment configuration file that disables guardrail models as shown in README.md (lines 403-417), then pass --deploy-config no_guardrails.yaml when starting the server. This reduces latency by skipping the safety classification steps.
What media formats does the OpenAI-compatible endpoint support?
The endpoint supports images via image_url content types (base64 or HTTPS URLs) and videos via video_url content types with configurable preprocessing. Use --media-io-kwargs to set video.num_frames (use -1 for full-length) or pass extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}} in client requests to control frame sampling rates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →