How to Deploy Cosmos 3 Reasoner Using vLLM for Production Inference
Deploy Cosmos 3 Reasoner using vLLM by installing the CUDA‑matched vLLM wheel alongside the vllm‑cosmos3 plugin, then launching the server with architecture overrides and multimodal I/O flags to expose an OpenAI‑compatible chat completions endpoint.
The NVIDIA Cosmos repository provides a production‑ready inference stack for the Cosmos 3 Reasoner family of multimodal language models. By leveraging the vllm‑cosmos3 plugin, you can serve these models through a standard OpenAI‑compatible API, enabling seamless integration with existing client libraries and production infrastructure.
Prerequisites and Environment Setup
Proper deployment begins with a CUDA‑matched Python environment and the correct plugin installation.
CUDA Version Compatibility
vLLM wheels are compiled against specific CUDA versions; mismatched drivers will cause torch.cuda.is_available() to return False. The NVIDIA Cosmos documentation specifies the following pairings:
- CUDA 13: Use
--torch-backend=cu130withvllm==0.21.0 - CUDA 12.8: Use
--torch-backend=cu128withvllm==0.19.1
Verify your driver version before proceeding, as this determines which vLLM build you must install.
Installing vLLM and the Cosmos3 Plugin
Create an isolated environment and install the dependencies from the cosmos-framework submodule:
# Create virtual environment with Python 3.13
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
# Install CUDA-matched vLLM and the Cosmos3 plugin
# Replace cu130/cu128 based on your driver
uv pip install --torch-backend=cu130 "vllm==0.21.0" \
"vllm-cosmos3 @ git+https://github.com/NVIDIA/cosmos-framework.git#subdirectory=packages/vllm-cosmos3"
The vllm‑cosmos3 package registers the Cosmos3ReasonerForConditionalGeneration architecture, which vLLM requires to load the model checkpoints correctly.
Launching the vLLM Server
The server initialization differs based on your hardware configuration and whether you are running the Nano or Super variant.
Single GPU Configuration (Nano)
For the nvidia/Cosmos3-Nano model on a single GPU, execute the following command from your repository root:
COSMOS3_MEDIA_ROOT=$(pwd)/cookbooks/cosmos3
VLLM_PORT=8000
setsid .venv/bin/vllm serve nvidia/Cosmos3-Nano \
--hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
--async-scheduling \
--allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
--media-io-kwargs '{"video": {"num_frames": -1}}' \
--port $VLLM_PORT > vllm_server.log 2>&1 &
disown
Key flags explained:
--hf-overrides: Explicitly sets the model architecture class toCosmos3ReasonerForConditionalGeneration--allowed-local-media-path: Permits the server to access media files viafile://URLs within the specified directory--media-io-kwargs: Configures video processing to load all frames (num_frames: -1)
Multi-GPU Configuration (Super)
For the larger nvidia/Cosmos3-Super model, enable tensor parallelism across multiple GPUs:
export CUDA_VISIBLE_DEVICES=0,1,2,3
vllm serve nvidia/Cosmos3-Super \
--hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
--tensor-parallel-size 4 \
--async-scheduling \
--allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
--port 8000
The --tensor-parallel-size 4 flag distributes the model across four GPUs, as demonstrated in the notebook cookbooks/cosmos3/reasoner/run_with_vllm.ipynb (lines 166‑172).
Disabling Guardrails (Optional)
To run without safety guardrail models, export a deployment configuration file and reference it at launch:
# Create config file (referencing README.md lines 403-417)
cat > no_guardrails.yaml << 'EOF'
deploy:
guardrails: false
EOF
vllm serve nvidia/Cosmos3-Nano \
--deploy-config no_guardrails.yaml \
--hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
--async-scheduling \
--port 8000
Sending Inference Requests
Once the server is running on port 8000, you can interact with it using standard OpenAI‑compatible clients.
Using curl for Image Inference
Send a base64‑encoded image with a text prompt:
IMAGE_URI="data:image/jpeg;base64,$(base64 -w 0 cookbooks/cosmos3/reasoner/assets/robot_153.jpg)"
curl -X POST http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/Cosmos3-Nano",
"messages": [
{"role":"system","content":"You are a helpful assistant."},
{"role":"user","content":[
{"type":"image_url","image_url":{"url":"'"$IMAGE_URI"'"}},
{"type":"text","text":"Describe what is happening in this image in one sentence."}
]}
],
"max_tokens":256,
"stream":false
}'
Using the OpenAI Python Client for Video Analysis
For video understanding, pass a video URL and specify frame sampling rates via extra_body:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
response = client.chat.completions.create(
model="nvidia/Cosmos3-Nano",
messages=[
{"role":"system","content":"You are a helpful assistant."},
{"role":"user","content":[
{"type":"video_url","video_url":{"url":"https://download.samplelib.com/mp4/sample-5s.mp4"}},
{"type":"text","text":"List the notable events with approximate timestamps."}
]}
],
max_tokens=256,
stream=False,
extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}}
)
print(response.choices[0].message.content)
Production Deployment Considerations
For production workloads, expose the vLLM server behind a reverse proxy (NGINX, Envoy, or cloud‑native Ingress) to handle SSL termination and load balancing. Enable health checks using the GET /v1/health endpoint to ensure uptime.
When scaling horizontally, run multiple vLLM instances on distinct ports and aggregate them behind a load balancer. Ensure all instances share the same --allowed-local-media-path or use object storage URLs to maintain consistency across replicas.
Summary
- Match CUDA and vLLM versions: Use
cu130with vLLM 0.21.0 for CUDA 13, orcu128with vLLM 0.19.1 for CUDA 12.8. - Install the plugin: The
vllm‑cosmos3package from thecosmos-frameworkrepository is required to register theCosmos3ReasonerForConditionalGenerationarchitecture. - Launch with correct flags: Always include
--hf-overrides,--async-scheduling, and--allowed-local-media-pathfor multimodal support. - Scale with tensor parallelism: Use
--tensor-parallel-sizefor multi‑GPU deployment of the Super model. - Consume via OpenAI API: Use standard clients pointing to
http://localhost:8000/v1with image or video payloads.
Frequently Asked Questions
What CUDA version should I use with Cosmos 3 Reasoner and vLLM?
You must match your CUDA driver version to the vLLM wheel. According to the NVIDIA Cosmos repository README, use --torch-backend=cu130 with vllm==0.21.0 for CUDA 13 systems, or --torch-backend=cu128 with vllm==0.19.1 for CUDA 12.8 systems. Mismatched versions will result in CUDA initialization failures where PyTorch cannot detect the GPU.
How do I enable tensor parallelism for the Cosmos 3 Super model?
Add the --tensor-parallel-size 4 flag to your vllm serve command and set CUDA_VISIBLE_DEVICES=0,1,2,3 to specify which GPUs to use. This configuration is documented in the cookbooks/cosmos3/reasoner/run_with_vllm.ipynb notebook and distributes the model weights across four GPUs for efficient inference of the larger Super checkpoint.
Can I use local media files with the vLLM Cosmos 3 Reasoner server?
Yes, but you must explicitly allow local file access using the --allowed-local-media-path flag when starting the server. Set this to the parent directory containing your media assets (e.g., cookbooks/cosmos3), then reference files using file:// URLs in your API requests. Without this flag, the server will reject local file paths for security reasons.
How do I disable the guardrail models in production?
Create a YAML deployment configuration file setting deploy.guardrails to false, then pass this file to the server using the --deploy-config flag. This bypasses the safety models entirely, reducing latency and memory overhead. Refer to the README.md lines 403‑417 in the NVIDIA Cosmos repository for the exact configuration syntax.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →