How to Run the Cosmos 3 Reasoner with vLLM for Video Understanding
The NVIDIA Cosmos 3 Reasoner can be served efficiently through vLLM, exposing an OpenAI-compatible REST API that accepts video URLs and frame-sampling parameters for multimodal inference.
The NVIDIA cosmos repository provides a production-ready pipeline for video understanding using the Cosmos 3 Reasoner family of models. By leveraging vLLM as the inference backend, you can run both the Nano and Super checkpoints with standard HTTP requests. This guide walks through the exact steps to launch the server and issue video understanding queries from a Python client.
Architecture Overview
The deployment stack consists of three layers:
- vLLM server — Loads the Cosmos 3 checkpoint (e.g.,
nvidia/cosmos3-super) and listens on a local port.8000is used for Nano, while8001is commonly used for Super. - OpenAI-compatible client — A thin Python wrapper built on the
openailibrary that constructs chat requests containing avideo_url, text prompt, and optionalmedia_io_kwargs. - Media handling — vLLM downloads or decodes the video, extracts frames according to the supplied sampling configuration, runs the visual encoder, and streams back a text response.
The request flows from the user client to the vllm.entrypoint.api_server, through the multimodal encoder, and finally to the LLM before returning captions or temporal reasoning results.
Environment Setup
Before launching the server, install vLLM and any remaining Cosmos 3 dependencies as described in the environment guide. The official cookbook at cookbooks/cosmos3/README.md contains the full walkthrough for preparing your Python environment.
You will also need the openai package in the client environment:
pip install openai
Launching the vLLM Server
Single-GPU Nano Setup
For smaller workloads or development, launch the Nano checkpoint on a single GPU:
CUDA_VISIBLE_DEVICES=0 \
python -m vllm.entrypoint.api_server \
--model nvidia/cosmos3-nano \
--port 8000
Multi-GPU Super Setup
For maximum reasoning quality, serve the Super checkpoint with tensor parallelism. As implemented in cookbooks/cosmos3/reasoner/README.md, allocate four GPUs and bind to port 8001:
CUDA_VISIBLE_DEVICES=0,1,2,3 \
python -m vllm.entrypoint.api_server \
--model nvidia/cosmos3-super \
--port 8001 \
--tensor-parallel-size 4
Wait until the server reports that the model weights are loaded and the HTTP endpoint is ready.
Running Video Understanding Requests via vLLM
Minimal Python Client Example
The repository notebook at cookbooks/cosmos3/reasoner/run_with_vllm.ipynb demonstrates the canonical client pattern. Below is a self-contained script that queries the local vLLM server:
import openai
video_url = (
"https://github.com/nvidia-cosmos/cosmos-dependencies/raw/refs/heads/assets/"
"cosmos3/inputs/video/temporal_localization_1.mp4"
)
client = openai.OpenAI(
api_key="EMPTY",
base_url="http://localhost:8001/v1"
)
response = client.chat.completions.create(
model=client.models.list().data[0].id,
messages=[
{
"role": "user",
"content": [
{"type": "video_url", "video_url": {"url": video_url}},
{"type": "text", "text": "Describe what is happening in this video."},
],
}
],
max_tokens=1024,
temperature=0.0,
)
print(response.choices[0].message.content)
The api_key is set to "EMPTY" because vLLM does not enforce authentication by default. The client automatically discovers the loaded checkpoint via client.models.list().data[0].id.
Controlling Frame Sampling with media_io_kwargs
You can adjust how many frames are extracted by passing media_io_kwargs inside the request's extra_body. This is useful for trading off inference speed against temporal granularity:
response = client.chat.completions.create(
model=client.models.list().data[0].id,
messages=[
{
"role": "user",
"content": [
{"type": "video_url", "video_url": {"url": video_url}},
{"type": "text", "text": "Describe what is happening in this video."},
],
}
],
max_tokens=1024,
temperature=0.0,
extra_body={
"media_io_kwargs": {
"video": {
"fps": 4.0,
"max_frames": 32,
"start_time": 0.0,
"end_time": 10.0,
}
}
},
)
vLLM uses these parameters during multimodal preprocessing to subsample the video before feeding frames to the Cosmos 3 visual encoder.
Advanced Configuration
Switching Checkpoints
Replace --model nvidia/cosmos3-super with nvidia/cosmos3-nano or a local path to a fine-tuned checkpoint. Update base_url in the client to match the new server port.
Batch Video Processing
Wrap the client.chat.completions.create call in a loop over multiple video_url values. Each request is handled independently by the vLLM asynchronous engine, so throughput scales with server concurrency limits.
Summary
- vLLM exposes the Cosmos 3 Reasoner as an OpenAI-compatible REST API on ports
8000(Nano) or8001(Super). - Python clients issue standard
chat.completions.createcalls withvideo_urlpayloads and text prompts. - Frame sampling is controlled through
extra_body={"media_io_kwargs": {"video": {...}}}. - Tensor parallelism via
--tensor-parallel-sizeenables serving the Super model across multiple GPUs.
Frequently Asked Questions
What port does the Cosmos 3 Reasoner vLLM server use by default?
Port 8000 is typical for the Nano checkpoint, while port 8001 is commonly assigned to the larger Super checkpoint in the official cookbook. You can override either with the --port flag when launching vllm.entrypoint.api_server.
How do I control video frame sampling when using vLLM?
Pass a media_io_kwargs dictionary inside the request's extra_body field. Keys such as fps, max_frames, start_time, and end_time inside the "video" nested dict let you subsample or trim the input before the visual encoder processes it.
Can I use the same client code for both image and video inputs?
Yes. The request structure is identical; simply change the content type from image_url to video_url in the message payload. The server handles modality-specific preprocessing automatically.
What hardware is required to run the Cosmos 3 Super model?
The Super checkpoint is designed for multi-GPU inference. The reference command in cookbooks/cosmos3/reasoner/README.md uses --tensor-parallel-size 4 across four GPUs. The Nano checkpoint can run on a single GPU.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →