How to Use the OpenAI-Compatible API Server in vLLM: Complete Setup Guide
vLLM provides a production-ready OpenAI-compatible HTTP server that exposes standard endpoints including /v1/completions, /v1/chat/completions, and /v1/models, enabling drop-in replacement for OpenAI's API while leveraging vLLM's high-throughput AsyncLLMEngine backend.
The vllm-project/vllm repository includes a complete OpenAI-compatible API implementation built on FastAPI. This server allows you to serve large language models locally or on dedicated infrastructure using the exact same request format and client libraries as OpenAI's official API, making migration seamless for existing applications.
Architecture of the OpenAI-Compatible Server
The vLLM OpenAI server consists of several integrated components that handle HTTP request parsing, tokenization, and distributed inference.
Core Components
vllm/entrypoints/openai/cli_args.py– Defines the command-line interface for thevllm servecommand, parsing arguments such as model path, tensor parallelism settings, and authentication tokens.vllm/entrypoints/openai/api_server.py– Contains therun_serverfunction that bootstraps the entire stack. This module creates the async engine client viabuild_async_engine_client, assembles the FastAPI application throughbuild_app, and launches the Uvicorn server.- Engine Client – A thin wrapper around
AsyncLLMEngine(instantiated viaAsyncEngineArgs.from_cli_args) that forwards generation requests to worker processes through ZMQ sockets. - FastAPI Routers – Modular endpoint handlers located in
vllm/entrypoints/openai/generate/api_router.py(for/v1/completionsand/v1/chat/completions) andvllm/entrypoints/openai/models/api_router.py(for/v1/models). - Middleware Stack – Optional
AuthenticationMiddlewareandScalingMiddlewareinjected duringbuild_appto handle request validation and autoscaling signals.
Request Processing Flow
When a client sends a generation request to the OpenAI-compatible API server, the following sequence occurs:
- The HTTP request arrives at
/v1/completionsor/v1/chat/completionsand routes through the appropriate FastAPI router. - The router validates the payload and processes tokenization through
OpenAIServingTokenizationinvllm/entrypoints/openai/tokenize/serving.py. - The engine client schedules the generation job across worker processes managed by
AsyncLLMEngine. - Results stream back to the client via Server-Sent Events (when
stream: true) or as a single JSON response.
Starting the OpenAI-Compatible Server
Basic Launch Command
The primary entry point is the vllm serve console script. A minimal single-GPU deployment requires only the model path:
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
--host 0.0.0.0 \
--port 8000
This command creates a single-process API server that internally spawns engine workers to perform inference. The server binds to all network interfaces on port 8000 and exposes the full OpenAI API specification at the /v1 prefix.
Multi-GPU and Security Configuration
For production deployments, configure tensor parallelism and authentication using flags defined in cli_args.py:
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
--tensor-parallel-size 2 \
--api-key sk-vllm-secret \
--ssl-keyfile /path/to/key.pem \
--ssl-certfile /path/to/cert.pem
Key parameters include:
--tensor-parallel-size– Splits the model across multiple GPUs for higher throughput.--api-key– Enforces bearer token authentication via theAuthorizationheader.--ssl-keyfile/--ssl-certfile– Enables HTTPS encryption for secure client connections.--enable-prompt-embeds– Allows passing pre-computed embeddings directly to the Completion API.
Calling the OpenAI-Compatible API
Raw HTTP Requests with curl
You can test the server immediately using standard HTTP tools without installing additional client libraries:
curl http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Meta-Llama-3-8B-Instruct",
"prompt": "Explain the difference between supervised and reinforcement learning.",
"max_tokens": 128,
"temperature": 0.7,
"stream": false
}'
Set "stream": true to receive a text/event-stream response where each chunk contains a partial completion delta.
Official OpenAI Python Client
The server is fully compatible with OpenAI's official Python SDK. Point the client to your local base URL:
import openai
client = openai.OpenAI(
base_url="http://localhost:8000/v1",
api_key="any-key-if-auth-enabled",
)
response = client.completions.create(
model="any-model", # Ignored by vLLM for single-model serving
prompt="Write a short poem about autumn.",
max_tokens=64,
temperature=0.8,
)
print(response.choices[0].text)
For conversational interfaces, use the chat completions endpoint which supports multi-turn dialogues and system prompts:
response = client.chat.completions.create(
model="any-model",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is quantum computing?"}
],
max_tokens=256
)
Streaming Real-Time Responses
Enable streaming to receive tokens as they are generated rather than waiting for the full completion:
stream = client.completions.create(
model="any-model",
prompt="List the first ten prime numbers.",
max_tokens=100,
temperature=0.0,
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
Multimodal Chat Completions
When serving vision-capable models, the /v1/chat/completions endpoint accepts image inputs encoded as base64 data URLs:
import base64
import pathlib
import openai
image_path = pathlib.Path("diagram.png")
image_b64 = base64.b64encode(image_path.read_bytes()).decode()
client = openai.OpenAI(base_url="http://localhost:8000/v1")
response = client.chat.completions.create(
model="any-model",
messages=[
{"role": "user", "content": [
{"type": "text", "text": "Explain this diagram:"},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image_b64}"}}
]}
],
max_tokens=200,
)
Advanced Server Features
Structured Outputs and Tool Calling
Enable advanced features by passing additional flags during server startup:
--enable-auto-tool-choice– Activates automatic tool selection for function-calling models.--tool-call-parser– Specifies the parser for extracting tool calls from model outputs.--tool-parser-plugin– Loads custom tool parsing logic from external plugins.
Validation for these arguments occurs in validate_api_server_args within api_server.py (lines 400–416).
Prompt Embeddings
For applications requiring embedding-based retrieval or cached prompt processing, start the server with --enable-prompt-embeds. This allows the Completion API to accept "prompt_embeds" arrays directly, bypassing the tokenization step for pre-computed embeddings.
Scaling Middleware
The server includes built-in ScalingMiddleware (located in vllm/entrypoints/serve/elastic_ep/middleware.py) that monitors external autoscaling signals. This enables integration with Kubernetes Horizontal Pod Autoscalers or custom cloud scaling logic without additional configuration flags.
Summary
- The OpenAI-compatible API server in vLLM provides drop-in replacements for
/v1/completions,/v1/chat/completions, and/v1/modelsendpoints using FastAPI and the AsyncLLMEngine backend. - Launch the server using
vllm serve <model>with optional flags for tensor parallelism (--tensor-parallel-size), authentication (--api-key), and SSL encryption. - Client code requires only a change to the
base_urlparameter in the OpenAI Python client or standard HTTP requests tolocalhost:8000/v1. - The server supports streaming responses, multimodal inputs (images/audio), structured outputs, and prompt embeddings through configuration flags and request payload options.
Frequently Asked Questions
How do I enable authentication on the vLLM OpenAI server?
Start the server with the --api-key <token> flag or set the VLLM_API_KEY environment variable. Clients must then include Authorization: Bearer <token> in their request headers. The AuthenticationMiddleware validates these tokens before processing requests.
Can I use the official OpenAI Python client with vLLM?
Yes. Configure the client with base_url="http://localhost:8000/v1" (adjusting host and port as needed) and set any string for api_key if authentication is enabled. The client methods completions.create() and chat.completions.create() function identically against the vLLM backend.
What is the difference between the /v1/completions and /v1/chat/completions endpoints?
The /v1/completions endpoint handles single-turn text generation with a raw prompt string, while /v1/chat/completions accepts structured message arrays with roles (system, user, assistant) and supports multi-modal content blocks. Both endpoints are implemented in vllm/entrypoints/openai/generate/api_router.py.
How do I configure tensor parallelism for the API server?
Pass --tensor-parallel-size N to the vllm serve command, where N matches your available GPU count. The server automatically shards the model across GPUs using the AsyncLLMEngine's distributed scheduling capabilities defined in the engine protocol layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →