# How to Use the OpenAI-Compatible API Server in vLLM: Complete Setup Guide

> Learn to set up vLLM's OpenAI-compatible API server for high-throughput LLM inference. Replace OpenAI's API seamlessly with vLLM's powerful backend.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**vLLM provides a production-ready OpenAI-compatible HTTP server that exposes standard endpoints including `/v1/completions`, `/v1/chat/completions`, and `/v1/models`, enabling drop-in replacement for OpenAI's API while leveraging vLLM's high-throughput AsyncLLMEngine backend.**

The vllm-project/vllm repository includes a complete OpenAI-compatible API implementation built on FastAPI. This server allows you to serve large language models locally or on dedicated infrastructure using the exact same request format and client libraries as OpenAI's official API, making migration seamless for existing applications.

## Architecture of the OpenAI-Compatible Server

The vLLM OpenAI server consists of several integrated components that handle HTTP request parsing, tokenization, and distributed inference.

### Core Components

- **[`vllm/entrypoints/openai/cli_args.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/cli_args.py)** – Defines the command-line interface for the `vllm serve` command, parsing arguments such as model path, tensor parallelism settings, and authentication tokens.
- **[`vllm/entrypoints/openai/api_server.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/api_server.py)** – Contains the `run_server` function that bootstraps the entire stack. This module creates the async engine client via `build_async_engine_client`, assembles the FastAPI application through `build_app`, and launches the Uvicorn server.
- **Engine Client** – A thin wrapper around `AsyncLLMEngine` (instantiated via `AsyncEngineArgs.from_cli_args`) that forwards generation requests to worker processes through ZMQ sockets.
- **FastAPI Routers** – Modular endpoint handlers located in [`vllm/entrypoints/openai/generate/api_router.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/generate/api_router.py) (for `/v1/completions` and `/v1/chat/completions`) and [`vllm/entrypoints/openai/models/api_router.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/models/api_router.py) (for `/v1/models`).
- **Middleware Stack** – Optional `AuthenticationMiddleware` and `ScalingMiddleware` injected during `build_app` to handle request validation and autoscaling signals.

### Request Processing Flow

When a client sends a generation request to the OpenAI-compatible API server, the following sequence occurs:

1. The HTTP request arrives at `/v1/completions` or `/v1/chat/completions` and routes through the appropriate FastAPI router.
2. The router validates the payload and processes tokenization through `OpenAIServingTokenization` in [`vllm/entrypoints/openai/tokenize/serving.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/tokenize/serving.py).
3. The engine client schedules the generation job across worker processes managed by `AsyncLLMEngine`.
4. Results stream back to the client via Server-Sent Events (when `stream: true`) or as a single JSON response.

## Starting the OpenAI-Compatible Server

### Basic Launch Command

The primary entry point is the `vllm serve` console script. A minimal single-GPU deployment requires only the model path:

```bash
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
    --host 0.0.0.0 \
    --port 8000

```

This command creates a single-process API server that internally spawns engine workers to perform inference. The server binds to all network interfaces on port 8000 and exposes the full OpenAI API specification at the `/v1` prefix.

### Multi-GPU and Security Configuration

For production deployments, configure tensor parallelism and authentication using flags defined in [`cli_args.py`](https://github.com/vllm-project/vllm/blob/main/cli_args.py):

```bash
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
    --tensor-parallel-size 2 \
    --api-key sk-vllm-secret \
    --ssl-keyfile /path/to/key.pem \
    --ssl-certfile /path/to/cert.pem

```

Key parameters include:
- **`--tensor-parallel-size`** – Splits the model across multiple GPUs for higher throughput.
- **`--api-key`** – Enforces bearer token authentication via the `Authorization` header.
- **`--ssl-keyfile` / `--ssl-certfile`** – Enables HTTPS encryption for secure client connections.
- **`--enable-prompt-embeds`** – Allows passing pre-computed embeddings directly to the Completion API.

## Calling the OpenAI-Compatible API

### Raw HTTP Requests with curl

You can test the server immediately using standard HTTP tools without installing additional client libraries:

```bash
curl http://localhost:8000/v1/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "meta-llama/Meta-Llama-3-8B-Instruct",
        "prompt": "Explain the difference between supervised and reinforcement learning.",
        "max_tokens": 128,
        "temperature": 0.7,
        "stream": false
      }'

```

Set `"stream": true` to receive a `text/event-stream` response where each chunk contains a partial completion delta.

### Official OpenAI Python Client

The server is fully compatible with OpenAI's official Python SDK. Point the client to your local base URL:

```python
import openai

client = openai.OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="any-key-if-auth-enabled",
)

response = client.completions.create(
    model="any-model",  # Ignored by vLLM for single-model serving

    prompt="Write a short poem about autumn.",
    max_tokens=64,
    temperature=0.8,
)

print(response.choices[0].text)

```

For conversational interfaces, use the chat completions endpoint which supports multi-turn dialogues and system prompts:

```python
response = client.chat.completions.create(
    model="any-model",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is quantum computing?"}
    ],
    max_tokens=256
)

```

### Streaming Real-Time Responses

Enable streaming to receive tokens as they are generated rather than waiting for the full completion:

```python
stream = client.completions.create(
    model="any-model",
    prompt="List the first ten prime numbers.",
    max_tokens=100,
    temperature=0.0,
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

```

### Multimodal Chat Completions

When serving vision-capable models, the `/v1/chat/completions` endpoint accepts image inputs encoded as base64 data URLs:

```python
import base64
import pathlib
import openai

image_path = pathlib.Path("diagram.png")
image_b64 = base64.b64encode(image_path.read_bytes()).decode()

client = openai.OpenAI(base_url="http://localhost:8000/v1")

response = client.chat.completions.create(
    model="any-model",
    messages=[
        {"role": "user", "content": [
            {"type": "text", "text": "Explain this diagram:"},
            {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image_b64}"}}
        ]}
    ],
    max_tokens=200,
)

```

## Advanced Server Features

### Structured Outputs and Tool Calling

Enable advanced features by passing additional flags during server startup:

- **`--enable-auto-tool-choice`** – Activates automatic tool selection for function-calling models.
- **`--tool-call-parser`** – Specifies the parser for extracting tool calls from model outputs.
- **`--tool-parser-plugin`** – Loads custom tool parsing logic from external plugins.

Validation for these arguments occurs in `validate_api_server_args` within [`api_server.py`](https://github.com/vllm-project/vllm/blob/main/api_server.py) (lines 400–416).

### Prompt Embeddings

For applications requiring embedding-based retrieval or cached prompt processing, start the server with `--enable-prompt-embeds`. This allows the Completion API to accept `"prompt_embeds"` arrays directly, bypassing the tokenization step for pre-computed embeddings.

### Scaling Middleware

The server includes built-in `ScalingMiddleware` (located in [`vllm/entrypoints/serve/elastic_ep/middleware.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/serve/elastic_ep/middleware.py)) that monitors external autoscaling signals. This enables integration with Kubernetes Horizontal Pod Autoscalers or custom cloud scaling logic without additional configuration flags.

## Summary

- The **OpenAI-compatible API server** in vLLM provides drop-in replacements for `/v1/completions`, `/v1/chat/completions`, and `/v1/models` endpoints using FastAPI and the AsyncLLMEngine backend.
- Launch the server using **`vllm serve <model>`** with optional flags for tensor parallelism (`--tensor-parallel-size`), authentication (`--api-key`), and SSL encryption.
- Client code requires only a change to the **`base_url`** parameter in the OpenAI Python client or standard HTTP requests to `localhost:8000/v1`.
- The server supports **streaming responses**, **multimodal inputs** (images/audio), **structured outputs**, and **prompt embeddings** through configuration flags and request payload options.

## Frequently Asked Questions

### How do I enable authentication on the vLLM OpenAI server?

Start the server with the `--api-key <token>` flag or set the `VLLM_API_KEY` environment variable. Clients must then include `Authorization: Bearer <token>` in their request headers. The `AuthenticationMiddleware` validates these tokens before processing requests.

### Can I use the official OpenAI Python client with vLLM?

Yes. Configure the client with `base_url="http://localhost:8000/v1"` (adjusting host and port as needed) and set any string for `api_key` if authentication is enabled. The client methods `completions.create()` and `chat.completions.create()` function identically against the vLLM backend.

### What is the difference between the `/v1/completions` and `/v1/chat/completions` endpoints?

The `/v1/completions` endpoint handles single-turn text generation with a raw prompt string, while `/v1/chat/completions` accepts structured message arrays with roles (system, user, assistant) and supports multi-modal content blocks. Both endpoints are implemented in [`vllm/entrypoints/openai/generate/api_router.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/generate/api_router.py).

### How do I configure tensor parallelism for the API server?

Pass `--tensor-parallel-size N` to the `vllm serve` command, where `N` matches your available GPU count. The server automatically shards the model across GPUs using the AsyncLLMEngine's distributed scheduling capabilities defined in the engine protocol layer.