# How to Use OpenAI-Compatible API Endpoints with VibeVoice vLLM: A Complete Guide

> Integrate VibeVoice vLLM with OpenAI compatible API endpoints. Stream audio and get transcriptions using familiar REST server formats. Get the complete guide now.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**VibeVoice exposes its streaming ASR model through a vLLM plugin that launches an OpenAI-compatible REST server, allowing you to send audio data via standard chat completion endpoints and receive streaming transcriptions using familiar request formats.**

The microsoft/VibeVoice repository provides a **vLLM plugin** that wraps the VibeVoice streaming ASR model into a fully OpenAI-compatible REST API. This integration lets developers interact with the speech recognition model using standard OpenAI chat completion protocols, eliminating the need for custom client libraries or specialized request formats.

## Starting the OpenAI-Compatible Server

The server launch process is handled by **[`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py)**, which constructs and executes a `vllm serve` command. At line 104, the script explicitly configures the chat template format for OpenAI compatibility by passing the argument `"--chat-template-content-format", "openai"`. This ensures that the vLLM engine interprets incoming messages according to the OpenAI chat schema.

The script automatically downloads the model, generates necessary tokenizer files, and binds to the specified port. Whether deploying on a single GPU or scaling across multiple instances behind an nginx reverse proxy, the startup process remains consistent.

## OpenAI-Compatible Endpoints and Request Format

Once running, the server exposes the standard OpenAI REST endpoints. The **`GET /v1/models`** endpoint lists the available model as `vibevoice`, while **`POST /v1/chat/completions`** accepts chat-style payloads for transcription tasks.

### Constructing the Audio Payload

The request body follows the OpenAI chat schema with a specific multimodal structure. The user message must contain two parts: an `audio_url` entry containing a **base64-encoded data URL**, and a text entry providing the transcription prompt.

As implemented in [`vllm_plugin/tests/test_api.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api.py) (lines 62-73), the payload structure requires:

- A system message defining the assistant's role
- A user message with a content array containing:
  - `{"type": "audio_url", "audio_url": {"url": "data:audio/wav;base64,..."}}`
  - `{"type": "text", "text": "Transcription instructions..."}`

### Handling Streaming Responses

The server returns responses in **Server-Sent Events (SSE)** format, where each line begins with `data: `. Clients must iterate over the response stream, parsing each JSON object to extract the `content` field from `choices[0].delta`.

## Advanced Configuration: Hot-Words and Parallel Deployment

### Injecting Hot-Words via Prompt Text

Unlike traditional ASR systems that accept separate hot-word parameters, VibeVoice integrates domain-specific vocabulary directly into the prompt text. Prepend hot-words to the transcription instructions (e.g., "with extra info: Microsoft,Azure") to condition the model during generation.

### Scaling with Data and Tensor Parallelism

The OpenAI-compatible API contract remains identical whether running a single instance or deploying a data-parallel cluster behind nginx. The [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) script handles distributed configuration without altering the endpoint behavior or request format.

## Implementation Examples

### Launching the Server

```bash
python3 -m vllm_plugin.scripts.start_server \
    --model microsoft/VibeVoice-ASR \
    --port 8000

```

### Python Client for Audio Transcription

```python
import base64, json, requests

# Load audio and encode as a data URL

with open("sample.wav", "rb") as f:
    audio_b64 = base64.b64encode(f.read()).decode()
data_url = f"data:audio/wav;base64,{audio_b64}"

payload = {
    "model": "vibevoice",
    "messages": [
        {"role": "system",
         "content": "You are a helpful assistant that transcribes audio into JSON."},
        {"role": "user",
         "content": [
             {"type": "audio_url", "audio_url": {"url": data_url}},
             {"type": "text",
              "text": "Please transcribe this audio with keys: Start time, End time, Speaker ID, Content"}
         ]}
    ],
    "max_tokens": 32768,
    "temperature": 0.0,
    "stream": True,
    "top_p": 1.0,
}

resp = requests.post("http://localhost:8000/v1/chat/completions",
                     json=payload, stream=True, timeout=12000)

for line in resp.iter_lines():
    if line and line.startswith(b"data: "):
        msg = json.loads(line[6:])
        delta = msg["choices"][0]["delta"]
        if "content" in delta:
            print(delta["content"], end="", flush=True)

```

### Adding Hot-Words to Your Request

```python
hotwords = "Microsoft,Azure,VibeVoice"
prompt = (f"This is a 12.34 seconds audio, with extra info: {hotwords}\n"
          "Please transcribe it with keys: Start time, End time, Speaker ID, Content")

# Insert `prompt` into the "text" field of the user message content array

```

### Testing with curl

```bash

# Verify model availability

curl -sS http://localhost:8000/v1/models | jq

# Send transcription request with streaming

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d @request.json --no-buffer

```

Note: The `--no-buffer` flag ensures curl outputs streaming chunks immediately as they arrive from the SSE stream.

## Summary

- The **vLLM plugin** in [`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py) launches an OpenAI-compatible server using `--chat-template-content-format openai` (line 104).
- Standard endpoints **`/v1/models`** and **`/v1/chat/completions`** handle model listing and audio transcription requests.
- Audio data must be embedded as **base64 data URLs** within the `audio_url` message content field, paired with text instructions.
- **Hot-words** are injected directly into the prompt text rather than passed as separate parameters.
- The API supports **streaming SSE responses** and scales from single-GPU deployments to data-parallel clusters without protocol changes.

## Frequently Asked Questions

### What file handles the OpenAI-compatible server startup in VibeVoice?

The **[`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py)** script manages the entire server initialization process. It builds the `vllm serve` command, downloads the model, generates tokenizer files, and explicitly sets `--chat-template-content-format openai` at line 104 to ensure OpenAI API compatibility.

### How do I format audio data for the VibeVoice OpenAI-compatible API?

Audio must be **base64-encoded** and wrapped in a data URL format (`data:audio/wav;base64,...`). This string is placed inside the `audio_url` object within the user message's content array, as demonstrated in [`vllm_plugin/tests/test_api.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api.py) (lines 62-73).

### Can I use standard OpenAI client libraries with VibeVoice?

Yes. Because the server implements the standard OpenAI REST protocol including the `/v1/chat/completions` endpoint with SSE streaming, any HTTP client or OpenAI SDK that supports streaming completions can connect to VibeVoice by pointing the base URL to your server (e.g., `http://localhost:8000`).

### How do I add custom vocabulary or hot-words to improve transcription accuracy?

Instead of using a dedicated hot-word parameter, prepend the vocabulary words to the **text portion** of your user message (e.g., "with extra info: Microsoft,Azure"). The model conditions on this context during generation to improve recognition of specific terms.