How to Use OpenAI-Compatible API Endpoints with VibeVoice vLLM: A Complete Guide
VibeVoice exposes its streaming ASR model through a vLLM plugin that launches an OpenAI-compatible REST server, allowing you to send audio data via standard chat completion endpoints and receive streaming transcriptions using familiar request formats.
The microsoft/VibeVoice repository provides a vLLM plugin that wraps the VibeVoice streaming ASR model into a fully OpenAI-compatible REST API. This integration lets developers interact with the speech recognition model using standard OpenAI chat completion protocols, eliminating the need for custom client libraries or specialized request formats.
Starting the OpenAI-Compatible Server
The server launch process is handled by vllm_plugin/scripts/start_server.py, which constructs and executes a vllm serve command. At line 104, the script explicitly configures the chat template format for OpenAI compatibility by passing the argument "--chat-template-content-format", "openai". This ensures that the vLLM engine interprets incoming messages according to the OpenAI chat schema.
The script automatically downloads the model, generates necessary tokenizer files, and binds to the specified port. Whether deploying on a single GPU or scaling across multiple instances behind an nginx reverse proxy, the startup process remains consistent.
OpenAI-Compatible Endpoints and Request Format
Once running, the server exposes the standard OpenAI REST endpoints. The GET /v1/models endpoint lists the available model as vibevoice, while POST /v1/chat/completions accepts chat-style payloads for transcription tasks.
Constructing the Audio Payload
The request body follows the OpenAI chat schema with a specific multimodal structure. The user message must contain two parts: an audio_url entry containing a base64-encoded data URL, and a text entry providing the transcription prompt.
As implemented in vllm_plugin/tests/test_api.py (lines 62-73), the payload structure requires:
- A system message defining the assistant's role
- A user message with a content array containing:
{"type": "audio_url", "audio_url": {"url": "data:audio/wav;base64,..."}}{"type": "text", "text": "Transcription instructions..."}
Handling Streaming Responses
The server returns responses in Server-Sent Events (SSE) format, where each line begins with data: . Clients must iterate over the response stream, parsing each JSON object to extract the content field from choices[0].delta.
Advanced Configuration: Hot-Words and Parallel Deployment
Injecting Hot-Words via Prompt Text
Unlike traditional ASR systems that accept separate hot-word parameters, VibeVoice integrates domain-specific vocabulary directly into the prompt text. Prepend hot-words to the transcription instructions (e.g., "with extra info: Microsoft,Azure") to condition the model during generation.
Scaling with Data and Tensor Parallelism
The OpenAI-compatible API contract remains identical whether running a single instance or deploying a data-parallel cluster behind nginx. The start_server.py script handles distributed configuration without altering the endpoint behavior or request format.
Implementation Examples
Launching the Server
python3 -m vllm_plugin.scripts.start_server \
--model microsoft/VibeVoice-ASR \
--port 8000
Python Client for Audio Transcription
import base64, json, requests
# Load audio and encode as a data URL
with open("sample.wav", "rb") as f:
audio_b64 = base64.b64encode(f.read()).decode()
data_url = f"data:audio/wav;base64,{audio_b64}"
payload = {
"model": "vibevoice",
"messages": [
{"role": "system",
"content": "You are a helpful assistant that transcribes audio into JSON."},
{"role": "user",
"content": [
{"type": "audio_url", "audio_url": {"url": data_url}},
{"type": "text",
"text": "Please transcribe this audio with keys: Start time, End time, Speaker ID, Content"}
]}
],
"max_tokens": 32768,
"temperature": 0.0,
"stream": True,
"top_p": 1.0,
}
resp = requests.post("http://localhost:8000/v1/chat/completions",
json=payload, stream=True, timeout=12000)
for line in resp.iter_lines():
if line and line.startswith(b"data: "):
msg = json.loads(line[6:])
delta = msg["choices"][0]["delta"]
if "content" in delta:
print(delta["content"], end="", flush=True)
Adding Hot-Words to Your Request
hotwords = "Microsoft,Azure,VibeVoice"
prompt = (f"This is a 12.34 seconds audio, with extra info: {hotwords}\n"
"Please transcribe it with keys: Start time, End time, Speaker ID, Content")
# Insert `prompt` into the "text" field of the user message content array
Testing with curl
# Verify model availability
curl -sS http://localhost:8000/v1/models | jq
# Send transcription request with streaming
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d @request.json --no-buffer
Note: The --no-buffer flag ensures curl outputs streaming chunks immediately as they arrive from the SSE stream.
Summary
- The vLLM plugin in
vllm_plugin/scripts/start_server.pylaunches an OpenAI-compatible server using--chat-template-content-format openai(line 104). - Standard endpoints
/v1/modelsand/v1/chat/completionshandle model listing and audio transcription requests. - Audio data must be embedded as base64 data URLs within the
audio_urlmessage content field, paired with text instructions. - Hot-words are injected directly into the prompt text rather than passed as separate parameters.
- The API supports streaming SSE responses and scales from single-GPU deployments to data-parallel clusters without protocol changes.
Frequently Asked Questions
What file handles the OpenAI-compatible server startup in VibeVoice?
The vllm_plugin/scripts/start_server.py script manages the entire server initialization process. It builds the vllm serve command, downloads the model, generates tokenizer files, and explicitly sets --chat-template-content-format openai at line 104 to ensure OpenAI API compatibility.
How do I format audio data for the VibeVoice OpenAI-compatible API?
Audio must be base64-encoded and wrapped in a data URL format (data:audio/wav;base64,...). This string is placed inside the audio_url object within the user message's content array, as demonstrated in vllm_plugin/tests/test_api.py (lines 62-73).
Can I use standard OpenAI client libraries with VibeVoice?
Yes. Because the server implements the standard OpenAI REST protocol including the /v1/chat/completions endpoint with SSE streaming, any HTTP client or OpenAI SDK that supports streaming completions can connect to VibeVoice by pointing the base URL to your server (e.g., http://localhost:8000).
How do I add custom vocabulary or hot-words to improve transcription accuracy?
Instead of using a dedicated hot-word parameter, prepend the vocabulary words to the text portion of your user message (e.g., "with extra info: Microsoft,Azure"). The model conditions on this context during generation to improve recognition of specific terms.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →