# How to Deploy VibeVoice-ASR with vLLM for High-Performance Inference: A Complete Guide

> Deploy VibeVoice-ASR with vLLM for high-performance inference. Use start_server.py to wrap the model as an OpenAI-compatible API for efficient throughput scaling.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**Deploy VibeVoice-ASR with vLLM by using the [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) launcher to wrap the model as an OpenAI-compatible API, supporting single-GPU, data-parallel, and tensor-parallel configurations for throughput scaling.**

The `microsoft/VibeVoice` repository provides a production-ready pipeline to deploy VibeVoice-ASR with vLLM for high-performance inference, enabling continuous batching and streaming transcription of audio files exceeding 60 minutes. This deployment leverages vLLM's fast inference engine to expose an OpenAI-compatible `/v1/chat/completions` endpoint with support for hot-words and automatic recovery from repetition loops.

## Model Preparation and Tokenizer Generation

Before starting the server, you must generate the six tokenizer artifacts required by vLLM to handle VibeVoice-specific audio tokens. The repository includes [`vllm_plugin/tools/generate_tokenizer_files.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tools/generate_tokenizer_files.py), which downloads the Qwen2.5 tokenizer and patches it with VibeVoice audio tokens to create [`vocab.json`](https://github.com/microsoft/VibeVoice/blob/main/vocab.json), [`merges.txt`](https://github.com/microsoft/VibeVoice/blob/main/merges.txt), [`tokenizer.json`](https://github.com/microsoft/VibeVoice/blob/main/tokenizer.json), [`tokenizer_config.json`](https://github.com/microsoft/VibeVoice/blob/main/tokenizer_config.json), [`added_tokens.json`](https://github.com/microsoft/VibeVoice/blob/main/added_tokens.json), and [`special_tokens_map.json`](https://github.com/microsoft/VibeVoice/blob/main/special_tokens_map.json).

The [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) launcher automates this step unless you pass the `--skip-tokenizer` flag, but you can run the tool manually if customizing the tokenizer configuration. This patching extends the context length support up to 131,072 tokens, critical for processing long-form audio.

## Deploying with Docker: Single-GPU Setup

The simplest deployment uses the official vLLM container image with GPU passthrough. The [`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py) orchestrates dependency installation, model weight download from Hugging Face (`microsoft/VibeVoice-ASR`), tokenizer generation, and server startup.

```bash
git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice

docker run -d --gpus all --name vibevoice-vllm \
  --ipc=host \
  -p 8000:8000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v $(pwd):/app \
  -w /app \
  --entrypoint bash \
  vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py"

```

The container mounts the repository at `/app`, allowing the launcher to access the source code and any audio files you place in the directory. The `VIBEVOICE_FFMPEG_MAX_CONCURRENCY` environment variable controls the worker pool size for audio decoding, while `PYTORCH_ALLOC_CONF=expandable_segments:True` optimizes GPU memory management for long sequences.

## Scaling to Multiple GPUs: Data Parallel and Tensor Parallel

For production workloads, the [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) launcher supports two distributed strategies: **Data Parallel (DP)** for throughput scaling across GPUs, and **Tensor Parallel (TP)** for model sharding when a single GPU cannot hold the full weights.

### Data Parallel (DP) for Throughput Scaling

Use the `--dp N` flag to launch N independent vLLM workers, each bound to a specific GPU subset, behind an nginx reverse proxy with round-robin load balancing. This configuration maximizes request throughput for concurrent clients.

```bash
docker run -d --gpus '"device=0,1,2,3"' --name vibevoice-vllm \
  --ipc=host -p 8000:8000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v $(pwd):/app -w /app \
  --entrypoint bash vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py --dp 4"

```

The launcher automatically writes an nginx configuration to distribute incoming requests across the four workers, each processing on a dedicated GPU.

### Tensor Parallel (TP) for Model Sharding

When the model exceeds single-GPU memory capacity, use `--tp N` to split the weight matrix across N GPUs using vLLM's tensor parallelism.

```bash
docker run -d --gpus '"device=0,1"' --name vibevoice-vllm \
  --ipc=host -p 8000:8000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v $(pwd):/app -w /app \
  --entrypoint bash vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py --tp 2"

```

### Hybrid DP × TP Configuration

Combine both strategies to create replicas of tensor-parallel groups. For example, `--dp 2 --tp 2` on a 4-GPU node creates two replicas, each splitting the model across two GPUs, yielding both memory capacity and throughput scaling.

```bash
docker run -d --gpus '"device=0,1,2,3"' --name vibevoice-vllm \
  --ipc=host -p 8000:8000 \
  -v $(pwd):/app -w /app \
  --entrypoint bash vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py --dp 2 --tp 2"

```

## Testing the Inference Endpoint

The repository includes test scripts in `vllm_plugin/tests/` that verify the OpenAI-compatible endpoint at `http://localhost:8000/v1/chat/completions`.

### Basic Transcription and Hot-Words

Run [`vllm_plugin/tests/test_api.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api.py) to perform end-to-end transcription testing. Pass the `--hotwords` argument to bias recognition toward domain-specific terminology like proper nouns.

```bash

# Basic transcription test

docker exec -it vibevoice-vllm \
  python3 vllm_plugin/tests/test_api.py /app/audio.wav

# With hot-word biasing for improved accuracy

docker exec -it vibevoice-vllm \
  python3 vllm_plugin/tests/test_api.py /app/audio.wav \
  --hotwords "Microsoft,VibeVoice"

```

### Auto-Recovery for Long-Form Audio

To verify the server's ability to recover from repetition loops that can occur during streaming of very long audio, use [`vllm_plugin/tests/test_api_auto_recover.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api_auto_recover.py). This test ensures stability when transcribing content exceeding 60 minutes.

```bash
docker exec -it vibevoice-vllm \
  python3 vllm_plugin/tests/test_api_auto_recover.py /app/audio.wav

```

## Performance Tuning Configuration

Optimize throughput and latency by adjusting vLLM runtime parameters passed through the launcher:

- **GPU Memory Utilization**: Set `--gpu-memory-utilization 0.9` to reserve 90% of GPU memory for the KV cache, maximizing throughput if memory permits.
- **Maximum Sequence Count**: Increase `--max-num-seqs` (default 64) to raise concurrency, trading GPU memory for higher batch processing capacity.
- **Model Context Length**: The default `--max-model-len` is 65536 tokens; for extended audio, raise this up to 131072 tokens after tokenizer patching.
- **FFmpeg Concurrency**: Tune `VIBEVOICE_FFMPEG_MAX_CONCURRENCY` to match your CPU core count, preventing audio decoding from becoming a bottleneck.

## Summary

- **Model Preparation**: Run [`vllm_plugin/tools/generate_tokenizer_files.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tools/generate_tokenizer_files.py) to patch the Qwen2.5 tokenizer with VibeVoice audio tokens before starting the server.
- **Server Launcher**: Use [`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py) to automate dependency installation, model download, and vLLM initialization with support for single-GPU, DP, and TP modes.
- **Docker Deployment**: Mount the repository into `vllm/vllm-openai:v0.14.1` and execute the launcher to expose the OpenAI-compatible API on port 8000.
- **Scaling Strategies**: Deploy `--dp N` for multi-GPU throughput with nginx load balancing, or `--tp N` for model sharding across GPUs when memory is constrained.
- **Testing**: Validate deployment using [`vllm_plugin/tests/test_api.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api.py) for standard transcription and [`vllm_plugin/tests/test_api_auto_recover.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api_auto_recover.py) for long-audio stability.

## Frequently Asked Questions

### What Docker image should I use to deploy VibeVoice-ASR with vLLM?

Use the official `vllm/vllm-openai:v0.14.1` image as the base container. Mount the VibeVoice repository into `/app` and execute [`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py) as the entrypoint command. This image provides the necessary CUDA runtime and Python environment to run the vLLM inference engine with GPU support.

### How do I enable multi-GPU scaling for VibeVoice-ASR inference?

Pass the `--dp N` flag to [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) for **Data Parallel** deployment, which creates N independent vLLM workers behind an nginx reverse proxy for load balancing. For **Tensor Parallel** model sharding across GPUs when memory is insufficient, use `--tp N`. You can combine both flags (e.g., `--dp 2 --tp 2` on a 4-GPU system) to create replicated, sharded model instances.

### What is the purpose of the tokenizer generation step in VibeVoice-ASR deployment?

The [`vllm_plugin/tools/generate_tokenizer_files.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tools/generate_tokenizer_files.py) script downloads the base Qwen2.5 tokenizer and patches it with VibeVoice-specific audio tokens, generating six required JSON and text files ([`tokenizer.json`](https://github.com/microsoft/VibeVoice/blob/main/tokenizer.json), [`vocab.json`](https://github.com/microsoft/VibeVoice/blob/main/vocab.json), etc.). This step is mandatory because vLLM requires these specific tokenizer artifacts to correctly encode audio inputs alongside text prompts for the ASR model.

### How does VibeVoice-ASR handle very long audio files without repetition loops?

The deployment includes an auto-recovery mechanism tested via [`vllm_plugin/tests/test_api_auto_recover.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api_auto_recover.py). When processing audio exceeding 60 minutes, the vLLM server can detect and recover from repetition loops that sometimes occur during streaming generation. The system maintains stability by managing KV cache allocation through `PYTORCH_ALLOC_CONF=expandable_segments:True` and supporting context lengths up to 131,072 tokens after tokenizer patching.