How to Deploy VibeVoice-ASR with vLLM for High-Performance Inference: A Complete Guide
Deploy VibeVoice-ASR with vLLM by using the start_server.py launcher to wrap the model as an OpenAI-compatible API, supporting single-GPU, data-parallel, and tensor-parallel configurations for throughput scaling.
The microsoft/VibeVoice repository provides a production-ready pipeline to deploy VibeVoice-ASR with vLLM for high-performance inference, enabling continuous batching and streaming transcription of audio files exceeding 60 minutes. This deployment leverages vLLM's fast inference engine to expose an OpenAI-compatible /v1/chat/completions endpoint with support for hot-words and automatic recovery from repetition loops.
Model Preparation and Tokenizer Generation
Before starting the server, you must generate the six tokenizer artifacts required by vLLM to handle VibeVoice-specific audio tokens. The repository includes vllm_plugin/tools/generate_tokenizer_files.py, which downloads the Qwen2.5 tokenizer and patches it with VibeVoice audio tokens to create vocab.json, merges.txt, tokenizer.json, tokenizer_config.json, added_tokens.json, and special_tokens_map.json.
The start_server.py launcher automates this step unless you pass the --skip-tokenizer flag, but you can run the tool manually if customizing the tokenizer configuration. This patching extends the context length support up to 131,072 tokens, critical for processing long-form audio.
Deploying with Docker: Single-GPU Setup
The simplest deployment uses the official vLLM container image with GPU passthrough. The vllm_plugin/scripts/start_server.py orchestrates dependency installation, model weight download from Hugging Face (microsoft/VibeVoice-ASR), tokenizer generation, and server startup.
git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice
docker run -d --gpus all --name vibevoice-vllm \
--ipc=host \
-p 8000:8000 \
-e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
-e PYTORCH_ALLOC_CONF=expandable_segments:True \
-v $(pwd):/app \
-w /app \
--entrypoint bash \
vllm/vllm-openai:v0.14.1 \
-c "python3 /app/vllm_plugin/scripts/start_server.py"
The container mounts the repository at /app, allowing the launcher to access the source code and any audio files you place in the directory. The VIBEVOICE_FFMPEG_MAX_CONCURRENCY environment variable controls the worker pool size for audio decoding, while PYTORCH_ALLOC_CONF=expandable_segments:True optimizes GPU memory management for long sequences.
Scaling to Multiple GPUs: Data Parallel and Tensor Parallel
For production workloads, the start_server.py launcher supports two distributed strategies: Data Parallel (DP) for throughput scaling across GPUs, and Tensor Parallel (TP) for model sharding when a single GPU cannot hold the full weights.
Data Parallel (DP) for Throughput Scaling
Use the --dp N flag to launch N independent vLLM workers, each bound to a specific GPU subset, behind an nginx reverse proxy with round-robin load balancing. This configuration maximizes request throughput for concurrent clients.
docker run -d --gpus '"device=0,1,2,3"' --name vibevoice-vllm \
--ipc=host -p 8000:8000 \
-e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
-e PYTORCH_ALLOC_CONF=expandable_segments:True \
-v $(pwd):/app -w /app \
--entrypoint bash vllm/vllm-openai:v0.14.1 \
-c "python3 /app/vllm_plugin/scripts/start_server.py --dp 4"
The launcher automatically writes an nginx configuration to distribute incoming requests across the four workers, each processing on a dedicated GPU.
Tensor Parallel (TP) for Model Sharding
When the model exceeds single-GPU memory capacity, use --tp N to split the weight matrix across N GPUs using vLLM's tensor parallelism.
docker run -d --gpus '"device=0,1"' --name vibevoice-vllm \
--ipc=host -p 8000:8000 \
-e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
-e PYTORCH_ALLOC_CONF=expandable_segments:True \
-v $(pwd):/app -w /app \
--entrypoint bash vllm/vllm-openai:v0.14.1 \
-c "python3 /app/vllm_plugin/scripts/start_server.py --tp 2"
Hybrid DP × TP Configuration
Combine both strategies to create replicas of tensor-parallel groups. For example, --dp 2 --tp 2 on a 4-GPU node creates two replicas, each splitting the model across two GPUs, yielding both memory capacity and throughput scaling.
docker run -d --gpus '"device=0,1,2,3"' --name vibevoice-vllm \
--ipc=host -p 8000:8000 \
-v $(pwd):/app -w /app \
--entrypoint bash vllm/vllm-openai:v0.14.1 \
-c "python3 /app/vllm_plugin/scripts/start_server.py --dp 2 --tp 2"
Testing the Inference Endpoint
The repository includes test scripts in vllm_plugin/tests/ that verify the OpenAI-compatible endpoint at http://localhost:8000/v1/chat/completions.
Basic Transcription and Hot-Words
Run vllm_plugin/tests/test_api.py to perform end-to-end transcription testing. Pass the --hotwords argument to bias recognition toward domain-specific terminology like proper nouns.
# Basic transcription test
docker exec -it vibevoice-vllm \
python3 vllm_plugin/tests/test_api.py /app/audio.wav
# With hot-word biasing for improved accuracy
docker exec -it vibevoice-vllm \
python3 vllm_plugin/tests/test_api.py /app/audio.wav \
--hotwords "Microsoft,VibeVoice"
Auto-Recovery for Long-Form Audio
To verify the server's ability to recover from repetition loops that can occur during streaming of very long audio, use vllm_plugin/tests/test_api_auto_recover.py. This test ensures stability when transcribing content exceeding 60 minutes.
docker exec -it vibevoice-vllm \
python3 vllm_plugin/tests/test_api_auto_recover.py /app/audio.wav
Performance Tuning Configuration
Optimize throughput and latency by adjusting vLLM runtime parameters passed through the launcher:
- GPU Memory Utilization: Set
--gpu-memory-utilization 0.9to reserve 90% of GPU memory for the KV cache, maximizing throughput if memory permits. - Maximum Sequence Count: Increase
--max-num-seqs(default 64) to raise concurrency, trading GPU memory for higher batch processing capacity. - Model Context Length: The default
--max-model-lenis 65536 tokens; for extended audio, raise this up to 131072 tokens after tokenizer patching. - FFmpeg Concurrency: Tune
VIBEVOICE_FFMPEG_MAX_CONCURRENCYto match your CPU core count, preventing audio decoding from becoming a bottleneck.
Summary
- Model Preparation: Run
vllm_plugin/tools/generate_tokenizer_files.pyto patch the Qwen2.5 tokenizer with VibeVoice audio tokens before starting the server. - Server Launcher: Use
vllm_plugin/scripts/start_server.pyto automate dependency installation, model download, and vLLM initialization with support for single-GPU, DP, and TP modes. - Docker Deployment: Mount the repository into
vllm/vllm-openai:v0.14.1and execute the launcher to expose the OpenAI-compatible API on port 8000. - Scaling Strategies: Deploy
--dp Nfor multi-GPU throughput with nginx load balancing, or--tp Nfor model sharding across GPUs when memory is constrained. - Testing: Validate deployment using
vllm_plugin/tests/test_api.pyfor standard transcription andvllm_plugin/tests/test_api_auto_recover.pyfor long-audio stability.
Frequently Asked Questions
What Docker image should I use to deploy VibeVoice-ASR with vLLM?
Use the official vllm/vllm-openai:v0.14.1 image as the base container. Mount the VibeVoice repository into /app and execute vllm_plugin/scripts/start_server.py as the entrypoint command. This image provides the necessary CUDA runtime and Python environment to run the vLLM inference engine with GPU support.
How do I enable multi-GPU scaling for VibeVoice-ASR inference?
Pass the --dp N flag to start_server.py for Data Parallel deployment, which creates N independent vLLM workers behind an nginx reverse proxy for load balancing. For Tensor Parallel model sharding across GPUs when memory is insufficient, use --tp N. You can combine both flags (e.g., --dp 2 --tp 2 on a 4-GPU system) to create replicated, sharded model instances.
What is the purpose of the tokenizer generation step in VibeVoice-ASR deployment?
The vllm_plugin/tools/generate_tokenizer_files.py script downloads the base Qwen2.5 tokenizer and patches it with VibeVoice-specific audio tokens, generating six required JSON and text files (tokenizer.json, vocab.json, etc.). This step is mandatory because vLLM requires these specific tokenizer artifacts to correctly encode audio inputs alongside text prompts for the ASR model.
How does VibeVoice-ASR handle very long audio files without repetition loops?
The deployment includes an auto-recovery mechanism tested via vllm_plugin/tests/test_api_auto_recover.py. When processing audio exceeding 60 minutes, the vLLM server can detect and recover from repetition loops that sometimes occur during streaming generation. The system maintains stability by managing KV cache allocation through PYTORCH_ALLOC_CONF=expandable_segments:True and supporting context lengths up to 131,072 tokens after tokenizer patching.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →