How to Integrate VibeVoice with Existing Speech Pipelines: A vLLM Multimodal Guide
Integrate VibeVoice into existing speech pipelines by loading 24 kHz mono audio via FFmpeg utilities, initializing a vLLM server with the VibeVoiceForCausalLM model class, and submitting requests that pair raw waveforms with text prompts containing the <|AUDIO|> placeholder token.
VibeVoice provides a native vLLM multimodal plugin that transforms raw audio waveforms into token embeddings compatible with any causal language model, such as Qwen-2.5. According to the microsoft/VibeVoice source code, the integration centers on the VibeVoiceMultiModalProcessor class, which automatically handles audio-token expansion and embedding fusion within the vLLM inference engine. This guide demonstrates how to integrate VibeVoice with existing speech pipelines using the actual implementation files and APIs found in the repository.
Prerequisites and Installation
Begin by cloning the repository and installing the package in editable mode to ensure you can modify configurations if needed. You must also pull the model weights using Git LFS.
# Clone the VibeVoice repository
git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice
# Install the package in editable mode
pip install -e .
# Pull model weights (example: 7B ASR variant)
git lfs install
git clone https://huggingface.co/microsoft/VibeVoice-ASR model_dir
All subsequent paths assume the repository root as your working directory.
Loading and Normalizing Audio Input
VibeVoice requires 24 kHz mono waveforms stored as float32 arrays in the range [-1, 1]. The vibevoice/processor/audio_utils.py file provides an FFmpeg-based loader that handles container decoding, resampling, and loudness normalization automatically.
from vibevoice.processor.audio_utils import load_audio_use_ffmpeg
# Load and resample to 24 kHz mono
audio_path = "my_audio.wav"
waveform, sr = load_audio_use_ffmpeg(
audio_path,
resample=True,
target_sr=24000
)
# waveform is now a NumPy float32 array ready for encoding
Implementation reference: vibevoice/processor/audio_utils.py defines load_audio_use_ffmpeg and AudioNormalizer (lines 24-78, 49-66).
Initializing the vLLM Server with VibeVoice
The VibeVoice plugin registers itself with vLLM via the @MULTIMODAL_REGISTRY.register_processor decorator found in vllm_plugin/model.py. When you initialize the LLM class with multimodal=True and trust_remote_code=True, vLLM automatically loads the VibeVoiceForCausalLM model class and its associated processor.
from vllm import LLM, SamplingParams
# Initialize the vLLM server with VibeVoice ASR
llm = LLM(
model="microsoft/VibeVoice-ASR", # HF repo or local directory
dtype="bfloat16", # Matches LM dtype; encoder remains fp32
trust_remote_code=True, # Required for custom model class
multimodal=True, # Enables multimodal support
)
sampling_params = SamplingParams(max_tokens=256)
Key registration: vllm_plugin/model.py lines 24-28 handle the MULTIMODAL_REGISTRY.register_processor call that binds the VibeVoiceMultiModalProcessor to the inference engine.
Constructing Multimodal Prompts with the Audio Placeholder
VibeVoice uses a single placeholder token, <|AUDIO|>, which the processor expands into a sequence of speech-start, speech-pad, and speech-end tokens. The number of pad tokens depends on the audio duration and the model's speech_tok_compress_ratio (default 3200).
# Construct prompt with exactly one audio placeholder
prompt = "Transcribe the following audio:\n<|AUDIO|>"
When the request reaches the server, VibeVoiceMultiModalProcessor._call_hf_processor (lines 46-84) performs three critical operations:
- Tokenizes the prompt without expanding the placeholder.
- Stacks raw audio tensors into
raw_audioand records lengths inraw_audio_lengths. - Appends a random salt to prevent cache collisions across requests.
Subsequently, _get_prompt_updates (lines 108-124) replaces <|AUDIO|> with the exact token count required for the specific audio segment.
Executing Inference and Embedding Merge
Submit requests as dictionaries containing the prompt string and the raw waveform. The waveform must be JSON-serializable, so convert NumPy arrays to Python lists before transmission.
import numpy as np
# Prepare JSON-friendly payload
request_payload = {
"prompt": prompt,
"audio": waveform.tolist() # vLLM reconverts to tensor internally
}
# Execute synchronous inference
outputs = llm.generate([request_payload], sampling_params)
transcription = outputs[0].outputs[0].text
print(transcription)
During server-side processing, the pipeline executes the following sequence defined in vllm_plugin/model.py:
embed_multimodal(lines 98-124) invokes theVibeVoiceAudioEncoderto convert the raw waveform into acoustic and semantic token embeddings.embed_input_ids(lines 149-164) merges these audio embeddings with the text embeddings at the correct positions before the language model's forward pass.
Handling Long-Form Audio with Streaming
For audio exceeding the default 60-second segment length, enable streaming in the audio encoder to process chunks sequentially and avoid hitting the 61-minute hard limit imposed by VibeVoiceAudioEncoder.get_mm_max_tokens_per_item.
import torch
# Enable streaming for long-form audio
embeds = model.audio_encoder(
torch.from_numpy(waveform).unsqueeze(0), # Shape: [1, T]
use_streaming=True, # Activates segmentation
segment_duration_s=60.0 # Chunk size in seconds
)
Configuration: Set enable_streaming and streaming_segment_duration in the encoder config, or pass them directly to VibeVoiceAudioEncoder.__init__ (lines 55-60).
End-to-End Integration Example
The following complete example demonstrates loading audio, initializing the vLLM server, and executing transcription in a single Python script:
import torch
from vllm import LLM, SamplingParams
from vibevoice.processor.audio_utils import load_audio_use_ffmpeg
# Step 1: Load and normalize audio to 24 kHz mono
waveform, _ = load_audio_use_ffmpeg(
"sample.wav",
resample=True,
target_sr=24000
)
# Step 2: Initialize vLLM with VibeVoice ASR capabilities
llm = LLM(
model="microsoft/VibeVoice-ASR",
dtype="bfloat16",
trust_remote_code=True,
multimodal=True,
)
params = SamplingParams(max_tokens=512)
# Step 3: Build prompt with audio placeholder
prompt = "Please transcribe the following audio:\n<|AUDIO|>"
# Step 4: Send request with raw waveform as list
request = {"prompt": prompt, "audio": waveform.tolist()}
result = llm.generate([request], params)
# Step 5: Extract and display transcription
print("=== Transcription ===")
print(result[0].outputs[0].text)
This implementation runs entirely server-side; the client only transmits the raw waveform and prompt string.
Key Source Files and Architecture
Understanding the source layout ensures accurate debugging and customization:
vibevoice/processor/audio_utils.py– Implementsload_audio_use_ffmpegandAudioNormalizerfor 24 kHz conversion and loudness normalization.vibevoice/modular/modular_vibevoice_tokenizer.py– Contains VAE-based acoustic and semantic tokenizers consumed by the encoder.vllm_plugin/model.py– Houses the core integration classes:VibeVoiceAudioEncoder(lines 68-140): Converts waveforms to embeddings with optional streaming.VibeVoiceMultiModalProcessor(lines 46-84, 108-124): Manages the<|AUDIO|>placeholder expansion and raw audio stacking.VibeVoiceForCausalLM(lines 24-28): Registers the processor with vLLM viaMULTIMODAL_REGISTRY.
demo/vibevoice_asr_gradio_demo.py– Reference implementation showing Gradio-based ASR inference.vibevoice/configs/qwen2.5_1.5b_64k.json– Configuration file specifyingspeech_tok_compress_ratioand other hyperparameters.
Summary
- Resample to 24 kHz mono using
load_audio_use_ffmpegfromvibevoice/processor/audio_utils.pyto meet encoder requirements. - Initialize vLLM with
multimodal=Trueandtrust_remote_code=Trueto activate theVibeVoiceMultiModalProcessorregistration. - Use the
<|AUDIO|>placeholder exactly once per audio segment; the processor automatically expands it based on thespeech_tok_compress_ratio(default 3200). - Transmit raw waveforms as JSON-serializable lists in the request payload; the server reconverts them to tensors via
embed_multimodal. - Enable streaming via
use_streaming=TrueinVibeVoiceAudioEncoderfor audio longer than 60 seconds to support transcriptions up to 61 minutes.
Frequently Asked Questions
What audio format does VibeVoice require for integration?
VibeVoice requires 24 kHz mono waveforms as float32 NumPy arrays in the range [-1, 1]. The load_audio_use_ffmpeg function in vibevoice/processor/audio_utils.py automatically handles resampling and loudness normalization from any common audio container.
How does the <|AUDIO|> placeholder work in prompts?
The VibeVoiceMultiModalProcessor._get_prompt_updates method (lines 108-124 in vllm_plugin/model.py) dynamically expands the single <|AUDIO|> token into a sequence of speech-start, speech-pad, and speech-end tokens. The pad token count calculates from the audio length divided by the speech_tok_compress_ratio (default 3200) defined in the model configuration.
Can VibeVoice handle long-form audio transcription beyond 60 seconds?
Yes. For audio exceeding 60 seconds, invoke the VibeVoiceAudioEncoder with use_streaming=True and specify segment_duration_s (default 60.0) to process audio in chunks. This streaming mechanism supports inputs up to the 61-minute limit defined by get_mm_max_tokens_per_item.
Is batch processing supported for multiple audio files?
Yes. The VibeVoiceMultiModalProcessor._call_hf_processor method accepts a list of waveforms and stacks them into a single raw_audio tensor with corresponding raw_audio_lengths. Send multiple requests in a list to the llm.generate() method to process batches efficiently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →