How to Integrate VibeVoice with Existing Speech Pipelines: A vLLM Multimodal Guide

Integrate VibeVoice into existing speech pipelines by loading 24 kHz mono audio via FFmpeg utilities, initializing a vLLM server with the VibeVoiceForCausalLM model class, and submitting requests that pair raw waveforms with text prompts containing the <|AUDIO|> placeholder token.

VibeVoice provides a native vLLM multimodal plugin that transforms raw audio waveforms into token embeddings compatible with any causal language model, such as Qwen-2.5. According to the microsoft/VibeVoice source code, the integration centers on the VibeVoiceMultiModalProcessor class, which automatically handles audio-token expansion and embedding fusion within the vLLM inference engine. This guide demonstrates how to integrate VibeVoice with existing speech pipelines using the actual implementation files and APIs found in the repository.

Prerequisites and Installation

Begin by cloning the repository and installing the package in editable mode to ensure you can modify configurations if needed. You must also pull the model weights using Git LFS.


# Clone the VibeVoice repository

git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice

# Install the package in editable mode

pip install -e .

# Pull model weights (example: 7B ASR variant)

git lfs install
git clone https://huggingface.co/microsoft/VibeVoice-ASR model_dir

All subsequent paths assume the repository root as your working directory.

Loading and Normalizing Audio Input

VibeVoice requires 24 kHz mono waveforms stored as float32 arrays in the range [-1, 1]. The vibevoice/processor/audio_utils.py file provides an FFmpeg-based loader that handles container decoding, resampling, and loudness normalization automatically.

from vibevoice.processor.audio_utils import load_audio_use_ffmpeg

# Load and resample to 24 kHz mono

audio_path = "my_audio.wav"
waveform, sr = load_audio_use_ffmpeg(
    audio_path, 
    resample=True, 
    target_sr=24000
)

# waveform is now a NumPy float32 array ready for encoding

Implementation reference: vibevoice/processor/audio_utils.py defines load_audio_use_ffmpeg and AudioNormalizer (lines 24-78, 49-66).

Initializing the vLLM Server with VibeVoice

The VibeVoice plugin registers itself with vLLM via the @MULTIMODAL_REGISTRY.register_processor decorator found in vllm_plugin/model.py. When you initialize the LLM class with multimodal=True and trust_remote_code=True, vLLM automatically loads the VibeVoiceForCausalLM model class and its associated processor.

from vllm import LLM, SamplingParams

# Initialize the vLLM server with VibeVoice ASR

llm = LLM(
    model="microsoft/VibeVoice-ASR",  # HF repo or local directory

    dtype="bfloat16",                   # Matches LM dtype; encoder remains fp32

    trust_remote_code=True,             # Required for custom model class

    multimodal=True,                    # Enables multimodal support

)

sampling_params = SamplingParams(max_tokens=256)

Key registration: vllm_plugin/model.py lines 24-28 handle the MULTIMODAL_REGISTRY.register_processor call that binds the VibeVoiceMultiModalProcessor to the inference engine.

Constructing Multimodal Prompts with the Audio Placeholder

VibeVoice uses a single placeholder token, <|AUDIO|>, which the processor expands into a sequence of speech-start, speech-pad, and speech-end tokens. The number of pad tokens depends on the audio duration and the model's speech_tok_compress_ratio (default 3200).


# Construct prompt with exactly one audio placeholder

prompt = "Transcribe the following audio:\n<|AUDIO|>"

When the request reaches the server, VibeVoiceMultiModalProcessor._call_hf_processor (lines 46-84) performs three critical operations:

  1. Tokenizes the prompt without expanding the placeholder.
  2. Stacks raw audio tensors into raw_audio and records lengths in raw_audio_lengths.
  3. Appends a random salt to prevent cache collisions across requests.

Subsequently, _get_prompt_updates (lines 108-124) replaces <|AUDIO|> with the exact token count required for the specific audio segment.

Executing Inference and Embedding Merge

Submit requests as dictionaries containing the prompt string and the raw waveform. The waveform must be JSON-serializable, so convert NumPy arrays to Python lists before transmission.

import numpy as np

# Prepare JSON-friendly payload

request_payload = {
    "prompt": prompt,
    "audio": waveform.tolist()  # vLLM reconverts to tensor internally

}

# Execute synchronous inference

outputs = llm.generate([request_payload], sampling_params)
transcription = outputs[0].outputs[0].text
print(transcription)

During server-side processing, the pipeline executes the following sequence defined in vllm_plugin/model.py:

  • embed_multimodal (lines 98-124) invokes the VibeVoiceAudioEncoder to convert the raw waveform into acoustic and semantic token embeddings.
  • embed_input_ids (lines 149-164) merges these audio embeddings with the text embeddings at the correct positions before the language model's forward pass.

Handling Long-Form Audio with Streaming

For audio exceeding the default 60-second segment length, enable streaming in the audio encoder to process chunks sequentially and avoid hitting the 61-minute hard limit imposed by VibeVoiceAudioEncoder.get_mm_max_tokens_per_item.

import torch

# Enable streaming for long-form audio

embeds = model.audio_encoder(
    torch.from_numpy(waveform).unsqueeze(0),  # Shape: [1, T]

    use_streaming=True,                         # Activates segmentation

    segment_duration_s=60.0                     # Chunk size in seconds

)

Configuration: Set enable_streaming and streaming_segment_duration in the encoder config, or pass them directly to VibeVoiceAudioEncoder.__init__ (lines 55-60).

End-to-End Integration Example

The following complete example demonstrates loading audio, initializing the vLLM server, and executing transcription in a single Python script:

import torch
from vllm import LLM, SamplingParams
from vibevoice.processor.audio_utils import load_audio_use_ffmpeg

# Step 1: Load and normalize audio to 24 kHz mono

waveform, _ = load_audio_use_ffmpeg(
    "sample.wav", 
    resample=True, 
    target_sr=24000
)

# Step 2: Initialize vLLM with VibeVoice ASR capabilities

llm = LLM(
    model="microsoft/VibeVoice-ASR",
    dtype="bfloat16",
    trust_remote_code=True,
    multimodal=True,
)
params = SamplingParams(max_tokens=512)

# Step 3: Build prompt with audio placeholder

prompt = "Please transcribe the following audio:\n<|AUDIO|>"

# Step 4: Send request with raw waveform as list

request = {"prompt": prompt, "audio": waveform.tolist()}
result = llm.generate([request], params)

# Step 5: Extract and display transcription

print("=== Transcription ===")
print(result[0].outputs[0].text)

This implementation runs entirely server-side; the client only transmits the raw waveform and prompt string.

Key Source Files and Architecture

Understanding the source layout ensures accurate debugging and customization:

  • vibevoice/processor/audio_utils.py – Implements load_audio_use_ffmpeg and AudioNormalizer for 24 kHz conversion and loudness normalization.
  • vibevoice/modular/modular_vibevoice_tokenizer.py – Contains VAE-based acoustic and semantic tokenizers consumed by the encoder.
  • vllm_plugin/model.py – Houses the core integration classes:
    • VibeVoiceAudioEncoder (lines 68-140): Converts waveforms to embeddings with optional streaming.
    • VibeVoiceMultiModalProcessor (lines 46-84, 108-124): Manages the <|AUDIO|> placeholder expansion and raw audio stacking.
    • VibeVoiceForCausalLM (lines 24-28): Registers the processor with vLLM via MULTIMODAL_REGISTRY.
  • demo/vibevoice_asr_gradio_demo.py – Reference implementation showing Gradio-based ASR inference.
  • vibevoice/configs/qwen2.5_1.5b_64k.json – Configuration file specifying speech_tok_compress_ratio and other hyperparameters.

Summary

  • Resample to 24 kHz mono using load_audio_use_ffmpeg from vibevoice/processor/audio_utils.py to meet encoder requirements.
  • Initialize vLLM with multimodal=True and trust_remote_code=True to activate the VibeVoiceMultiModalProcessor registration.
  • Use the <|AUDIO|> placeholder exactly once per audio segment; the processor automatically expands it based on the speech_tok_compress_ratio (default 3200).
  • Transmit raw waveforms as JSON-serializable lists in the request payload; the server reconverts them to tensors via embed_multimodal.
  • Enable streaming via use_streaming=True in VibeVoiceAudioEncoder for audio longer than 60 seconds to support transcriptions up to 61 minutes.

Frequently Asked Questions

What audio format does VibeVoice require for integration?

VibeVoice requires 24 kHz mono waveforms as float32 NumPy arrays in the range [-1, 1]. The load_audio_use_ffmpeg function in vibevoice/processor/audio_utils.py automatically handles resampling and loudness normalization from any common audio container.

How does the <|AUDIO|> placeholder work in prompts?

The VibeVoiceMultiModalProcessor._get_prompt_updates method (lines 108-124 in vllm_plugin/model.py) dynamically expands the single <|AUDIO|> token into a sequence of speech-start, speech-pad, and speech-end tokens. The pad token count calculates from the audio length divided by the speech_tok_compress_ratio (default 3200) defined in the model configuration.

Can VibeVoice handle long-form audio transcription beyond 60 seconds?

Yes. For audio exceeding 60 seconds, invoke the VibeVoiceAudioEncoder with use_streaming=True and specify segment_duration_s (default 60.0) to process audio in chunks. This streaming mechanism supports inputs up to the 61-minute limit defined by get_mm_max_tokens_per_item.

Is batch processing supported for multiple audio files?

Yes. The VibeVoiceMultiModalProcessor._call_hf_processor method accepts a list of waveforms and stacks them into a single raw_audio tensor with corresponding raw_audio_lengths. Send multiple requests in a list to the llm.generate() method to process batches efficiently.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →