# How to Integrate VibeVoice with Existing Speech Pipelines: A vLLM Multimodal Guide

> Integrate VibeVoice into speech pipelines using vLLM. Load 24 kHz audio via FFmpeg, initialize vLLM server with VibeVoiceForCausalLM, and submit requests pairing waveforms with text prompts containing the <|AUDIO|> token.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**Integrate VibeVoice into existing speech pipelines by loading 24 kHz mono audio via FFmpeg utilities, initializing a vLLM server with the `VibeVoiceForCausalLM` model class, and submitting requests that pair raw waveforms with text prompts containing the `<|AUDIO|>` placeholder token.**

VibeVoice provides a native vLLM multimodal plugin that transforms raw audio waveforms into token embeddings compatible with any causal language model, such as Qwen-2.5. According to the `microsoft/VibeVoice` source code, the integration centers on the `VibeVoiceMultiModalProcessor` class, which automatically handles audio-token expansion and embedding fusion within the vLLM inference engine. This guide demonstrates how to integrate VibeVoice with existing speech pipelines using the actual implementation files and APIs found in the repository.

## Prerequisites and Installation

Begin by cloning the repository and installing the package in editable mode to ensure you can modify configurations if needed. You must also pull the model weights using Git LFS.

```bash

# Clone the VibeVoice repository

git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice

# Install the package in editable mode

pip install -e .

# Pull model weights (example: 7B ASR variant)

git lfs install
git clone https://huggingface.co/microsoft/VibeVoice-ASR model_dir

```

All subsequent paths assume the repository root as your working directory.

## Loading and Normalizing Audio Input

VibeVoice requires **24 kHz mono** waveforms stored as **float32** arrays in the range `[-1, 1]`. The [`vibevoice/processor/audio_utils.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/audio_utils.py) file provides an FFmpeg-based loader that handles container decoding, resampling, and loudness normalization automatically.

```python
from vibevoice.processor.audio_utils import load_audio_use_ffmpeg

# Load and resample to 24 kHz mono

audio_path = "my_audio.wav"
waveform, sr = load_audio_use_ffmpeg(
    audio_path, 
    resample=True, 
    target_sr=24000
)

# waveform is now a NumPy float32 array ready for encoding

```

*Implementation reference:* [`vibevoice/processor/audio_utils.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/audio_utils.py) defines `load_audio_use_ffmpeg` and `AudioNormalizer` (lines 24-78, 49-66).

## Initializing the vLLM Server with VibeVoice

The VibeVoice plugin registers itself with vLLM via the `@MULTIMODAL_REGISTRY.register_processor` decorator found in [`vllm_plugin/model.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/model.py). When you initialize the `LLM` class with `multimodal=True` and `trust_remote_code=True`, vLLM automatically loads the `VibeVoiceForCausalLM` model class and its associated processor.

```python
from vllm import LLM, SamplingParams

# Initialize the vLLM server with VibeVoice ASR

llm = LLM(
    model="microsoft/VibeVoice-ASR",  # HF repo or local directory

    dtype="bfloat16",                   # Matches LM dtype; encoder remains fp32

    trust_remote_code=True,             # Required for custom model class

    multimodal=True,                    # Enables multimodal support

)

sampling_params = SamplingParams(max_tokens=256)

```

*Key registration:* [`vllm_plugin/model.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/model.py) lines 24-28 handle the `MULTIMODAL_REGISTRY.register_processor` call that binds the `VibeVoiceMultiModalProcessor` to the inference engine.

## Constructing Multimodal Prompts with the Audio Placeholder

VibeVoice uses a single placeholder token, `<|AUDIO|>`, which the processor expands into a sequence of speech-start, speech-pad, and speech-end tokens. The number of pad tokens depends on the audio duration and the model's `speech_tok_compress_ratio` (default **3200**).

```python

# Construct prompt with exactly one audio placeholder

prompt = "Transcribe the following audio:\n<|AUDIO|>"

```

When the request reaches the server, `VibeVoiceMultiModalProcessor._call_hf_processor` (lines 46-84) performs three critical operations:

1. **Tokenizes** the prompt without expanding the placeholder.
2. **Stacks** raw audio tensors into `raw_audio` and records lengths in `raw_audio_lengths`.
3. **Appends a random salt** to prevent cache collisions across requests.

Subsequently, `_get_prompt_updates` (lines 108-124) replaces `<|AUDIO|>` with the exact token count required for the specific audio segment.

## Executing Inference and Embedding Merge

Submit requests as dictionaries containing the prompt string and the raw waveform. The waveform must be JSON-serializable, so convert NumPy arrays to Python lists before transmission.

```python
import numpy as np

# Prepare JSON-friendly payload

request_payload = {
    "prompt": prompt,
    "audio": waveform.tolist()  # vLLM reconverts to tensor internally

}

# Execute synchronous inference

outputs = llm.generate([request_payload], sampling_params)
transcription = outputs[0].outputs[0].text
print(transcription)

```

During server-side processing, the pipeline executes the following sequence defined in [`vllm_plugin/model.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/model.py):

- **`embed_multimodal`** (lines 98-124) invokes the `VibeVoiceAudioEncoder` to convert the raw waveform into acoustic and semantic token embeddings.
- **`embed_input_ids`** (lines 149-164) merges these audio embeddings with the text embeddings at the correct positions before the language model's forward pass.

## Handling Long-Form Audio with Streaming

For audio exceeding the default 60-second segment length, enable streaming in the audio encoder to process chunks sequentially and avoid hitting the 61-minute hard limit imposed by `VibeVoiceAudioEncoder.get_mm_max_tokens_per_item`.

```python
import torch

# Enable streaming for long-form audio

embeds = model.audio_encoder(
    torch.from_numpy(waveform).unsqueeze(0),  # Shape: [1, T]

    use_streaming=True,                         # Activates segmentation

    segment_duration_s=60.0                     # Chunk size in seconds

)

```

*Configuration:* Set `enable_streaming` and `streaming_segment_duration` in the encoder config, or pass them directly to `VibeVoiceAudioEncoder.__init__` (lines 55-60).

## End-to-End Integration Example

The following complete example demonstrates loading audio, initializing the vLLM server, and executing transcription in a single Python script:

```python
import torch
from vllm import LLM, SamplingParams
from vibevoice.processor.audio_utils import load_audio_use_ffmpeg

# Step 1: Load and normalize audio to 24 kHz mono

waveform, _ = load_audio_use_ffmpeg(
    "sample.wav", 
    resample=True, 
    target_sr=24000
)

# Step 2: Initialize vLLM with VibeVoice ASR capabilities

llm = LLM(
    model="microsoft/VibeVoice-ASR",
    dtype="bfloat16",
    trust_remote_code=True,
    multimodal=True,
)
params = SamplingParams(max_tokens=512)

# Step 3: Build prompt with audio placeholder

prompt = "Please transcribe the following audio:\n<|AUDIO|>"

# Step 4: Send request with raw waveform as list

request = {"prompt": prompt, "audio": waveform.tolist()}
result = llm.generate([request], params)

# Step 5: Extract and display transcription

print("=== Transcription ===")
print(result[0].outputs[0].text)

```

This implementation runs entirely server-side; the client only transmits the raw waveform and prompt string.

## Key Source Files and Architecture

Understanding the source layout ensures accurate debugging and customization:

- **[`vibevoice/processor/audio_utils.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/audio_utils.py)** – Implements `load_audio_use_ffmpeg` and `AudioNormalizer` for 24 kHz conversion and loudness normalization.
- **[`vibevoice/modular/modular_vibevoice_tokenizer.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modular_vibevoice_tokenizer.py)** – Contains VAE-based acoustic and semantic tokenizers consumed by the encoder.
- **[`vllm_plugin/model.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/model.py)** – Houses the core integration classes:
  - `VibeVoiceAudioEncoder` (lines 68-140): Converts waveforms to embeddings with optional streaming.
  - `VibeVoiceMultiModalProcessor` (lines 46-84, 108-124): Manages the `<|AUDIO|>` placeholder expansion and raw audio stacking.
  - `VibeVoiceForCausalLM` (lines 24-28): Registers the processor with vLLM via `MULTIMODAL_REGISTRY`.
- **[`demo/vibevoice_asr_gradio_demo.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_asr_gradio_demo.py)** – Reference implementation showing Gradio-based ASR inference.
- **[`vibevoice/configs/qwen2.5_1.5b_64k.json`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/configs/qwen2.5_1.5b_64k.json)** – Configuration file specifying `speech_tok_compress_ratio` and other hyperparameters.

## Summary

- **Resample to 24 kHz mono** using `load_audio_use_ffmpeg` from [`vibevoice/processor/audio_utils.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/audio_utils.py) to meet encoder requirements.
- **Initialize vLLM** with `multimodal=True` and `trust_remote_code=True` to activate the `VibeVoiceMultiModalProcessor` registration.
- **Use the `<|AUDIO|>` placeholder** exactly once per audio segment; the processor automatically expands it based on the `speech_tok_compress_ratio` (default 3200).
- **Transmit raw waveforms** as JSON-serializable lists in the request payload; the server reconverts them to tensors via `embed_multimodal`.
- **Enable streaming** via `use_streaming=True` in `VibeVoiceAudioEncoder` for audio longer than 60 seconds to support transcriptions up to 61 minutes.

## Frequently Asked Questions

### What audio format does VibeVoice require for integration?

VibeVoice requires **24 kHz mono** waveforms as **float32** NumPy arrays in the range `[-1, 1]`. The `load_audio_use_ffmpeg` function in [`vibevoice/processor/audio_utils.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/audio_utils.py) automatically handles resampling and loudness normalization from any common audio container.

### How does the `<|AUDIO|>` placeholder work in prompts?

The `VibeVoiceMultiModalProcessor._get_prompt_updates` method (lines 108-124 in [`vllm_plugin/model.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/model.py)) dynamically expands the single `<|AUDIO|>` token into a sequence of speech-start, speech-pad, and speech-end tokens. The pad token count calculates from the audio length divided by the `speech_tok_compress_ratio` (default 3200) defined in the model configuration.

### Can VibeVoice handle long-form audio transcription beyond 60 seconds?

Yes. For audio exceeding 60 seconds, invoke the `VibeVoiceAudioEncoder` with `use_streaming=True` and specify `segment_duration_s` (default 60.0) to process audio in chunks. This streaming mechanism supports inputs up to the 61-minute limit defined by `get_mm_max_tokens_per_item`.

### Is batch processing supported for multiple audio files?

Yes. The `VibeVoiceMultiModalProcessor._call_hf_processor` method accepts a list of waveforms and stacks them into a single `raw_audio` tensor with corresponding `raw_audio_lengths`. Send multiple requests in a list to the `llm.generate()` method to process batches efficiently.