# How to Integrate VibeVoice with Hugging Face Transformers: Complete Setup and ASR Guide

> Integrate VibeVoice with Hugging Face Transformers using auto factory methods. Get a complete setup and ASR guide for seamless integration.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**You can integrate VibeVoice with Hugging Face Transformers by using the standard `Auto*` factory methods—such as `AutoModelForCausalLM.from_pretrained()`—which automatically resolve to VibeVoice-specific classes thanks to registry entries in [`vllm_plugin/__init__.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/__init__.py).**

The Microsoft VibeVoice repository ships as a fully Hugging Face Transformers-compatible package, providing composite configs, custom tokenizers, and processor classes that enable seamless speech-to-text inference. Whether you are building batch ASR pipelines or real-time streaming applications, you can load and run VibeVoice using the familiar Transformers APIs without custom wrappers.

## Core Integration Components

VibeVoice exposes four primary classes that plug into the Transformers ecosystem. These components are registered automatically when you import the package, allowing seamless use of `AutoConfig`, `AutoTokenizer`, `AutoProcessor`, and `AutoModelForCausalLM`.

- **VibeVoiceConfig** – Defined in [`vibevoice/modular/configuration_vibevoice.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/configuration_vibevoice.py), this composite `PretrainedConfig` aggregates acoustic, semantic, and Qwen2 decoder sub-configurations. It tells the model how to initialize the acoustic tokenizer and diffusion head.

- **VibeVoiceTextTokenizerFast** – Located in [`vibevoice/modular/modular_vibevoice_text_tokenizer.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modular_vibevoice_text_tokenizer.py), this fast tokenizer inherits from `Qwen2TokenizerFast` and injects three speech-specific special tokens: `<|vision_start|>`, `<|vision_end|>`, and `<|vision_pad|>`.

- **VibeVoiceProcessor** and **VibeVoiceASRProcessor** – Implemented in [`vibevoice/processor/vibevoice_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_processor.py), these processors wrap the acoustic tokenizer and handle audio normalization. They combine feature extraction with text tokenization to produce model-ready inputs.

- **VibeVoiceForConditionalGeneration** – Found in [`vibevoice/modular/modeling_vibevoice.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice.py), this is the core model class that performs both text generation and speech-to-text (ASR) when acoustic tensors are supplied.

## Loading Models with Auto Factories

Because the package registers its components via [`vllm_plugin/__init__.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/__init__.py), you can load the entire stack using standard Transformers factories. This eliminates the need to import concrete classes manually unless you require specific functionality.

```python
from transformers import AutoConfig, AutoTokenizer, AutoProcessor, AutoModelForCausalLM

model_id = "microsoft/VibeVoice-ASR"

config = AutoConfig.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

```

The `AutoConfig` call resolves to `VibeVoiceConfig`, while `AutoTokenizer` returns `VibeVoiceASRTextTokenizerFast` (or the base `VibeVoiceTextTokenizerFast` depending on the checkpoint). The processor resolves to `VibeVoiceProcessor` and the model instantiates `VibeVoiceForConditionalGeneration` with automatic mixed-precision and device placement.

## Performing Speech-to-Text Inference

To run ASR, you must pass acoustic features via the `speech_tensors` parameter. The processor converts raw audio into the expected feature format, while the model requires a text prompt containing the speech-start token to align acoustic embeddings with text embeddings.

```python
import torch
import torchaudio
from pathlib import Path

# Load audio

audio_path = Path("sample.wav")
waveform, sr = torchaudio.load(audio_path)

# Pre-process: extract acoustic tokens and normalize

inputs = processor(
    audio=waveform.squeeze(0),
    sampling_rate=sr,
    return_tensors="pt",
)

# Prepare text prompt with speech-start token

input_ids = tokenizer("<|vision_start|>", return_tensors="pt")["input_ids"]

# Forward pass

outputs = model(
    input_ids=input_ids,
    speech_tensors=inputs["input_features"],
    speech_masks=inputs["attention_mask"].bool(),
    acoustic_input_mask=torch.arange(input_ids.shape[1]) == tokenizer.speech_start_id,
    return_dict=True,
)

# Decode to text

generated_ids = outputs.logits.argmax(-1)
transcription = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
print(transcription)

```

**Key parameters explained:**

- `speech_tensors` – The acoustic token sequence produced by `VibeVoiceTokenizerProcessor`.
- `acoustic_input_mask` – A boolean mask indicating which positions in `input_ids` should be replaced with acoustic embeddings (typically the position of `<|vision_start|>`).
- `speech_masks` – Attention mask for the acoustic features, ensuring the model ignores padded frames.

## Real-Time Streaming Inference

For low-latency applications, VibeVoice provides a streaming processor that maintains internal caches, mirroring the KV-cache mechanism used in language models. This allows chunk-wise processing of live audio without recomputing features for the entire history.

```python
from vibevoice.processor.vibevoice_streaming_processor import VibeVoiceStreamingProcessor

streamer = VibeVoiceStreamingProcessor.from_pretrained(model_id)

for chunk in microphone_stream():
    batch = streamer(chunk, sampling_rate=sr)
    
    out = model(
        input_ids=batch["input_ids"],
        speech_tensors=batch["input_features"],
        speech_masks=batch["attention_mask"].bool(),
        acoustic_input_mask=batch["acoustic_input_mask"],
        past_key_values=streamer.past_key_values,
        use_cache=True,
    )
    
    streamer.update_cache(out.past_key_values)
    new_tokens = out.logits.argmax(-1)
    print(tokenizer.decode(new_tokens[0], skip_special_tokens=True), end="", flush=True)

```

The `VibeVoiceStreamingProcessor` handles padding alignment and cache management automatically, returning ready-to-use tensors for each audio fragment.

## Using the Transformers Pipeline

You can also deploy VibeVoice through the high-level `pipeline` abstraction for automatic speech recognition. Because `VibeVoiceProcessor` implements the standard feature-extractor interface, it works out-of-the-box with the pipeline API.

```python
from transformers import pipeline

asr_pipe = pipeline(
    "automatic-speech-recognition",
    model=model,
    tokenizer=tokenizer,
    feature_extractor=processor,
    device=0,
)

result = asr_pipe("sample.wav")
print(result["text"])

```

This approach is ideal for rapid prototyping or batch processing, as the pipeline handles batching and device placement internally.

## Summary

- **Registry-based loading**: Import the vibevoice package to automatically register `VibeVoiceConfig`, tokenizer, processor, and model classes with Transformers `Auto*` factories.
- **Audio input**: Use `VibeVoiceProcessor` to convert raw audio to `speech_tensors`, then pass these alongside text prompts containing `<|vision_start|>`.
- **Inference modes**: Choose between standard batch inference for file-based ASR or `VibeVoiceStreamingProcessor` for real-time, low-latency transcription.
- **Pipeline compatibility**: Deploy immediately using `pipeline("automatic-speech-recognition")` by supplying the model, tokenizer, and processor as initialized components.

## Frequently Asked Questions

### Can I load VibeVoice using standard Hugging Face Auto classes?

Yes. Once you install the package, `AutoConfig.from_pretrained()`, `AutoTokenizer.from_pretrained()`, and `AutoModelForCausalLM.from_pretrained()` will resolve to `VibeVoiceConfig`, `VibeVoiceTextTokenizerFast`, and `VibeVoiceForConditionalGeneration` respectively. This registration happens automatically through the plugin initialization in [`vllm_plugin/__init__.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/__init__.py).

### What special tokens does VibeVoice use for speech processing?

The tokenizer reserves three speech-specific tokens: `<|vision_start|>` marks the insertion point for acoustic embeddings, `<|vision_end|>` terminates the speech segment, and `<|vision_pad|>` handles padding. You must include `<|vision_start|>` in your text prompt so the model knows where to inject the `speech_tensors`.

### How do I handle streaming or real-time audio input?

Use `VibeVoiceStreamingProcessor` from [`vibevoice/processor/vibevoice_streaming_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_streaming_processor.py). This class maintains an internal cache and accepts audio chunks incrementally. Pass `past_key_values=streamer.past_key_values` and `use_cache=True` during generation to enable efficient incremental decoding without recomputing attention over the full audio history.

### Is VibeVoice compatible with vLLM for accelerated serving?

Yes. The repository includes a vLLM plugin that registers the model architecture with the vLLM engine. When you import the package, `register_vibevoice()` is called automatically through the `vllm.general_plugins` entry point, making VibeVoice discoverable by the vLLM server for optimized batch inference and continuous batching.