How to Integrate VibeVoice with Hugging Face Transformers: Complete Setup and ASR Guide
You can integrate VibeVoice with Hugging Face Transformers by using the standard Auto* factory methods—such as AutoModelForCausalLM.from_pretrained()—which automatically resolve to VibeVoice-specific classes thanks to registry entries in vllm_plugin/__init__.py.
The Microsoft VibeVoice repository ships as a fully Hugging Face Transformers-compatible package, providing composite configs, custom tokenizers, and processor classes that enable seamless speech-to-text inference. Whether you are building batch ASR pipelines or real-time streaming applications, you can load and run VibeVoice using the familiar Transformers APIs without custom wrappers.
Core Integration Components
VibeVoice exposes four primary classes that plug into the Transformers ecosystem. These components are registered automatically when you import the package, allowing seamless use of AutoConfig, AutoTokenizer, AutoProcessor, and AutoModelForCausalLM.
-
VibeVoiceConfig – Defined in
vibevoice/modular/configuration_vibevoice.py, this compositePretrainedConfigaggregates acoustic, semantic, and Qwen2 decoder sub-configurations. It tells the model how to initialize the acoustic tokenizer and diffusion head. -
VibeVoiceTextTokenizerFast – Located in
vibevoice/modular/modular_vibevoice_text_tokenizer.py, this fast tokenizer inherits fromQwen2TokenizerFastand injects three speech-specific special tokens:<|vision_start|>,<|vision_end|>, and<|vision_pad|>. -
VibeVoiceProcessor and VibeVoiceASRProcessor – Implemented in
vibevoice/processor/vibevoice_processor.py, these processors wrap the acoustic tokenizer and handle audio normalization. They combine feature extraction with text tokenization to produce model-ready inputs. -
VibeVoiceForConditionalGeneration – Found in
vibevoice/modular/modeling_vibevoice.py, this is the core model class that performs both text generation and speech-to-text (ASR) when acoustic tensors are supplied.
Loading Models with Auto Factories
Because the package registers its components via vllm_plugin/__init__.py, you can load the entire stack using standard Transformers factories. This eliminates the need to import concrete classes manually unless you require specific functionality.
from transformers import AutoConfig, AutoTokenizer, AutoProcessor, AutoModelForCausalLM
model_id = "microsoft/VibeVoice-ASR"
config = AutoConfig.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
The AutoConfig call resolves to VibeVoiceConfig, while AutoTokenizer returns VibeVoiceASRTextTokenizerFast (or the base VibeVoiceTextTokenizerFast depending on the checkpoint). The processor resolves to VibeVoiceProcessor and the model instantiates VibeVoiceForConditionalGeneration with automatic mixed-precision and device placement.
Performing Speech-to-Text Inference
To run ASR, you must pass acoustic features via the speech_tensors parameter. The processor converts raw audio into the expected feature format, while the model requires a text prompt containing the speech-start token to align acoustic embeddings with text embeddings.
import torch
import torchaudio
from pathlib import Path
# Load audio
audio_path = Path("sample.wav")
waveform, sr = torchaudio.load(audio_path)
# Pre-process: extract acoustic tokens and normalize
inputs = processor(
audio=waveform.squeeze(0),
sampling_rate=sr,
return_tensors="pt",
)
# Prepare text prompt with speech-start token
input_ids = tokenizer("<|vision_start|>", return_tensors="pt")["input_ids"]
# Forward pass
outputs = model(
input_ids=input_ids,
speech_tensors=inputs["input_features"],
speech_masks=inputs["attention_mask"].bool(),
acoustic_input_mask=torch.arange(input_ids.shape[1]) == tokenizer.speech_start_id,
return_dict=True,
)
# Decode to text
generated_ids = outputs.logits.argmax(-1)
transcription = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
print(transcription)
Key parameters explained:
speech_tensors– The acoustic token sequence produced byVibeVoiceTokenizerProcessor.acoustic_input_mask– A boolean mask indicating which positions ininput_idsshould be replaced with acoustic embeddings (typically the position of<|vision_start|>).speech_masks– Attention mask for the acoustic features, ensuring the model ignores padded frames.
Real-Time Streaming Inference
For low-latency applications, VibeVoice provides a streaming processor that maintains internal caches, mirroring the KV-cache mechanism used in language models. This allows chunk-wise processing of live audio without recomputing features for the entire history.
from vibevoice.processor.vibevoice_streaming_processor import VibeVoiceStreamingProcessor
streamer = VibeVoiceStreamingProcessor.from_pretrained(model_id)
for chunk in microphone_stream():
batch = streamer(chunk, sampling_rate=sr)
out = model(
input_ids=batch["input_ids"],
speech_tensors=batch["input_features"],
speech_masks=batch["attention_mask"].bool(),
acoustic_input_mask=batch["acoustic_input_mask"],
past_key_values=streamer.past_key_values,
use_cache=True,
)
streamer.update_cache(out.past_key_values)
new_tokens = out.logits.argmax(-1)
print(tokenizer.decode(new_tokens[0], skip_special_tokens=True), end="", flush=True)
The VibeVoiceStreamingProcessor handles padding alignment and cache management automatically, returning ready-to-use tensors for each audio fragment.
Using the Transformers Pipeline
You can also deploy VibeVoice through the high-level pipeline abstraction for automatic speech recognition. Because VibeVoiceProcessor implements the standard feature-extractor interface, it works out-of-the-box with the pipeline API.
from transformers import pipeline
asr_pipe = pipeline(
"automatic-speech-recognition",
model=model,
tokenizer=tokenizer,
feature_extractor=processor,
device=0,
)
result = asr_pipe("sample.wav")
print(result["text"])
This approach is ideal for rapid prototyping or batch processing, as the pipeline handles batching and device placement internally.
Summary
- Registry-based loading: Import the vibevoice package to automatically register
VibeVoiceConfig, tokenizer, processor, and model classes with TransformersAuto*factories. - Audio input: Use
VibeVoiceProcessorto convert raw audio tospeech_tensors, then pass these alongside text prompts containing<|vision_start|>. - Inference modes: Choose between standard batch inference for file-based ASR or
VibeVoiceStreamingProcessorfor real-time, low-latency transcription. - Pipeline compatibility: Deploy immediately using
pipeline("automatic-speech-recognition")by supplying the model, tokenizer, and processor as initialized components.
Frequently Asked Questions
Can I load VibeVoice using standard Hugging Face Auto classes?
Yes. Once you install the package, AutoConfig.from_pretrained(), AutoTokenizer.from_pretrained(), and AutoModelForCausalLM.from_pretrained() will resolve to VibeVoiceConfig, VibeVoiceTextTokenizerFast, and VibeVoiceForConditionalGeneration respectively. This registration happens automatically through the plugin initialization in vllm_plugin/__init__.py.
What special tokens does VibeVoice use for speech processing?
The tokenizer reserves three speech-specific tokens: <|vision_start|> marks the insertion point for acoustic embeddings, <|vision_end|> terminates the speech segment, and <|vision_pad|> handles padding. You must include <|vision_start|> in your text prompt so the model knows where to inject the speech_tensors.
How do I handle streaming or real-time audio input?
Use VibeVoiceStreamingProcessor from vibevoice/processor/vibevoice_streaming_processor.py. This class maintains an internal cache and accepts audio chunks incrementally. Pass past_key_values=streamer.past_key_values and use_cache=True during generation to enable efficient incremental decoding without recomputing attention over the full audio history.
Is VibeVoice compatible with vLLM for accelerated serving?
Yes. The repository includes a vLLM plugin that registers the model architecture with the vLLM engine. When you import the package, register_vibevoice() is called automatically through the vllm.general_plugins entry point, making VibeVoice discoverable by the vLLM server for optimized batch inference and continuous batching.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →