How to Integrate VibeVoice with Hugging Face Transformers: Complete Setup and ASR Guide

You can integrate VibeVoice with Hugging Face Transformers by using the standard Auto* factory methods—such as AutoModelForCausalLM.from_pretrained()—which automatically resolve to VibeVoice-specific classes thanks to registry entries in vllm_plugin/__init__.py.

The Microsoft VibeVoice repository ships as a fully Hugging Face Transformers-compatible package, providing composite configs, custom tokenizers, and processor classes that enable seamless speech-to-text inference. Whether you are building batch ASR pipelines or real-time streaming applications, you can load and run VibeVoice using the familiar Transformers APIs without custom wrappers.

Core Integration Components

VibeVoice exposes four primary classes that plug into the Transformers ecosystem. These components are registered automatically when you import the package, allowing seamless use of AutoConfig, AutoTokenizer, AutoProcessor, and AutoModelForCausalLM.

  • VibeVoiceConfig – Defined in vibevoice/modular/configuration_vibevoice.py, this composite PretrainedConfig aggregates acoustic, semantic, and Qwen2 decoder sub-configurations. It tells the model how to initialize the acoustic tokenizer and diffusion head.

  • VibeVoiceTextTokenizerFast – Located in vibevoice/modular/modular_vibevoice_text_tokenizer.py, this fast tokenizer inherits from Qwen2TokenizerFast and injects three speech-specific special tokens: <|vision_start|>, <|vision_end|>, and <|vision_pad|>.

  • VibeVoiceProcessor and VibeVoiceASRProcessor – Implemented in vibevoice/processor/vibevoice_processor.py, these processors wrap the acoustic tokenizer and handle audio normalization. They combine feature extraction with text tokenization to produce model-ready inputs.

  • VibeVoiceForConditionalGeneration – Found in vibevoice/modular/modeling_vibevoice.py, this is the core model class that performs both text generation and speech-to-text (ASR) when acoustic tensors are supplied.

Loading Models with Auto Factories

Because the package registers its components via vllm_plugin/__init__.py, you can load the entire stack using standard Transformers factories. This eliminates the need to import concrete classes manually unless you require specific functionality.

from transformers import AutoConfig, AutoTokenizer, AutoProcessor, AutoModelForCausalLM

model_id = "microsoft/VibeVoice-ASR"

config = AutoConfig.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

The AutoConfig call resolves to VibeVoiceConfig, while AutoTokenizer returns VibeVoiceASRTextTokenizerFast (or the base VibeVoiceTextTokenizerFast depending on the checkpoint). The processor resolves to VibeVoiceProcessor and the model instantiates VibeVoiceForConditionalGeneration with automatic mixed-precision and device placement.

Performing Speech-to-Text Inference

To run ASR, you must pass acoustic features via the speech_tensors parameter. The processor converts raw audio into the expected feature format, while the model requires a text prompt containing the speech-start token to align acoustic embeddings with text embeddings.

import torch
import torchaudio
from pathlib import Path

# Load audio

audio_path = Path("sample.wav")
waveform, sr = torchaudio.load(audio_path)

# Pre-process: extract acoustic tokens and normalize

inputs = processor(
    audio=waveform.squeeze(0),
    sampling_rate=sr,
    return_tensors="pt",
)

# Prepare text prompt with speech-start token

input_ids = tokenizer("<|vision_start|>", return_tensors="pt")["input_ids"]

# Forward pass

outputs = model(
    input_ids=input_ids,
    speech_tensors=inputs["input_features"],
    speech_masks=inputs["attention_mask"].bool(),
    acoustic_input_mask=torch.arange(input_ids.shape[1]) == tokenizer.speech_start_id,
    return_dict=True,
)

# Decode to text

generated_ids = outputs.logits.argmax(-1)
transcription = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
print(transcription)

Key parameters explained:

  • speech_tensors – The acoustic token sequence produced by VibeVoiceTokenizerProcessor.
  • acoustic_input_mask – A boolean mask indicating which positions in input_ids should be replaced with acoustic embeddings (typically the position of <|vision_start|>).
  • speech_masks – Attention mask for the acoustic features, ensuring the model ignores padded frames.

Real-Time Streaming Inference

For low-latency applications, VibeVoice provides a streaming processor that maintains internal caches, mirroring the KV-cache mechanism used in language models. This allows chunk-wise processing of live audio without recomputing features for the entire history.

from vibevoice.processor.vibevoice_streaming_processor import VibeVoiceStreamingProcessor

streamer = VibeVoiceStreamingProcessor.from_pretrained(model_id)

for chunk in microphone_stream():
    batch = streamer(chunk, sampling_rate=sr)
    
    out = model(
        input_ids=batch["input_ids"],
        speech_tensors=batch["input_features"],
        speech_masks=batch["attention_mask"].bool(),
        acoustic_input_mask=batch["acoustic_input_mask"],
        past_key_values=streamer.past_key_values,
        use_cache=True,
    )
    
    streamer.update_cache(out.past_key_values)
    new_tokens = out.logits.argmax(-1)
    print(tokenizer.decode(new_tokens[0], skip_special_tokens=True), end="", flush=True)

The VibeVoiceStreamingProcessor handles padding alignment and cache management automatically, returning ready-to-use tensors for each audio fragment.

Using the Transformers Pipeline

You can also deploy VibeVoice through the high-level pipeline abstraction for automatic speech recognition. Because VibeVoiceProcessor implements the standard feature-extractor interface, it works out-of-the-box with the pipeline API.

from transformers import pipeline

asr_pipe = pipeline(
    "automatic-speech-recognition",
    model=model,
    tokenizer=tokenizer,
    feature_extractor=processor,
    device=0,
)

result = asr_pipe("sample.wav")
print(result["text"])

This approach is ideal for rapid prototyping or batch processing, as the pipeline handles batching and device placement internally.

Summary

  • Registry-based loading: Import the vibevoice package to automatically register VibeVoiceConfig, tokenizer, processor, and model classes with Transformers Auto* factories.
  • Audio input: Use VibeVoiceProcessor to convert raw audio to speech_tensors, then pass these alongside text prompts containing <|vision_start|>.
  • Inference modes: Choose between standard batch inference for file-based ASR or VibeVoiceStreamingProcessor for real-time, low-latency transcription.
  • Pipeline compatibility: Deploy immediately using pipeline("automatic-speech-recognition") by supplying the model, tokenizer, and processor as initialized components.

Frequently Asked Questions

Can I load VibeVoice using standard Hugging Face Auto classes?

Yes. Once you install the package, AutoConfig.from_pretrained(), AutoTokenizer.from_pretrained(), and AutoModelForCausalLM.from_pretrained() will resolve to VibeVoiceConfig, VibeVoiceTextTokenizerFast, and VibeVoiceForConditionalGeneration respectively. This registration happens automatically through the plugin initialization in vllm_plugin/__init__.py.

What special tokens does VibeVoice use for speech processing?

The tokenizer reserves three speech-specific tokens: <|vision_start|> marks the insertion point for acoustic embeddings, <|vision_end|> terminates the speech segment, and <|vision_pad|> handles padding. You must include <|vision_start|> in your text prompt so the model knows where to inject the speech_tensors.

How do I handle streaming or real-time audio input?

Use VibeVoiceStreamingProcessor from vibevoice/processor/vibevoice_streaming_processor.py. This class maintains an internal cache and accepts audio chunks incrementally. Pass past_key_values=streamer.past_key_values and use_cache=True during generation to enable efficient incremental decoding without recomputing attention over the full audio history.

Is VibeVoice compatible with vLLM for accelerated serving?

Yes. The repository includes a vLLM plugin that registers the model architecture with the vLLM engine. When you import the package, register_vibevoice() is called automatically through the vllm.general_plugins entry point, making VibeVoice discoverable by the vLLM server for optimized batch inference and continuous batching.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →