How to Use Hotwords in VibeVoice-ASR for Better Recognition
Pass a comma-separated string of domain-specific terms to the context_info parameter in VibeVoiceASRProcessor, or include "with extra info: {hotwords}" in the user prompt when calling the vLLM API, to bias the model toward correctly recognizing those words.
VibeVoice-ASR is an open-source speech recognition system by Microsoft that supports hotwords—custom vocabulary hints that improve transcription accuracy for technical jargon, names, and brand terms. By injecting these terms as textual context before the acoustic data, the model biases its language model toward the specified tokens during generation.
How Hotwords Work in the VibeVoice-ASR Processor
The context_info Parameter Flow
According to the source code in microsoft/VibeVoice, hotwords enter the system through the context_info argument in vibevoice/processor/vibevoice_asr_processor.py.
In VibeVoiceASRProcessor.__call__ (lines 199–208), the context_info string is forwarded to the internal _process_single_audio method. Here, if the string is non-empty, the processor appends it to the user-side prompt using the phrase "with extra info: {hotwords}" (lines 361–364). This insertion occurs before the list of JSON keys the model must output, ensuring the hotwords act as a clear linguistic hint.
Prompt Structure and Speech Tokens
VibeVoice-ASR uses a chat-style prompt format. The system message provides instructions, while the user message contains speech placeholders surrounded by special tokens (<|speech_start|> … <|speech_end|>) and the hotword context. By placing the hotwords as explicit text before the acoustic representation, the transformer can adjust its token probabilities to favor the supplied terms during transcription generation.
Three Ways to Use Hotwords in VibeVoice-ASR
1. Python API: VibeVoiceASRProcessor
For programmatic access, instantiate the processor and pass hotwords via context_info:
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("microsoft/VibeVoice-ASR")
processor = VibeVoiceASRProcessor(tokenizer=tokenizer)
hotwords = "VibeVoice,Microsoft,Azure"
encoding = processor(
audio="demo/asr_demo/demo3-hotwords.wav",
return_tensors="pt",
context_info=hotwords,
)
The processor automatically transforms the input into the prompt fragment: This is a 12.34 seconds audio, with extra info: VibeVoice,Microsoft,Azure.
2. Gradio Web Interface
Launch the official demo to use hotwords interactively:
pip install -e .
python demo/vibevoice_asr_gradio_demo.py --model_path microsoft/VibeVoice-ASR --share
In the interface:
- Upload your audio file.
- Enter comma-separated terms in the Hotwords / context textbox (e.g.,
OpenAI,TensorFlow,CUDA). - Click Transcribe.
The demo passes the textbox value as context_info to the processor (see demo/vibevoice_asr_gradio_demo.py, lines 19–32).
3. vLLM HTTP API
For production deployments using the vLLM plugin, embed hotwords directly in the user message:
AUDIO_B64=$(base64 -w 0 demo/asr_demo/demo3-hotwords.wav)
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "vibevoice",
"messages": [
{"role":"system","content":"You are a helpful assistant that transcribes audio input into text output in JSON format."},
{"role":"user","content":[
{"type":"audio_url","audio_url":{"url":"data:audio/wav;base64,'"${AUDIO_B64}"'"}},
{"type":"text","text":"This is a 12.34 seconds audio, with extra info: VibeVoice,Microsoft,Azure\n\nPlease transcribe it with these keys: Start time, End time, Speaker ID, Content"}
]}
],
"max_tokens": 32768,
"temperature": 0
}'
Reference vllm_plugin/tests/test_api.py (lines 42–55) for the programmatic implementation of test_transcription_with_hotwords.
Summary
- Hotwords in VibeVoice-ASR are comma-separated strings passed via the
context_infoparameter. - The processor injects them into the user prompt as "with extra info: {hotwords}" before the speech tokens and output keys.
- You can supply hotwords through the Python
VibeVoiceASRProcessorAPI, the Gradio demo interface, or direct HTTP calls to the vLLM endpoint. - This technique biases the language model toward specific tokens, improving recognition accuracy for domain-specific terminology.
Frequently Asked Questions
What is the correct format for hotwords in VibeVoice-ASR?
Enter hotwords as a plain comma-separated string without spaces after commas (e.g., VibeVoice,Microsoft,Azure). The processor in vibevoice_asr_processor.py inserts this string verbatim after the phrase "with extra info:".
Can I use hotwords when deploying with vLLM?
Yes. The vLLM plugin fully supports hotwords. Include the "with extra info: {hotwords}" text inside the user message's text content block, as demonstrated in vllm_plugin/tests/test_api.py and the curl example above.
Where exactly do hotwords appear in the model's prompt?
Hotwords appear in the user-side prompt text, positioned after the audio duration statement and before the instruction listing the required JSON output keys. This placement ensures the model processes the vocabulary hints before generating the transcription.
Does the hotwords feature support multi-word phrases or just single words?
The context_info parameter accepts any text string, so you can include multi-word phrases (e.g., Visual Studio Code,Azure DevOps). The model treats the entire injected string as contextual bias, regardless of word boundaries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →