# How to Use Hotwords in VibeVoice-ASR for Better Recognition

> Improve VibeVoice-ASR recognition accuracy by using hotwords. Learn how to pass domain-specific terms via context info or user prompts to bias the model for better performance.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**Pass a comma-separated string of domain-specific terms to the `context_info` parameter in `VibeVoiceASRProcessor`, or include "with extra info: {hotwords}" in the user prompt when calling the vLLM API, to bias the model toward correctly recognizing those words.**

VibeVoice-ASR is an open-source speech recognition system by Microsoft that supports **hotwords**—custom vocabulary hints that improve transcription accuracy for technical jargon, names, and brand terms. By injecting these terms as textual context before the acoustic data, the model biases its language model toward the specified tokens during generation.

## How Hotwords Work in the VibeVoice-ASR Processor

### The `context_info` Parameter Flow

According to the source code in `microsoft/VibeVoice`, hotwords enter the system through the `context_info` argument in [`vibevoice/processor/vibevoice_asr_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_asr_processor.py).

In `VibeVoiceASRProcessor.__call__` (lines 199–208), the `context_info` string is forwarded to the internal `_process_single_audio` method. Here, if the string is non-empty, the processor appends it to the user-side prompt using the phrase "with extra info: {hotwords}" (lines 361–364). This insertion occurs **before** the list of JSON keys the model must output, ensuring the hotwords act as a clear linguistic hint.

### Prompt Structure and Speech Tokens

VibeVoice-ASR uses a chat-style prompt format. The system message provides instructions, while the user message contains speech placeholders surrounded by special tokens (`<|speech_start|> … <|speech_end|>`) and the hotword context. By placing the hotwords as explicit text before the acoustic representation, the transformer can adjust its token probabilities to favor the supplied terms during transcription generation.

## Three Ways to Use Hotwords in VibeVoice-ASR

### 1. Python API: `VibeVoiceASRProcessor`

For programmatic access, instantiate the processor and pass hotwords via `context_info`:

```python
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("microsoft/VibeVoice-ASR")
processor = VibeVoiceASRProcessor(tokenizer=tokenizer)

hotwords = "VibeVoice,Microsoft,Azure"

encoding = processor(
    audio="demo/asr_demo/demo3-hotwords.wav",
    return_tensors="pt",
    context_info=hotwords,
)

```

The processor automatically transforms the input into the prompt fragment: `This is a 12.34 seconds audio, with extra info: VibeVoice,Microsoft,Azure`.

### 2. Gradio Web Interface

Launch the official demo to use hotwords interactively:

```bash
pip install -e .
python demo/vibevoice_asr_gradio_demo.py --model_path microsoft/VibeVoice-ASR --share

```

In the interface:

- Upload your audio file.
- Enter comma-separated terms in the **Hotwords / context** textbox (e.g., `OpenAI,TensorFlow,CUDA`).
- Click **Transcribe**.

The demo passes the textbox value as `context_info` to the processor (see [`demo/vibevoice_asr_gradio_demo.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_asr_gradio_demo.py), lines 19–32).

### 3. vLLM HTTP API

For production deployments using the vLLM plugin, embed hotwords directly in the user message:

```bash
AUDIO_B64=$(base64 -w 0 demo/asr_demo/demo3-hotwords.wav)

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "vibevoice",
    "messages": [
      {"role":"system","content":"You are a helpful assistant that transcribes audio input into text output in JSON format."},
      {"role":"user","content":[
        {"type":"audio_url","audio_url":{"url":"data:audio/wav;base64,'"${AUDIO_B64}"'"}},
        {"type":"text","text":"This is a 12.34 seconds audio, with extra info: VibeVoice,Microsoft,Azure\n\nPlease transcribe it with these keys: Start time, End time, Speaker ID, Content"}
      ]}
    ],
    "max_tokens": 32768,
    "temperature": 0
  }'

```

Reference [`vllm_plugin/tests/test_api.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api.py) (lines 42–55) for the programmatic implementation of `test_transcription_with_hotwords`.

## Summary

- Hotwords in VibeVoice-ASR are comma-separated strings passed via the `context_info` parameter.
- The processor injects them into the user prompt as "with extra info: {hotwords}" before the speech tokens and output keys.
- You can supply hotwords through the Python `VibeVoiceASRProcessor` API, the Gradio demo interface, or direct HTTP calls to the vLLM endpoint.
- This technique biases the language model toward specific tokens, improving recognition accuracy for domain-specific terminology.

## Frequently Asked Questions

### What is the correct format for hotwords in VibeVoice-ASR?

Enter hotwords as a plain comma-separated string without spaces after commas (e.g., `VibeVoice,Microsoft,Azure`). The processor in [`vibevoice_asr_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice_asr_processor.py) inserts this string verbatim after the phrase "with extra info:".

### Can I use hotwords when deploying with vLLM?

Yes. The vLLM plugin fully supports hotwords. Include the "with extra info: {hotwords}" text inside the user message's text content block, as demonstrated in [`vllm_plugin/tests/test_api.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tests/test_api.py) and the curl example above.

### Where exactly do hotwords appear in the model's prompt?

Hotwords appear in the user-side prompt text, positioned after the audio duration statement and before the instruction listing the required JSON output keys. This placement ensures the model processes the vocabulary hints before generating the transcription.

### Does the hotwords feature support multi-word phrases or just single words?

The `context_info` parameter accepts any text string, so you can include multi-word phrases (e.g., `Visual Studio Code,Azure DevOps`). The model treats the entire injected string as contextual bias, regardless of word boundaries.