How VibeVoice-ASR Handles Speaker Diarization and Timestamps: A Prompt-Based Approach
TL;DR: VibeVoice-ASR performs speaker diarization and timestamping without a separate clustering module by prompting the multimodal model to emit a JSON list containing Start time, End time, Speaker ID, and Content, which the VibeVoiceASRProcessor then parses into normalized Python dictionaries.
The microsoft/VibeVoice repository implements an end-to-end automatic speech recognition (ASR) system that unifies transcription, timestamping, and speaker diarization into a single generative process. Unlike traditional pipelines that rely on external diarization engines or forced alignment, VibeVoice-ASR treats speaker identification and temporal boundaries as structured outputs predicted directly by the language model. This article examines the prompt engineering strategy, post-processing logic, and source code implementation that enable VibeVoice-ASR to deliver timestamped, speaker-attributed transcripts from raw audio.
Architectural Overview
The diarization mechanism in VibeVoice-ASR consists of three coordinated components: the processor that constructs specialized prompts, the generative model that outputs structured JSON, and the post-processor that normalizes the results. The system is implemented across vibevoice/processor/vibevoice_asr_processor.py and demo/vibevoice_asr_gradio_demo.py.
Prompt Engineering for Structured Output
In vibevoice/processor/vibevoice_asr_processor.py, the _process_single_audio method (lines 60-66) dynamically constructs a user prompt that explicitly instructs the model to return specific keys. The processor appends a suffix to the speech placeholder tokens that defines the required output format:
show_keys = ['Start time', 'End time', 'Speaker ID', 'Content']
user_suffix = (
f"This is a {audio_duration:.2f} seconds audio, please transcribe it with these keys: "
+ ", ".join(show_keys)
)
user_input_string = speech_placeholder + "\n" + user_suffix
The system prompt (SYSTEM_PROMPT) further instructs the model to produce valid JSON. During inference, the model generates an autoregressive text completion that includes a JSON array where each element contains the four specified fields. Because the model learns this structure from its training data, it predicts speaker turns and temporal boundaries directly without requiring a separate clustering algorithm.
Post-Processing and Key Normalization
After generation, the post_process_transcription method (lines 90-115) extracts and validates the structured output. This method handles three common output variations: plain JSON, markdown code blocks (```json), and truncated snippets. It then normalizes potentially varying key spellings into canonical field names using a mapping dictionary:
key_mapping = {
"Start time": "start_time",
"Start": "start_time",
"End time": "end_time",
"End": "end_time",
"Speaker ID": "speaker_id",
"Speaker": "speaker_id",
"Content": "text",
}
The method filters out any items missing the expected fields and returns a clean list of dictionaries with the standardized keys start_time, end_time, speaker_id, and text. This normalization ensures downstream consumers receive a consistent data structure regardless of minor variations in the model's raw output.
End-to-End Transcription Flow
The complete inference pipeline illustrates how VibeVoice-ASR integrates diarization into the standard transcription workflow:
- Input Preparation: The
VibeVoiceASRProcessortokenizes the audio and constructs the prompt with the four required keys. - Model Generation: The model receives speech tokens and the textual prompt, then generates a JSON string containing segments with timestamps and speaker identifiers.
- Segment Extraction: The processor's
post_process_transcriptionparses the JSON, normalizes keys, and filters invalid entries. - UI Rendering: The
VibeVoiceASRInference.transcribemethod (lines 92-106 indemo/vibevoice_asr_gradio_demo.py) returns the segments, which the Gradio interface renders as timestamped speaker labels with per-segment audio playback.
This design eliminates the need for complex multi-pass algorithms or external speaker embedding models, reducing latency and simplifying deployment.
Code Implementation Examples
Processing Audio with VibeVoiceASRProcessor
To prepare audio for diarization-aware transcription without running the model (useful for debugging prompts), initialize the processor and encode the audio:
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
from transformers import AutoTokenizer
import numpy as np
# Load the tokenizer used for VibeVoice-ASR
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B")
# Build the processor
processor = VibeVoiceASRProcessor.from_pretrained(
pretrained_model_name_or_path="path/to/vibevoice_asr_checkpoint",
tokenizer=tokenizer,
)
# Prepare raw audio (numpy array) - synthesize a 2-second tone for demonstration
sr = 24000
duration_sec = 2.0
audio = np.sin(2 * np.pi * 440 * np.arange(sr * duration_sec) / sr).astype(np.float32)
# Encode the audio (builds the prompt with diarization instructions)
encoding = processor(
audio=audio,
sampling_rate=sr,
return_tensors="pt",
add_generation_prompt=True,
context_info="John, Mary are speaking",
)
print("Input token IDs length:", encoding["input_ids"].shape[-1])
print("Acoustic mask sum (speech tokens):", encoding["acoustic_input_mask"].sum())
The returned BatchEncoding contains the speech placeholder tokens combined with the user suffix that demands the four output keys, ready for model inference.
Running Full Inference with Diarization
For complete transcription including speaker attribution, use the VibeVoiceASRInference wrapper:
from demo.vibevoice_asr_gradio_demo import VibeVoiceASRInference
# Initialize the inference wrapper (loads model and processor)
asr = VibeVoiceASRInference(
model_path="path/to/vibevoice_asr_checkpoint",
device="cpu", # or "cuda"
)
# Transcribe an audio file (WAV, MP3, OGG supported)
result = asr.transcribe(
audio_path="sample.wav",
max_new_tokens=1024,
temperature=0.0, # deterministic output
top_p=1.0,
do_sample=False,
)
# Access the diarized segments
for seg in result["segments"]:
print(
f"[Speaker {seg['speaker_id']}] "
f"{seg['start_time']:.2f}s-{seg['end_time']:.2f}s: {seg['text']}"
)
Sample Output:
[Speaker 1] 0.00s-3.45s: Hello everyone
[Speaker 2] 3.46s-7.12s: Great to be here
The underlying model generated these speaker IDs and timestamps as part of its text completion, which the processor extracted and normalized into the segments list.
Visualizing Results in the Gradio Demo
To launch the interactive interface that renders speaker diarization visually:
python demo/vibevoice_asr_gradio_demo.py \
--model_path path/to/vibevoice_asr_checkpoint \
--device cuda
The demo interface (implemented in demo/vibevoice_asr_gradio_demo.py, lines 393-410) displays the Audio Segments tab only when valid segments are present. Each entry renders as Segment 1: [0.00s - 3.45s] Speaker 1 with an embedded HTML5 audio player clipped to the exact time boundaries, providing immediate visual and auditory verification of the diarization accuracy.
Summary
- Unified Generation: VibeVoice-ASR treats speaker diarization and timestamps as part of the language modeling task, eliminating the need for separate clustering or alignment modules.
- Prompt-Driven Structure: The processor instructs the model to output specific JSON keys (
Start time,End time,Speaker ID,Content) through engineered prompts in_process_single_audio. - Robust Parsing: The
post_process_transcriptionmethod handles markdown code blocks, truncated JSON, and key variants, returning normalized dictionaries with canonical field names. - End-to-End Integration: The
VibeVoiceASRInference.transcribemethod exposes these segments directly to applications, enabling the Gradio demo to render interactive, per-speaker audio clips.
Frequently Asked Questions
Does VibeVoice-ASR require a separate speaker diarization model?
No. According to the source code in vibevoice/processor/vibevoice_asr_processor.py, the system performs diarization through prompt engineering and the base model's generative capabilities. There is no external diarization module; the model predicts speaker identifiers and timestamps directly during the autoregressive text generation phase.
What audio formats does the Gradio demo support?
The VibeVoiceASRInference.transcribe method supports standard formats including WAV, MP3, and OGG, as indicated by the audio loading utilities in demo/vibevoice_asr_gradio_demo.py. The demo processes these formats uniformly before tokenizing the audio for the model.
How does the model handle truncated or malformed JSON output?
The post_process_transcription method (lines 90-115) implements defensive parsing logic that detects and extracts JSON from markdown code blocks (```json ... ```), plain text, or truncated snippets. It then validates the presence of required fields and maps various key spellings to canonical names before returning the final segment list.
Can I customize the speaker labels or timestamp format?
While the prompt currently requests Start time, End time, Speaker ID, and Content, you can modify the show_keys list and key_mapping dictionary in vibevoice_asr_processor.py to accommodate different field names. However, the model's effectiveness depends on its training data; significant deviations from the expected format may require fine-tuning to maintain output quality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →