# How VibeVoice-ASR Handles Speaker Diarization and Timestamps: A Prompt-Based Approach

> Discover how VibeVoice-ASR achieves speaker diarization and timestamps using a prompt-based approach. Learn how the multimodal model generates structured output for efficient audio processing.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: deep-dive
- Published: 2026-03-28

---

**TL;DR:** VibeVoice-ASR performs speaker diarization and timestamping without a separate clustering module by prompting the multimodal model to emit a JSON list containing `Start time`, `End time`, `Speaker ID`, and `Content`, which the `VibeVoiceASRProcessor` then parses into normalized Python dictionaries.

The **microsoft/VibeVoice** repository implements an end-to-end automatic speech recognition (ASR) system that unifies transcription, timestamping, and speaker diarization into a single generative process. Unlike traditional pipelines that rely on external diarization engines or forced alignment, VibeVoice-ASR treats speaker identification and temporal boundaries as structured outputs predicted directly by the language model. This article examines the prompt engineering strategy, post-processing logic, and source code implementation that enable VibeVoice-ASR to deliver timestamped, speaker-attributed transcripts from raw audio.

## Architectural Overview

The diarization mechanism in VibeVoice-ASR consists of three coordinated components: the processor that constructs specialized prompts, the generative model that outputs structured JSON, and the post-processor that normalizes the results. The system is implemented across [`vibevoice/processor/vibevoice_asr_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_asr_processor.py) and [`demo/vibevoice_asr_gradio_demo.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_asr_gradio_demo.py).

### Prompt Engineering for Structured Output

In [`vibevoice/processor/vibevoice_asr_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_asr_processor.py), the `_process_single_audio` method (lines 60-66) dynamically constructs a user prompt that explicitly instructs the model to return specific keys. The processor appends a suffix to the speech placeholder tokens that defines the required output format:

```python
show_keys = ['Start time', 'End time', 'Speaker ID', 'Content']
user_suffix = (
    f"This is a {audio_duration:.2f} seconds audio, please transcribe it with these keys: "
    + ", ".join(show_keys)
)
user_input_string = speech_placeholder + "\n" + user_suffix

```

The **system prompt** (`SYSTEM_PROMPT`) further instructs the model to produce valid JSON. During inference, the model generates an autoregressive text completion that includes a JSON array where each element contains the four specified fields. Because the model learns this structure from its training data, it predicts speaker turns and temporal boundaries directly without requiring a separate clustering algorithm.

### Post-Processing and Key Normalization

After generation, the `post_process_transcription` method (lines 90-115) extracts and validates the structured output. This method handles three common output variations: plain JSON, markdown code blocks (`` ```json ``), and truncated snippets. It then normalizes potentially varying key spellings into canonical field names using a mapping dictionary:

```python
key_mapping = {
    "Start time": "start_time",
    "Start": "start_time",
    "End time": "end_time",
    "End": "end_time",
    "Speaker ID": "speaker_id",
    "Speaker": "speaker_id",
    "Content": "text",
}

```

The method filters out any items missing the expected fields and returns a clean list of dictionaries with the standardized keys `start_time`, `end_time`, `speaker_id`, and `text`. This normalization ensures downstream consumers receive a consistent data structure regardless of minor variations in the model's raw output.

## End-to-End Transcription Flow

The complete inference pipeline illustrates how VibeVoice-ASR integrates diarization into the standard transcription workflow:

1. **Input Preparation**: The `VibeVoiceASRProcessor` tokenizes the audio and constructs the prompt with the four required keys.
2. **Model Generation**: The model receives speech tokens and the textual prompt, then generates a JSON string containing segments with timestamps and speaker identifiers.
3. **Segment Extraction**: The processor's `post_process_transcription` parses the JSON, normalizes keys, and filters invalid entries.
4. **UI Rendering**: The `VibeVoiceASRInference.transcribe` method (lines 92-106 in `demo/vibevoice_asr_gradio_demo.py`) returns the segments, which the Gradio interface renders as timestamped speaker labels with per-segment audio playback.

This design eliminates the need for complex multi-pass algorithms or external speaker embedding models, reducing latency and simplifying deployment.

## Code Implementation Examples

### Processing Audio with VibeVoiceASRProcessor

To prepare audio for diarization-aware transcription without running the model (useful for debugging prompts), initialize the processor and encode the audio:

```python
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
from transformers import AutoTokenizer
import numpy as np

# Load the tokenizer used for VibeVoice-ASR

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B")

# Build the processor

processor = VibeVoiceASRProcessor.from_pretrained(
    pretrained_model_name_or_path="path/to/vibevoice_asr_checkpoint",
    tokenizer=tokenizer,
)

# Prepare raw audio (numpy array) - synthesize a 2-second tone for demonstration

sr = 24000
duration_sec = 2.0
audio = np.sin(2 * np.pi * 440 * np.arange(sr * duration_sec) / sr).astype(np.float32)

# Encode the audio (builds the prompt with diarization instructions)

encoding = processor(
    audio=audio,
    sampling_rate=sr,
    return_tensors="pt",
    add_generation_prompt=True,
    context_info="John, Mary are speaking",
)

print("Input token IDs length:", encoding["input_ids"].shape[-1])
print("Acoustic mask sum (speech tokens):", encoding["acoustic_input_mask"].sum())

```

The returned `BatchEncoding` contains the speech placeholder tokens combined with the user suffix that demands the four output keys, ready for model inference.

### Running Full Inference with Diarization

For complete transcription including speaker attribution, use the `VibeVoiceASRInference` wrapper:

```python
from demo.vibevoice_asr_gradio_demo import VibeVoiceASRInference

# Initialize the inference wrapper (loads model and processor)

asr = VibeVoiceASRInference(
    model_path="path/to/vibevoice_asr_checkpoint",
    device="cpu",  # or "cuda"

)

# Transcribe an audio file (WAV, MP3, OGG supported)

result = asr.transcribe(
    audio_path="sample.wav",
    max_new_tokens=1024,
    temperature=0.0,  # deterministic output

    top_p=1.0,
    do_sample=False,
)

# Access the diarized segments

for seg in result["segments"]:
    print(
        f"[Speaker {seg['speaker_id']}] "
        f"{seg['start_time']:.2f}s-{seg['end_time']:.2f}s: {seg['text']}"
    )

```

**Sample Output:**

```

[Speaker 1] 0.00s-3.45s: Hello everyone
[Speaker 2] 3.46s-7.12s: Great to be here

```

The underlying model generated these speaker IDs and timestamps as part of its text completion, which the processor extracted and normalized into the `segments` list.

### Visualizing Results in the Gradio Demo

To launch the interactive interface that renders speaker diarization visually:

```bash
python demo/vibevoice_asr_gradio_demo.py \
    --model_path path/to/vibevoice_asr_checkpoint \
    --device cuda

```

The demo interface (implemented in `demo/vibevoice_asr_gradio_demo.py`, lines 393-410) displays the **Audio Segments** tab only when valid segments are present. Each entry renders as `Segment 1: [0.00s - 3.45s] Speaker 1` with an embedded HTML5 audio player clipped to the exact time boundaries, providing immediate visual and auditory verification of the diarization accuracy.

## Summary

- **Unified Generation**: VibeVoice-ASR treats speaker diarization and timestamps as part of the language modeling task, eliminating the need for separate clustering or alignment modules.
- **Prompt-Driven Structure**: The processor instructs the model to output specific JSON keys (`Start time`, `End time`, `Speaker ID`, `Content`) through engineered prompts in `_process_single_audio`.
- **Robust Parsing**: The `post_process_transcription` method handles markdown code blocks, truncated JSON, and key variants, returning normalized dictionaries with canonical field names.
- **End-to-End Integration**: The `VibeVoiceASRInference.transcribe` method exposes these segments directly to applications, enabling the Gradio demo to render interactive, per-speaker audio clips.

## Frequently Asked Questions

### Does VibeVoice-ASR require a separate speaker diarization model?

No. According to the source code in `vibevoice/processor/vibevoice_asr_processor.py`, the system performs diarization through prompt engineering and the base model's generative capabilities. There is no external diarization module; the model predicts speaker identifiers and timestamps directly during the autoregressive text generation phase.

### What audio formats does the Gradio demo support?

The `VibeVoiceASRInference.transcribe` method supports standard formats including WAV, MP3, and OGG, as indicated by the audio loading utilities in `demo/vibevoice_asr_gradio_demo.py`. The demo processes these formats uniformly before tokenizing the audio for the model.

### How does the model handle truncated or malformed JSON output?

The `post_process_transcription` method (lines 90-115) implements defensive parsing logic that detects and extracts JSON from markdown code blocks (`` ```json ... ``` ``), plain text, or truncated snippets. It then validates the presence of required fields and maps various key spellings to canonical names before returning the final segment list.

### Can I customize the speaker labels or timestamp format?

While the prompt currently requests `Start time`, `End time`, `Speaker ID`, and `Content`, you can modify the `show_keys` list and `key_mapping` dictionary in [`vibevoice_asr_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice_asr_processor.py) to accommodate different field names. However, the model's effectiveness depends on its training data; significant deviations from the expected format may require fine-tuning to maintain output quality.