# VibeVoice Performance Benchmarks: ASR, Real-Time TTS, and Long-Form Metrics

> Discover VibeVoice performance benchmarks: sub-5% diarization error rates for multilingual ASR, 2.00% WER for real-time TTS, and SOTA quality for 90-min multi-speaker generation. See the data.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: performance
- Published: 2026-03-28

---

**Microsoft VibeVoice achieves sub-5% diarization error rates on multilingual ASR, 2.00% WER on real-time zero-shot TTS, and state-of-the-art quality on 90-minute multi-speaker generation according to the repository evaluation tables.**

The microsoft/VibeVoice repository provides comprehensive performance benchmarks for three distinct model families: automatic speech recognition (ASR), real-time streaming TTS, and long-form multi-speaker synthesis. These metrics are documented in the repository's markdown files and demonstrate competitive results against modern large-scale speech models on datasets like MLC-Challenge, LibriSpeech, and AISHELL-4.

## VibeVoice-ASR Performance Metrics

The ASR model, defined in [`vibevoice/modular/modeling_vibevoice_asr.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_asr.py), supports transcription of up to 60 minutes of audio in a single pass while maintaining diarization error rates (DER) below 5% for most languages. Evaluation tables in [`docs/vibevoice-asr.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-asr.md) report four key metrics: **DER** (diarization error rate), **cpWER** (concatenated minimum-permutation word error rate), **tcpWER** (time-constrained cpWER), and standard **WER**.

### Multilingual Results on MLC-Challenge

The MLC-Challenge benchmark spans multiple languages with the following average performance:

| Dataset | Language | DER (↓) | cpWER (↓) | tcpWER (↓) | WER (↓) |
|---------|----------|---------|-----------|------------|---------|
| MLC-Challenge | English | 4.28 | 11.48 | 13.02 | 7.99 |
| MLC-Challenge | French | 3.80 | 18.80 | 19.64 | 15.21 |
| MLC-Challenge | German | 1.04 | 17.10 | 17.26 | 16.30 |
| **Average (MLC-Challenge)** | – | **3.42** | **14.81** | **15.66** | **12.07** |

These results position VibeVoice-ASR competitively with modern large-scale ASR systems, particularly excelling in German with a 1.04% DER.

### Meeting and Conversation Benchmarks

On challenging multi-speaker meeting datasets, the model shows robust performance:

| Dataset | Language | DER (↓) | cpWER (↓) | tcpWER (↓) | WER (↓) |
|---------|----------|---------|-----------|------------|---------|
| AISHELL-4 | Chinese | 6.77 | 24.99 | 25.35 | 21.40 |
| AMI-IHM | English | 11.92 | 20.41 | 20.82 | 18.81 |
| AMI-SDM | English | 13.43 | 28.82 | 29.80 | 24.65 |
| AliMeeting | Chinese | 10.92 | 29.33 | 29.51 | 27.40 |

The higher DER values on AMI and AliMeeting reflect the complexity of far-field microphone arrays and overlapping speech in natural meeting environments.

## Real-Time TTS Performance (0.5B Model)

VibeVoice-Realtime-0.5B targets zero-shot streaming text-to-speech with low latency. According to [`docs/vibevoice-realtime-0.5b.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-realtime-0.5b.md), the model achieves **2.00% WER** on LibriSpeech test-clean with a **0.695 speaker similarity** score, outperforming comparable systems like VALL-E 2 (2.40% WER) while maintaining competitive quality with VoiceBox (1.90% WER).

| Benchmark | Model | WER (↓) | Speaker Similarity (↑) |
|-----------|-------|---------|------------------------|
| LibriSpeech test-clean | VibeVoice-Realtime-0.5B | 2.00% | 0.695 |
| SEED test-en | VibeVoice-Realtime-0.5B | 2.05% | 0.633 |

### Latency Benchmarks

The real-time implementation delivers the first audible audio chunk in approximately **200 ms** when running on a T4 GPU. This measurement comes from the streaming inference server defined in [`demo/vibevoice_realtime_demo.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_realtime_demo.py), which uses websocket-based chunk delivery for interactive applications.

## Long-Form Multi-Speaker TTS Results

For extended content generation, VibeVoice-TTS handles 90-minute multi-speaker dialogues while preserving speaker identity consistency. The documentation in [`docs/vibevoice-tts.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-tts.md) states that this model "achieves state-of-the-art performance on long-form multi-speaker speech generation tasks." Unlike the ASR and real-time TTS components, the repository does not expose numeric tables for this model family; exact figures are available in the linked technical report.

## Reproducing the Benchmarks

You can validate these performance benchmarks for VibeVoice using the provided inference scripts and demo files.

### ASR Inference and Metric Calculation

Run the reference script [`demo/vibevoice_asr_inference_from_file.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_asr_inference_from_file.py) to generate transcripts for WER and DER calculation:

```python

# demo/vibevoice_asr_inference_from_file.py

# Example usage (replace with your own audio path)

# python demo/vibevoice_asr_inference_from_file.py \

#   --model_path microsoft/VibeVoice-ASR \

#   --audio_files demo/asr_demo/demo1-chat.mp4

```

The script outputs JSON transcripts compatible with `jiwer.wer` and `pyannote.metrics.diarization` for official metric computation.

### Real-Time TTS Latency Testing

Launch the streaming server and measure end-to-end latency:

```bash

# Launch the real‑time inference server

python demo/vibevoice_realtime_demo.py \
  --model_path microsoft/VibeVoice-Realtime-0.5B

# In another terminal, send a short prompt via the websocket client

python -c "
import websockets, asyncio, json
async def demo():
    uri = 'ws://localhost:8000'
    async with websockets.connect(uri) as ws:
        await ws.send(json.dumps({'text':'Hello, this is a latency test.'}))
        while True:
            chunk = await ws.recv()
            if not chunk: break
            print('audio chunk received')
asyncio.run(demo())
"

```

The websocket implementation streams audio chunks as they are synthesized, enabling the sub-200ms time-to-first-audio characteristic.

### Long-Form TTS Generation

For multi-speaker long-form synthesis:

```bash
pip install -e .[tts]   # install with TTS extras

python demo/vibevoice_asr_gradio_demo.py \
  --model_path microsoft/VibeVoice-1.5B \
  --speaker_names Alice Bob Carol Dave \
  --text_file demo/text_examples/1p_vibevoice.txt

```

This produces continuous audio up to 90 minutes while maintaining consistent speaker characteristics across the dialogue.

## Summary

- **VibeVoice-ASR** achieves a 3.42% average DER and 12.07% WER across 8+ languages on the MLC-Challenge benchmark, with specific results available in [`docs/vibevoice-asr.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-asr.md).
- **VibeVoice-Realtime-0.5B** delivers 2.00% WER on LibriSpeech with 0.695 speaker similarity, offering sub-200ms latency on T4 GPUs as documented in [`docs/vibevoice-realtime-0.5b.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-realtime-0.5b.md).
- **VibeVoice-TTS** claims state-of-the-art performance on 90-minute multi-speaker generation tasks, though numeric details reside in the external technical report referenced in [`docs/vibevoice-tts.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-tts.md).
- All benchmarks can be reproduced using the inference scripts in the `demo/` directory, with core model implementations located in `vibevoice/modular/`.

## Frequently Asked Questions

### What metrics does VibeVoice use to evaluate ASR performance?

VibeVoice-ASR reports four primary metrics: **DER** (diarization error rate) for speaker separation accuracy, **cpWER** (concatenated minimum-permutation word error rate) for multi-speaker transcription quality, **tcpWER** (time-constrained cpWER) for temporal alignment, and standard **WER** (word error rate). These metrics are calculated on datasets including MLC-Challenge, AISHELL-4, AMI, and AliMeeting as shown in [`docs/vibevoice-asr.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-asr.md).

### How does VibeVoice real-time TTS compare to VALL-E 2 and VoiceBox?

According to the evaluation tables in [`docs/vibevoice-realtime-0.5b.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-realtime-0.5b.md), VibeVoice-Realtime-0.5B achieves a **2.00% WER** on LibriSpeech test-clean, compared to VALL-E 2's 2.40% and VoiceBox's 1.90%. However, VibeVoice attains a higher speaker similarity score of **0.695**, indicating more faithful voice reproduction than competing systems while maintaining competitive intelligibility.

### What hardware achieves the reported 200ms latency for real-time TTS?

The **200ms time-to-first-audio** latency is achieved when running [`demo/vibevoice_realtime_demo.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_realtime_demo.py) on an NVIDIA T4 GPU. This measurement represents the duration from text input to the first audio chunk output via the websocket streaming interface, making it suitable for interactive conversational applications.

### Where can I find specific benchmark numbers for the long-form TTS model?

The repository documentation in [`docs/vibevoice-tts.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-tts.md) states that VibeVoice-TTS achieves state-of-the-art performance on long-form multi-speaker tasks but does not provide detailed numeric tables within the codebase. For exact quantitative results on the 90-minute generation benchmarks, consult the technical report linked in the documentation rather than the repository source files.