VibeVoice Performance Benchmarks: ASR, Real-Time TTS, and Long-Form Metrics

Microsoft VibeVoice achieves sub-5% diarization error rates on multilingual ASR, 2.00% WER on real-time zero-shot TTS, and state-of-the-art quality on 90-minute multi-speaker generation according to the repository evaluation tables.

The microsoft/VibeVoice repository provides comprehensive performance benchmarks for three distinct model families: automatic speech recognition (ASR), real-time streaming TTS, and long-form multi-speaker synthesis. These metrics are documented in the repository's markdown files and demonstrate competitive results against modern large-scale speech models on datasets like MLC-Challenge, LibriSpeech, and AISHELL-4.

VibeVoice-ASR Performance Metrics

The ASR model, defined in vibevoice/modular/modeling_vibevoice_asr.py, supports transcription of up to 60 minutes of audio in a single pass while maintaining diarization error rates (DER) below 5% for most languages. Evaluation tables in docs/vibevoice-asr.md report four key metrics: DER (diarization error rate), cpWER (concatenated minimum-permutation word error rate), tcpWER (time-constrained cpWER), and standard WER.

Multilingual Results on MLC-Challenge

The MLC-Challenge benchmark spans multiple languages with the following average performance:

Dataset Language DER (↓) cpWER (↓) tcpWER (↓) WER (↓)
MLC-Challenge English 4.28 11.48 13.02 7.99
MLC-Challenge French 3.80 18.80 19.64 15.21
MLC-Challenge German 1.04 17.10 17.26 16.30
Average (MLC-Challenge) – 3.42 14.81 15.66 12.07

These results position VibeVoice-ASR competitively with modern large-scale ASR systems, particularly excelling in German with a 1.04% DER.

Meeting and Conversation Benchmarks

On challenging multi-speaker meeting datasets, the model shows robust performance:

Dataset Language DER (↓) cpWER (↓) tcpWER (↓) WER (↓)
AISHELL-4 Chinese 6.77 24.99 25.35 21.40
AMI-IHM English 11.92 20.41 20.82 18.81
AMI-SDM English 13.43 28.82 29.80 24.65
AliMeeting Chinese 10.92 29.33 29.51 27.40

The higher DER values on AMI and AliMeeting reflect the complexity of far-field microphone arrays and overlapping speech in natural meeting environments.

Real-Time TTS Performance (0.5B Model)

VibeVoice-Realtime-0.5B targets zero-shot streaming text-to-speech with low latency. According to docs/vibevoice-realtime-0.5b.md, the model achieves 2.00% WER on LibriSpeech test-clean with a 0.695 speaker similarity score, outperforming comparable systems like VALL-E 2 (2.40% WER) while maintaining competitive quality with VoiceBox (1.90% WER).

Benchmark Model WER (↓) Speaker Similarity (↑)
LibriSpeech test-clean VibeVoice-Realtime-0.5B 2.00% 0.695
SEED test-en VibeVoice-Realtime-0.5B 2.05% 0.633

Latency Benchmarks

The real-time implementation delivers the first audible audio chunk in approximately 200 ms when running on a T4 GPU. This measurement comes from the streaming inference server defined in demo/vibevoice_realtime_demo.py, which uses websocket-based chunk delivery for interactive applications.

Long-Form Multi-Speaker TTS Results

For extended content generation, VibeVoice-TTS handles 90-minute multi-speaker dialogues while preserving speaker identity consistency. The documentation in docs/vibevoice-tts.md states that this model "achieves state-of-the-art performance on long-form multi-speaker speech generation tasks." Unlike the ASR and real-time TTS components, the repository does not expose numeric tables for this model family; exact figures are available in the linked technical report.

Reproducing the Benchmarks

You can validate these performance benchmarks for VibeVoice using the provided inference scripts and demo files.

ASR Inference and Metric Calculation

Run the reference script demo/vibevoice_asr_inference_from_file.py to generate transcripts for WER and DER calculation:


# demo/vibevoice_asr_inference_from_file.py

# Example usage (replace with your own audio path)

# python demo/vibevoice_asr_inference_from_file.py \

#   --model_path microsoft/VibeVoice-ASR \

#   --audio_files demo/asr_demo/demo1-chat.mp4

The script outputs JSON transcripts compatible with jiwer.wer and pyannote.metrics.diarization for official metric computation.

Real-Time TTS Latency Testing

Launch the streaming server and measure end-to-end latency:


# Launch the real‑time inference server

python demo/vibevoice_realtime_demo.py \
  --model_path microsoft/VibeVoice-Realtime-0.5B

# In another terminal, send a short prompt via the websocket client

python -c "
import websockets, asyncio, json
async def demo():
    uri = 'ws://localhost:8000'
    async with websockets.connect(uri) as ws:
        await ws.send(json.dumps({'text':'Hello, this is a latency test.'}))
        while True:
            chunk = await ws.recv()
            if not chunk: break
            print('audio chunk received')
asyncio.run(demo())
"

The websocket implementation streams audio chunks as they are synthesized, enabling the sub-200ms time-to-first-audio characteristic.

Long-Form TTS Generation

For multi-speaker long-form synthesis:

pip install -e .[tts]   # install with TTS extras

python demo/vibevoice_asr_gradio_demo.py \
  --model_path microsoft/VibeVoice-1.5B \
  --speaker_names Alice Bob Carol Dave \
  --text_file demo/text_examples/1p_vibevoice.txt

This produces continuous audio up to 90 minutes while maintaining consistent speaker characteristics across the dialogue.

Summary

  • VibeVoice-ASR achieves a 3.42% average DER and 12.07% WER across 8+ languages on the MLC-Challenge benchmark, with specific results available in docs/vibevoice-asr.md.
  • VibeVoice-Realtime-0.5B delivers 2.00% WER on LibriSpeech with 0.695 speaker similarity, offering sub-200ms latency on T4 GPUs as documented in docs/vibevoice-realtime-0.5b.md.
  • VibeVoice-TTS claims state-of-the-art performance on 90-minute multi-speaker generation tasks, though numeric details reside in the external technical report referenced in docs/vibevoice-tts.md.
  • All benchmarks can be reproduced using the inference scripts in the demo/ directory, with core model implementations located in vibevoice/modular/.

Frequently Asked Questions

What metrics does VibeVoice use to evaluate ASR performance?

VibeVoice-ASR reports four primary metrics: DER (diarization error rate) for speaker separation accuracy, cpWER (concatenated minimum-permutation word error rate) for multi-speaker transcription quality, tcpWER (time-constrained cpWER) for temporal alignment, and standard WER (word error rate). These metrics are calculated on datasets including MLC-Challenge, AISHELL-4, AMI, and AliMeeting as shown in docs/vibevoice-asr.md.

How does VibeVoice real-time TTS compare to VALL-E 2 and VoiceBox?

According to the evaluation tables in docs/vibevoice-realtime-0.5b.md, VibeVoice-Realtime-0.5B achieves a 2.00% WER on LibriSpeech test-clean, compared to VALL-E 2's 2.40% and VoiceBox's 1.90%. However, VibeVoice attains a higher speaker similarity score of 0.695, indicating more faithful voice reproduction than competing systems while maintaining competitive intelligibility.

What hardware achieves the reported 200ms latency for real-time TTS?

The 200ms time-to-first-audio latency is achieved when running demo/vibevoice_realtime_demo.py on an NVIDIA T4 GPU. This measurement represents the duration from text input to the first audio chunk output via the websocket streaming interface, making it suitable for interactive conversational applications.

Where can I find specific benchmark numbers for the long-form TTS model?

The repository documentation in docs/vibevoice-tts.md states that VibeVoice-TTS achieves state-of-the-art performance on long-form multi-speaker tasks but does not provide detailed numeric tables within the codebase. For exact quantitative results on the 90-minute generation benchmarks, consult the technical report linked in the documentation rather than the repository source files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →