What Are the Training Data Sources for VibeVoice? A Complete Guide to the Microsoft Speech Corpus
VibeVoice models are trained on publicly available multilingual speech corpora including MLC-Challenge, AISHELL-4, AMI, and AliMeeting, supporting over 50 languages with robust handling of noisy, natural audio conditions.
Microsoft VibeVoice is an open-source speech processing framework that powers both automatic speech recognition (ASR) and text-to-speech (TTS) models across more than 50 languages. Understanding the VibeVoice training data sources is essential for developers looking to fine-tune the models or evaluate their performance on specific acoustic conditions. According to the microsoft/VibeVoice repository, the training pipeline leverages continuous acoustic tokenizers and JSON-annotated datasets to ingest long-form audio without chunking.
VibeVoice-ASR Training Data Sources
The VibeVoice-ASR model documented in docs/vibevoice-asr.md draws from several key public benchmarks that provide multilingual coverage and diverse acoustic environments.
MLC-Challenge Multilingual Benchmark
The primary evaluation and training source is the MLC-Challenge, a multilingual long-form speech benchmark covering English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese. This corpus enables the model to handle extended audio sequences across diverse linguistic contexts.
Meeting Speech Corpora
For robust conversational ASR, VibeVoice incorporates meeting-style recordings:
- AISHELL-4: Mandarin meeting recordings providing Chinese conversational speech
- AMI-IHM / AMI-SDM: English meeting recordings captured with both individual headset microphones (IHM) and single distant microphones (SDM)
- AliMeeting: Chinese meeting recordings offering additional Mandarin conversational data
Streaming Support for HuggingFace Datasets
The inference pipeline in demo/vibevoice_asr_inference_from_file.py (lines 278-302) demonstrates support for streaming any HuggingFace dataset, including openslr/librispeech_asr and Common Voice variants. This architecture allows researchers to supplement training with additional public corpora like LibriSpeech without modifying the core training scripts.
VibeVoice-TTS Training Data Sources
As documented in docs/vibevoice-tts.md (lines 123-129), the TTS training data follows a specific curation strategy: the dataset does not contain music and was intentionally left "noisy" to preserve natural acoustic conditions including background sounds. While the exact corpus list isn't hard-coded in the repository, the model was built on large public collections such as LibriTTS and Common Voice, focusing on natural conversational audio rather than denoised studio recordings.
How VibeVoice Consumes Training Data
VibeVoice employs a unique data ingestion pipeline that combines continuous tokenization with structured JSON annotations.
Dataset Format and JSON Schema
The VibeVoiceASRDataset class in finetuning-asr/lora_finetune.py (lines 55-73, 78-96) expects samples formatted as JSON files containing:
- Audio file paths
segmentsarrays with speaker IDs, timestamps, and transcriptions- Optional
customized_contextfields for domain adaptation
Tokenization and Long-Form Processing
The models use continuous acoustic and semantic tokenizers operating at 7.5 Hz, enabling ingestion of audio streams up to one hour in length without chunking. This architecture supports the long-form nature of the MLC-Challenge and meeting corpora.
Fine-Tuning with Custom Datasets
You can fine-tune VibeVoice-ASR on custom data using LoRA adapters. The script expects the same JSON schema used in the public training sets:
# finetuning-asr/lora_finetune.py
from vibevoice.modular.modeling_vibevoice_asr import VibeVoiceASRForConditionalGeneration
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
# Dataset loader expects JSON with audio paths and segment annotations
# See VibeVoiceASRDataset implementation (lines 55-73)
Run the fine-tuning script:
python finetuning-asr/lora_finetune.py \
--model_path microsoft/VibeVoice-ASR \
--data_dir ./my_custom_dataset \
--output_dir ./asr_lora_ckpt \
--per_device_train_batch_size 2 \
--num_train_epochs 3 \
--learning_rate 5e-5 \
--lora_r 16 \
--lora_alpha 32
Loading Public Datasets for Inference
The inference script supports streaming public datasets for evaluation or preprocessing:
# demo/vibevoice_asr_inference_from_file.py
from datasets import load_dataset
import torch
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
from vibevoice.modular.modeling_vibevoice_asr import VibeVoiceASRForConditionalGeneration
# Load public training data source (e.g., LibriSpeech)
dataset = load_dataset(
"openslr/librispeech_asr",
split="test",
streaming=True,
)
# Concatenate utterances into long-form chunks (up to 1 hour)
# Implementation reference: lines 278-312
def concatenate(samples, max_sec=3600):
# Concatenation logic for long-form processing
pass
long_audio = concatenate(dataset, max_sec=1800)
# Initialize model
processor = VibeVoiceASRProcessor.from_pretrained(
"microsoft/VibeVoice-ASR",
language_model_pretrained_name="Qwen/Qwen2.5-7B",
)
model = VibeVoiceASRForConditionalGeneration.from_pretrained(
"microsoft/VibeVoice-ASR",
device_map="auto",
torch_dtype=torch.bfloat16,
)
# Generate transcription
outputs = model.generate(
**processor(long_audio, return_tensors="pt", sampling_rate=24000),
max_new_tokens=1024,
)
print(processor.decode(outputs[0], skip_special_tokens=True))
Summary
- VibeVoice-ASR trains on the MLC-Challenge benchmark (11 languages) plus meeting corpora (AISHELL-4, AMI, AliMeeting) for conversational speech robustness.
- VibeVoice-TTS uses large public collections like LibriTTS and Common Voice, intentionally preserving natural noise and excluding music.
- The training pipeline in
finetuning-asr/lora_finetune.pyconsumes JSON-annotated audio with speaker timestamps and supports LoRA fine-tuning at 7.5 Hz tokenization. - Any HuggingFace dataset can be streamed via the inference script, enabling flexible evaluation on additional public speech corpora.
Frequently Asked Questions
Does VibeVoice use proprietary training data?
No. According to the microsoft/VibeVoice source code, all referenced VibeVoice training data sources are publicly available corpora. The ASR model specifically lists MLC-Challenge, AISHELL-4, AMI, and AliMeeting, while the inference script can stream any public HuggingFace dataset like LibriSpeech or Common Voice.
Can I train VibeVoice on my own speech dataset?
Yes. The finetuning-asr/lora_finetune.py script supports custom datasets using the same JSON format as the public training data. You must provide audio files paired with JSON annotations containing speaker IDs, timestamps, and transcriptions. The implementation uses LoRA adapters (lora_r=16, lora_alpha=32) for efficient fine-tuning.
Why does VibeVoice-TTS training data exclude music?
As documented in docs/vibevoice-tts.md, the TTS training set excludes music to focus on natural speech patterns. The data was intentionally left noisy (preserving background sounds and BGM) to train the model on realistic acoustic conditions rather than studio-clean audio.
What tokenizer frequency does VibeVoice use for processing training audio?
VibeVoice processes training audio using continuous acoustic and semantic tokenizers operating at 7.5 Hz, as implemented in the dataset loaders. This enables the model to handle long-form audio up to one hour without chunking, supporting the extended sequences found in the MLC-Challenge and meeting corpora.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →