How to Configure Multi-Speaker TTS with VibeVoice: Complete Implementation Guide
VibeVoice supports synthesizing dialogue with up to four distinct speakers in a single generation pass by parsing speaker-tagged scripts and optionally conditioning the model on voice samples through the VibeVoiceProcessor class.
Microsoft's open-source VibeVoice repository enables high-fidelity multi-speaker text-to-speech (TTS) synthesis through a script-based configuration system. To configure multi-speaker TTS with VibeVoice, developers must format input text with speaker identifiers and leverage the processor's built-in normalization and voice conditioning capabilities. This guide covers the complete implementation based on the actual source code in microsoft/VibeVoice.
Core Components of Multi-Speaker Configuration
The VibeVoiceProcessor Entry Point
The VibeVoiceProcessor class in vibevoice/processor/vibevoice_processor.py serves as the primary interface for multi-speaker TTS configuration. According to the source code, this processor handles script parsing, speaker ID normalization, and the construction of token sequences required by the diffusion-based model.
When you call the processor, it executes three critical operations:
- Script Parsing: The
_parse_scriptmethod recognizes lines formatted asSpeaker <id>: <text>using regex logic defined in lines 600–630, normalizing all speaker IDs to a zero-based index. - System Prompt Injection: It automatically prepends a static instruction that forces the model to maintain distinct voices per speaker (lines 41–42).
- Voice Sample Integration: If provided, it processes voice samples through
_create_voice_promptto condition the model on specific timbres (lines 75–80).
Speaker ID Parsing and Normalization
VibeVoice requires strict formatting for speaker identification. The parser expects utterances prefixed with Speaker <id>:, where the ID can be any integer. The _parse_script method automatically remaps these IDs to consecutive integers starting at 0, ensuring compatibility with the model's embedding layers.
The processor also accepts JSON-formatted scripts containing objects with speaker and text keys. When the input file path ends with .json, _process_single automatically detects the format and parses the structured dialogue (lines 55–60).
Voice Sample Conditioning
For consistent speaker characteristics across utterances, provide a short audio exemplar for each distinct speaker via the voice_samples parameter. The processor accepts file paths (.wav, .pt, .npy) or NumPy arrays.
If voice_samples=None, the model falls back to its learned speaker embeddings, generating distinct but synthetic voices for each ID. When samples are provided, the _create_voice_prompt function constructs additional token sequences that bias the diffusion head toward the provided timbres.
System Prompt Engineering
The multi-speaker capability relies on a hardcoded system prompt initialized in VibeVoiceProcessor.__init__:
" Transform the text provided by various speakers into speech output, utilizing the distinct voice of each respective speaker.\n"
This instruction, encoded at lines 71–74, ensures the LLM component of the architecture recognizes that multiple voices must be rendered in the output stream.
Implementation Methods
Basic Multi-Speaker Generation Without Voice Samples
For quick prototyping without voice cloning, use internal speaker embeddings by passing voice_samples=None:
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor
# Initialize the processor (auto-downloads tokenizer)
processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-1.5B")
# Format dialogue with Speaker prefixes (any integer IDs work)
script = """
Speaker 1: Hello, I'm Alice.
Speaker 2: Hi Alice, I'm Bob.
Speaker 1: Nice to meet you, Bob.
Speaker 3: And I'm Carol, joining the conversation.
"""
# Generate token batch
batch = processor(
text=script,
voice_samples=None,
padding=True,
return_tensors="pt"
)
# batch contains input_ids, attention_mask, speech_input_mask
# Pass to model for waveform generation
The processor normalizes the IDs 1, 2, 3 to 0, 1, 2 internally, allowing up to four distinct speakers per the model's architectural constraints.
Conditioning on Speaker Voice Samples
To maintain consistent, recognizable voices across turns, provide one sample per speaker. The order of samples must match the normalized speaker IDs:
import numpy as np
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor
processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-1.5B")
script = """
Speaker 0: Good morning, everyone.
Speaker 1: Good morning! I'm excited for today.
Speaker 0: Let's get started.
"""
# Load voice checkpoints or raw audio files
alice_sample = "demo/voices/streaming_model/sp-Alice_woman.pt"
bob_sample = "demo/voices/streaming_model/sp-Bob_man.pt"
# Order matches speaker IDs in script: 0 -> Alice, 1 -> Bob
voice_samples = [alice_sample, bob_sample]
batch = processor(
text=script,
voice_samples=voice_samples,
padding=True,
return_tensors="pt"
)
The _process_single method pairs each sample with its corresponding speaker ID via _create_voice_prompt, constructing the voice_tokens, voice_speech_inputs, and voice_speech_masks tensors required by the model's conditioning mechanism.
JSON-Style Script Configuration
For programmatic pipelines, use structured JSON instead of plain text:
[
{"speaker": 0, "text": "Welcome to the podcast."},
{"speaker": 1, "text": "Thanks for having me!"},
{"speaker": 0, "text": "Let's discuss the latest research."}
]
Process the file directly:
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor
processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-1.5B")
# Pass the file path; processor detects .json suffix
batch = processor(
text="path/to/dialogue.json",
voice_samples=None,
padding=True,
return_tensors="pt"
)
The suffix detection logic in _process_single routes JSON inputs through a specialized parser that extracts the speaker-text pairs without requiring manual string formatting.
Using the Streaming Demo with Named Speakers
For command-line experimentation, the repository includes a convenience wrapper that maps human-readable names to pretrained checkpoints. The VoiceMapper class in demo/realtime_model_inference_from_file.py (lines 38–84) scans the demo/voices/streaming_model directory and resolves names to file paths:
python demo/realtime_model_inference_from_file.py \
--model_path microsoft/VibeVoice-Realtime-0.5B \
--txt_path demo/text_examples/1p_vibevoice.txt \
--speaker_name "Wayne"
The VoiceMapper.get_voice_path method performs fuzzy matching against available voice files, then passes the resolved checkpoint as a voice sample to the streaming inference engine.
Summary
- VibeVoice supports up to four speakers per generation pass through the
VibeVoiceProcessorclass invibevoice/processor/vibevoice_processor.py. - Script formatting requires either
Speaker <id>: <text>prefixes or JSON objects withspeakerandtextkeys. - Speaker IDs are normalized to zero-based indices automatically by the
_parse_scriptmethod (lines 600–630). - Voice samples are optional but recommended for consistency; provide one sample per speaker in the order of their normalized IDs.
- System prompt injection happens automatically to enforce distinct voice generation for each speaker turn.
- Command-line demos support human-readable speaker names via the
VoiceMapperutility in the demo scripts.
Frequently Asked Questions
How many speakers can VibeVoice handle in one generation?
VibeVoice supports up to four distinct speakers in a single generation pass. The _parse_script method in vibevoice/processor/vibevoice_processor.py normalizes speaker IDs to integers 0–3, and the underlying model architecture is optimized for this maximum cardinality.
What audio formats work for voice samples?
The processor accepts voice samples as file paths (.wav, .pt, .npy) or NumPy arrays. The _create_voice_prompt method handles format detection and tensor conversion internally. Samples should be short audio clips (typically 3–10 seconds) that clearly capture the target speaker's timbre.
Can I use speaker names instead of numbers in scripts?
For the core VibeVoiceProcessor, you must use numeric IDs (e.g., Speaker 1:) or JSON format. However, the streaming demo in demo/realtime_model_inference_from_file.py provides a VoiceMapper class that translates human-readable names like "Wayne" to numeric indices and corresponding voice checkpoints automatically.
Why does the processor normalize speaker IDs to start at zero?
The normalization ensures compatibility with the model's embedding layers, which expect consecutive integer indices beginning at 0. The regex and remapping logic in _parse_script (lines 602–610) handles this automatically, so you can use any integer IDs in your input script without manual reordering.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →