# How to Configure Multi-Speaker TTS with VibeVoice: Complete Implementation Guide

> Learn how to configure multi-speaker TTS with VibeVoice. This guide shows you how to synthesize dialogue with up to four distinct speakers in one pass using speaker-tagged scripts and voice samples.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**VibeVoice supports synthesizing dialogue with up to four distinct speakers in a single generation pass by parsing speaker-tagged scripts and optionally conditioning the model on voice samples through the `VibeVoiceProcessor` class.**

Microsoft's open-source **VibeVoice** repository enables high-fidelity multi-speaker text-to-speech (TTS) synthesis through a script-based configuration system. To configure multi-speaker TTS with VibeVoice, developers must format input text with speaker identifiers and leverage the processor's built-in normalization and voice conditioning capabilities. This guide covers the complete implementation based on the actual source code in `microsoft/VibeVoice`.

## Core Components of Multi-Speaker Configuration

### The VibeVoiceProcessor Entry Point

The **`VibeVoiceProcessor`** class in [`vibevoice/processor/vibevoice_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_processor.py) serves as the primary interface for multi-speaker TTS configuration. According to the source code, this processor handles script parsing, speaker ID normalization, and the construction of token sequences required by the diffusion-based model.

When you call the processor, it executes three critical operations:

- **Script Parsing**: The `_parse_script` method recognizes lines formatted as `Speaker <id>: <text>` using regex logic defined in lines 600–630, normalizing all speaker IDs to a zero-based index.
- **System Prompt Injection**: It automatically prepends a static instruction that forces the model to maintain distinct voices per speaker (lines 41–42).
- **Voice Sample Integration**: If provided, it processes voice samples through `_create_voice_prompt` to condition the model on specific timbres (lines 75–80).

### Speaker ID Parsing and Normalization

VibeVoice requires strict formatting for speaker identification. The parser expects utterances prefixed with `Speaker <id>:`, where the ID can be any integer. The `_parse_script` method automatically remaps these IDs to consecutive integers starting at 0, ensuring compatibility with the model's embedding layers.

The processor also accepts **JSON-formatted scripts** containing objects with `speaker` and `text` keys. When the input file path ends with `.json`, `_process_single` automatically detects the format and parses the structured dialogue (lines 55–60).

### Voice Sample Conditioning

For consistent speaker characteristics across utterances, provide a short audio exemplar for each distinct speaker via the `voice_samples` parameter. The processor accepts file paths (`.wav`, `.pt`, `.npy`) or NumPy arrays.

If `voice_samples=None`, the model falls back to its learned speaker embeddings, generating distinct but synthetic voices for each ID. When samples are provided, the `_create_voice_prompt` function constructs additional token sequences that bias the diffusion head toward the provided timbres.

### System Prompt Engineering

The multi-speaker capability relies on a hardcoded system prompt initialized in `VibeVoiceProcessor.__init__`:

```python
" Transform the text provided by various speakers into speech output, utilizing the distinct voice of each respective speaker.\n"

```

This instruction, encoded at lines 71–74, ensures the LLM component of the architecture recognizes that multiple voices must be rendered in the output stream.

## Implementation Methods

### Basic Multi-Speaker Generation Without Voice Samples

For quick prototyping without voice cloning, use internal speaker embeddings by passing `voice_samples=None`:

```python
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor

# Initialize the processor (auto-downloads tokenizer)

processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-1.5B")

# Format dialogue with Speaker prefixes (any integer IDs work)

script = """
Speaker 1: Hello, I'm Alice.
Speaker 2: Hi Alice, I'm Bob.
Speaker 1: Nice to meet you, Bob.
Speaker 3: And I'm Carol, joining the conversation.
"""

# Generate token batch

batch = processor(
    text=script, 
    voice_samples=None, 
    padding=True, 
    return_tensors="pt"
)

# batch contains input_ids, attention_mask, speech_input_mask

# Pass to model for waveform generation

```

The processor normalizes the IDs 1, 2, 3 to 0, 1, 2 internally, allowing up to four distinct speakers per the model's architectural constraints.

### Conditioning on Speaker Voice Samples

To maintain consistent, recognizable voices across turns, provide one sample per speaker. The order of samples must match the normalized speaker IDs:

```python
import numpy as np
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor

processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-1.5B")

script = """
Speaker 0: Good morning, everyone.
Speaker 1: Good morning! I'm excited for today.
Speaker 0: Let's get started.
"""

# Load voice checkpoints or raw audio files

alice_sample = "demo/voices/streaming_model/sp-Alice_woman.pt"
bob_sample = "demo/voices/streaming_model/sp-Bob_man.pt"

# Order matches speaker IDs in script: 0 -> Alice, 1 -> Bob

voice_samples = [alice_sample, bob_sample]

batch = processor(
    text=script, 
    voice_samples=voice_samples, 
    padding=True, 
    return_tensors="pt"
)

```

The `_process_single` method pairs each sample with its corresponding speaker ID via `_create_voice_prompt`, constructing the `voice_tokens`, `voice_speech_inputs`, and `voice_speech_masks` tensors required by the model's conditioning mechanism.

### JSON-Style Script Configuration

For programmatic pipelines, use structured JSON instead of plain text:

```json
[
  {"speaker": 0, "text": "Welcome to the podcast."},
  {"speaker": 1, "text": "Thanks for having me!"},
  {"speaker": 0, "text": "Let's discuss the latest research."}
]

```

Process the file directly:

```python
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor

processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-1.5B")

# Pass the file path; processor detects .json suffix

batch = processor(
    text="path/to/dialogue.json",
    voice_samples=None,
    padding=True,
    return_tensors="pt"
)

```

The suffix detection logic in `_process_single` routes JSON inputs through a specialized parser that extracts the speaker-text pairs without requiring manual string formatting.

### Using the Streaming Demo with Named Speakers

For command-line experimentation, the repository includes a convenience wrapper that maps human-readable names to pretrained checkpoints. The `VoiceMapper` class in [`demo/realtime_model_inference_from_file.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/realtime_model_inference_from_file.py) (lines 38–84) scans the `demo/voices/streaming_model` directory and resolves names to file paths:

```bash
python demo/realtime_model_inference_from_file.py \
    --model_path microsoft/VibeVoice-Realtime-0.5B \
    --txt_path demo/text_examples/1p_vibevoice.txt \
    --speaker_name "Wayne"

```

The `VoiceMapper.get_voice_path` method performs fuzzy matching against available voice files, then passes the resolved checkpoint as a voice sample to the streaming inference engine.

## Summary

- **VibeVoice supports up to four speakers** per generation pass through the `VibeVoiceProcessor` class in [`vibevoice/processor/vibevoice_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_processor.py).
- **Script formatting requires** either `Speaker <id>: <text>` prefixes or JSON objects with `speaker` and `text` keys.
- **Speaker IDs are normalized** to zero-based indices automatically by the `_parse_script` method (lines 600–630).
- **Voice samples are optional** but recommended for consistency; provide one sample per speaker in the order of their normalized IDs.
- **System prompt injection** happens automatically to enforce distinct voice generation for each speaker turn.
- **Command-line demos** support human-readable speaker names via the `VoiceMapper` utility in the demo scripts.

## Frequently Asked Questions

### How many speakers can VibeVoice handle in one generation?

VibeVoice supports up to **four distinct speakers** in a single generation pass. The `_parse_script` method in [`vibevoice/processor/vibevoice_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_processor.py) normalizes speaker IDs to integers 0–3, and the underlying model architecture is optimized for this maximum cardinality.

### What audio formats work for voice samples?

The processor accepts voice samples as **file paths** (`.wav`, `.pt`, `.npy`) or **NumPy arrays**. The `_create_voice_prompt` method handles format detection and tensor conversion internally. Samples should be short audio clips (typically 3–10 seconds) that clearly capture the target speaker's timbre.

### Can I use speaker names instead of numbers in scripts?

For the core `VibeVoiceProcessor`, you must use numeric IDs (e.g., `Speaker 1:`) or JSON format. However, the **streaming demo** in [`demo/realtime_model_inference_from_file.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/realtime_model_inference_from_file.py) provides a `VoiceMapper` class that translates human-readable names like "Wayne" to numeric indices and corresponding voice checkpoints automatically.

### Why does the processor normalize speaker IDs to start at zero?

The normalization ensures compatibility with the model's **embedding layers**, which expect consecutive integer indices beginning at 0. The regex and remapping logic in `_parse_script` (lines 602–610) handles this automatically, so you can use any integer IDs in your input script without manual reordering.