# How to Customize Voice Prompts in VibeVoice-Realtime: A Complete Technical Guide

> Learn to customize voice prompts in VibeVoice-Realtime. This guide details how VibeVoice encodes audio snippets into the KV cache for consistent speaker characteristics during streaming inference.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**VibeVoice-Realtime encodes a short audio snippet into the model's KV cache via the `_create_voice_prompt` method in [`vibevoice/processor/vibevoice_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_processor.py), allowing you to pre-compute a speaker-specific cache file that seeds the diffusion decoder for consistent voice characteristics during streaming inference.**

The **microsoft/VibeVoice** repository implements voice cloning through a unique voice prompt mechanism that embeds speaker identity directly into the model's key-value cache before text generation begins. To **customize voice prompts in VibeVoice-Realtime**, you must generate a pre-filled KV cache containing your reference audio tokens, which the streaming inference engine then uses to condition the diffusion-based speech synthesis pipeline.

## Understanding the Voice Prompt Architecture

VibeVoice-Realtime processes voice prompts as special token sequences that prime the acoustic diffusion model. The architecture involves three core components working across specific source files.

### Voice Prompt Creation via `_create_voice_prompt`

The foundation of voice customization lies in the `_create_voice_prompt` method located at [`vibevoice/processor/vibevoice_processor.py#L400-L460`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_processor.py#L400-L460). This private method accepts one or more audio samples, converts them into VAE tokens, and constructs a structured token sequence following the format:

```

Voice input:
 Speaker N: <speech_start> <vae_token…> <speech_end>

```

The method returns three critical tensors: **token IDs**, **raw audio tensors**, and a **speech mask** indicating which tokens correspond to audio versus text.

### Tokenisation and KV Cache Integration

Once created, voice prompts merge with user text through `process_input_with_cached_prompt` in [`vibevoice/processor/vibevoice_streaming_processor.py#L170-L220`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_streaming_processor.py#L170-L220). This method builds a `BatchEncoding` containing `input_ids`, `attention_mask`, and `speech_input_mask`, preparing the combined representation for the transformer model.

### Streaming Inference Consumption

The [`vibevoice/modular/modeling_vibevoice_streaming_inference.py#L622-L676`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_streaming_inference.py#L622-L676) file contains the core generation loop. Here, `model.generate` receives the encoded inputs plus a **cached_prompt** parameter containing the pre-filled KV cache. The voice prompt tokens seed the acoustic diffusion, ensuring generated speech inherits the reference speaker's timbre during windowed text prefill and incremental generation.

## Step-by-Step Guide to Creating a Custom Voice Prompt

To generate a production-ready custom voice prompt, follow this workflow that captures and persists the KV cache for reuse across streaming sessions.

1. **Record clean reference audio** – Capture 3–5 seconds of high-quality speech for your target speaker at the model's expected sampling rate (typically 16 kHz).

2. **Load and process the sample** – Use the processor's audio loading utilities to convert your WAV file into the required tensor format.

3. **Generate voice tokens** – Invoke `_create_voice_prompt` to build the token sequence, audio tensors, and speech masks from your reference sample.

4. **Execute a prefill pass** – Run `model.generate` with `max_new_tokens=0` to process the voice prompt through the transformer layers, capturing the resulting KV cache.

5. **Persist the cache** – Save the output dictionary containing hidden states and cache tensors as a PyTorch checkpoint (`.pt` file).

6. **Deploy in streaming** – Load the saved checkpoint as a `cached_prompt` in your streaming service configuration.

## Code Implementation: Generating a Custom Voice Prompt

The following implementation demonstrates the complete pipeline from audio loading through cache persistence:

```python
import torch
import numpy as np
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor
from vibevoice.modular.modeling_vibevoice_streaming import VibeVoiceStreamingModel

# 1️⃣ Initialise processor & model (use the streaming checkpoint)

processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")
model = VibeVoiceStreamingModel.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")

# 2️⃣ Load your own speaker audio (wav → np.ndarray)

#    Here `sample.wav` is a 16 kHz mono file.

wav, sr = processor.audio_processor._load_audio_from_path("sample.wav")
assert sr == processor.speech_sampling_rate

# 3️⃣ Build the voice‑prompt tokens, audio tensors & mask

voice_tokens, voice_inputs, voice_masks = processor._create_voice_prompt([wav])

# 4️⃣ Encode a dummy script (the prompt itself) – we only need the voice part.

#    The processor will prepend a system prompt automatically.

dummy_script = " Voice input:\n"  # minimal script – actual speech comes from voice_inputs

encoding = processor.encode(
    dummy_script,
    speech_inputs=voice_inputs,
    speech_input_masks=voice_masks,
    return_tensors="pt",
)

# 5️⃣ Run a *prefill* pass to capture KV cache

outputs = model.generate(
    **encoding,
    max_new_tokens=0,                # no new text generation yet

    return_dict_in_generate=True,
    output_hidden_states=True,
)

# 6️⃣ Save the cached KV cache (prefilled prompt) for later streaming use

prefilled_path = "my_custom_prompt.pt"
torch.save(outputs, prefilled_path)
print(f"Custom voice prompt saved to {prefilled_path}")

```

## Deploying Custom Prompts in Streaming Applications

After generating your `.pt` checkpoint, integrate it into the streaming inference pipeline. The reference implementation in [`demo/web/app.py#L160-L187`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py#L160-L187) demonstrates loading pre-computed voice caches into a `StreamingTTSService` instance.

```python
from demo.web.app import StreamingTTSService

service = StreamingTTSService(
    model_path="microsoft/VibeVoice-Realtime-0.5B",
    voice_key="my_custom",                # arbitrary identifier

    voice_presets={"my_custom": "my_custom_prompt.pt"},
)

# Now `service.stream("Hello world!")` will synthesize with your voice.

```

This pattern allows you to maintain a registry of voice presets (stored in `demo/voices/streaming_model/` in the reference implementation) that load instantly without recomputing the KV cache for each request.

## Key Source Files for Voice Prompt Customization

Understanding these specific files enables advanced customization beyond the standard API:

- **[`vibevoice/processor/vibevoice_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_processor.py)** – Contains `_create_voice_prompt` and audio preprocessing logic for converting raw waveforms into VAE token sequences.

- **[`vibevoice/processor/vibevoice_streaming_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_streaming_processor.py)** – Implements `process_input_with_cached_prompt`, which merges live text inputs with pre-computed voice prompt caches.

- **[`vibevoice/modular/modeling_vibevoice_streaming_inference.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_streaming_inference.py)** – Houses the streaming generation loop that consumes prefixed KV caches and performs windowed diffusion-based speech synthesis.

- **[`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py)** – Reference Flask-style service demonstrating production deployment patterns for cached voice prompts.

- **`demo/voices/streaming_model/*.pt`** – Directory containing shipped example prompts (e.g., `en-Carter_man.pt`) that demonstrate the expected checkpoint format.

## Summary

- **Voice prompts** in VibeVoice-Realtime are encoded into the KV cache before streaming begins, not processed in real-time during generation.
- Use **`_create_voice_prompt`** in [`vibevoice_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice_processor.py) to convert reference audio into the required token and mask representation.
- **Prefill the model** with `max_new_tokens=0` to capture the voice-specific KV cache, then save it as a `.pt` checkpoint for instant loading.
- The **streaming processor** merges user text with cached voice prompts via `process_input_with_cached_prompt`, enabling zero-latency voice switching.
- Deploy custom voices by referencing saved checkpoints in your `StreamingTTSService` configuration, following the pattern established in [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py).

## Frequently Asked Questions

### How long should a reference audio sample be for optimal voice prompt quality?

A clean recording of **3–5 seconds** provides sufficient speaker characteristics for the VAE tokenization pipeline while maintaining manageable cache file sizes. The `_create_voice_prompt` method processes these samples into compact token sequences that capture timbre and prosodic features without requiring extended utterances.

### Can I use multiple speaker samples to create a single voice prompt?

Yes. The `_create_voice_prompt` method accepts a **list of audio arrays** as its input parameter, allowing you to concatenate multiple reference utterances from the same speaker. This improves robustness by averaging speaker characteristics across different phonetic contexts, though the method returns unified token and mask tensors representing the combined voice identity.

### Why does VibeVoice-Realtime use KV cache prefill instead of processing voice prompts during generation?

The **KV cache prefill** strategy eliminates inference latency for voice conditioning. By pre-computing the transformer states for voice prompt tokens and persisting them as `.pt` files, the streaming pipeline in [`modeling_vibevoice_streaming_inference.py`](https://github.com/microsoft/VibeVoice/blob/main/modeling_vibevoice_streaming_inference.py) can immediately begin text processing without recomputing speaker embeddings. This architecture enables the real-time performance required for interactive applications while maintaining consistent speaker identity across long-form generation.

### Where are the example voice prompt files located in the repository?

Pre-generated voice prompts ship as PyTorch checkpoints in the [`demo/voices/streaming_model/`](https://github.com/microsoft/VibeVoice/tree/main/demo/voices/streaming_model) directory. These files (such as `en-Carter_man.pt`) demonstrate the expected output format from the prefill process and can be loaded directly via the `voice_presets` parameter in `StreamingTTSService` as shown in [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py).