How to Customize Voice Prompts in VibeVoice-Realtime: A Complete Technical Guide

VibeVoice-Realtime encodes a short audio snippet into the model's KV cache via the _create_voice_prompt method in vibevoice/processor/vibevoice_processor.py, allowing you to pre-compute a speaker-specific cache file that seeds the diffusion decoder for consistent voice characteristics during streaming inference.

The microsoft/VibeVoice repository implements voice cloning through a unique voice prompt mechanism that embeds speaker identity directly into the model's key-value cache before text generation begins. To customize voice prompts in VibeVoice-Realtime, you must generate a pre-filled KV cache containing your reference audio tokens, which the streaming inference engine then uses to condition the diffusion-based speech synthesis pipeline.

Understanding the Voice Prompt Architecture

VibeVoice-Realtime processes voice prompts as special token sequences that prime the acoustic diffusion model. The architecture involves three core components working across specific source files.

Voice Prompt Creation via _create_voice_prompt

The foundation of voice customization lies in the _create_voice_prompt method located at vibevoice/processor/vibevoice_processor.py#L400-L460. This private method accepts one or more audio samples, converts them into VAE tokens, and constructs a structured token sequence following the format:


Voice input:
 Speaker N: <speech_start> <vae_token…> <speech_end>

The method returns three critical tensors: token IDs, raw audio tensors, and a speech mask indicating which tokens correspond to audio versus text.

Tokenisation and KV Cache Integration

Once created, voice prompts merge with user text through process_input_with_cached_prompt in vibevoice/processor/vibevoice_streaming_processor.py#L170-L220. This method builds a BatchEncoding containing input_ids, attention_mask, and speech_input_mask, preparing the combined representation for the transformer model.

Streaming Inference Consumption

The vibevoice/modular/modeling_vibevoice_streaming_inference.py#L622-L676 file contains the core generation loop. Here, model.generate receives the encoded inputs plus a cached_prompt parameter containing the pre-filled KV cache. The voice prompt tokens seed the acoustic diffusion, ensuring generated speech inherits the reference speaker's timbre during windowed text prefill and incremental generation.

Step-by-Step Guide to Creating a Custom Voice Prompt

To generate a production-ready custom voice prompt, follow this workflow that captures and persists the KV cache for reuse across streaming sessions.

  1. Record clean reference audio – Capture 3–5 seconds of high-quality speech for your target speaker at the model's expected sampling rate (typically 16 kHz).

  2. Load and process the sample – Use the processor's audio loading utilities to convert your WAV file into the required tensor format.

  3. Generate voice tokens – Invoke _create_voice_prompt to build the token sequence, audio tensors, and speech masks from your reference sample.

  4. Execute a prefill pass – Run model.generate with max_new_tokens=0 to process the voice prompt through the transformer layers, capturing the resulting KV cache.

  5. Persist the cache – Save the output dictionary containing hidden states and cache tensors as a PyTorch checkpoint (.pt file).

  6. Deploy in streaming – Load the saved checkpoint as a cached_prompt in your streaming service configuration.

Code Implementation: Generating a Custom Voice Prompt

The following implementation demonstrates the complete pipeline from audio loading through cache persistence:

import torch
import numpy as np
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor
from vibevoice.modular.modeling_vibevoice_streaming import VibeVoiceStreamingModel

# 1️⃣ Initialise processor & model (use the streaming checkpoint)

processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")
model = VibeVoiceStreamingModel.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")

# 2️⃣ Load your own speaker audio (wav → np.ndarray)

#    Here `sample.wav` is a 16 kHz mono file.

wav, sr = processor.audio_processor._load_audio_from_path("sample.wav")
assert sr == processor.speech_sampling_rate

# 3️⃣ Build the voice‑prompt tokens, audio tensors & mask

voice_tokens, voice_inputs, voice_masks = processor._create_voice_prompt([wav])

# 4️⃣ Encode a dummy script (the prompt itself) – we only need the voice part.

#    The processor will prepend a system prompt automatically.

dummy_script = " Voice input:\n"  # minimal script – actual speech comes from voice_inputs

encoding = processor.encode(
    dummy_script,
    speech_inputs=voice_inputs,
    speech_input_masks=voice_masks,
    return_tensors="pt",
)

# 5️⃣ Run a *prefill* pass to capture KV cache

outputs = model.generate(
    **encoding,
    max_new_tokens=0,                # no new text generation yet

    return_dict_in_generate=True,
    output_hidden_states=True,
)

# 6️⃣ Save the cached KV cache (prefilled prompt) for later streaming use

prefilled_path = "my_custom_prompt.pt"
torch.save(outputs, prefilled_path)
print(f"Custom voice prompt saved to {prefilled_path}")

Deploying Custom Prompts in Streaming Applications

After generating your .pt checkpoint, integrate it into the streaming inference pipeline. The reference implementation in demo/web/app.py#L160-L187 demonstrates loading pre-computed voice caches into a StreamingTTSService instance.

from demo.web.app import StreamingTTSService

service = StreamingTTSService(
    model_path="microsoft/VibeVoice-Realtime-0.5B",
    voice_key="my_custom",                # arbitrary identifier

    voice_presets={"my_custom": "my_custom_prompt.pt"},
)

# Now `service.stream("Hello world!")` will synthesize with your voice.

This pattern allows you to maintain a registry of voice presets (stored in demo/voices/streaming_model/ in the reference implementation) that load instantly without recomputing the KV cache for each request.

Key Source Files for Voice Prompt Customization

Understanding these specific files enables advanced customization beyond the standard API:

Summary

  • Voice prompts in VibeVoice-Realtime are encoded into the KV cache before streaming begins, not processed in real-time during generation.
  • Use _create_voice_prompt in vibevoice_processor.py to convert reference audio into the required token and mask representation.
  • Prefill the model with max_new_tokens=0 to capture the voice-specific KV cache, then save it as a .pt checkpoint for instant loading.
  • The streaming processor merges user text with cached voice prompts via process_input_with_cached_prompt, enabling zero-latency voice switching.
  • Deploy custom voices by referencing saved checkpoints in your StreamingTTSService configuration, following the pattern established in demo/web/app.py.

Frequently Asked Questions

How long should a reference audio sample be for optimal voice prompt quality?

A clean recording of 3–5 seconds provides sufficient speaker characteristics for the VAE tokenization pipeline while maintaining manageable cache file sizes. The _create_voice_prompt method processes these samples into compact token sequences that capture timbre and prosodic features without requiring extended utterances.

Can I use multiple speaker samples to create a single voice prompt?

Yes. The _create_voice_prompt method accepts a list of audio arrays as its input parameter, allowing you to concatenate multiple reference utterances from the same speaker. This improves robustness by averaging speaker characteristics across different phonetic contexts, though the method returns unified token and mask tensors representing the combined voice identity.

Why does VibeVoice-Realtime use KV cache prefill instead of processing voice prompts during generation?

The KV cache prefill strategy eliminates inference latency for voice conditioning. By pre-computing the transformer states for voice prompt tokens and persisting them as .pt files, the streaming pipeline in modeling_vibevoice_streaming_inference.py can immediately begin text processing without recomputing speaker embeddings. This architecture enables the real-time performance required for interactive applications while maintaining consistent speaker identity across long-form generation.

Where are the example voice prompt files located in the repository?

Pre-generated voice prompts ship as PyTorch checkpoints in the demo/voices/streaming_model/ directory. These files (such as en-Carter_man.pt) demonstrate the expected output format from the prefill process and can be loaded directly via the voice_presets parameter in StreamingTTSService as shown in demo/web/app.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →