How to Customize Voice Prompts in VibeVoice-Realtime: A Complete Technical Guide
VibeVoice-Realtime encodes a short audio snippet into the model's KV cache via the _create_voice_prompt method in vibevoice/processor/vibevoice_processor.py, allowing you to pre-compute a speaker-specific cache file that seeds the diffusion decoder for consistent voice characteristics during streaming inference.
The microsoft/VibeVoice repository implements voice cloning through a unique voice prompt mechanism that embeds speaker identity directly into the model's key-value cache before text generation begins. To customize voice prompts in VibeVoice-Realtime, you must generate a pre-filled KV cache containing your reference audio tokens, which the streaming inference engine then uses to condition the diffusion-based speech synthesis pipeline.
Understanding the Voice Prompt Architecture
VibeVoice-Realtime processes voice prompts as special token sequences that prime the acoustic diffusion model. The architecture involves three core components working across specific source files.
Voice Prompt Creation via _create_voice_prompt
The foundation of voice customization lies in the _create_voice_prompt method located at vibevoice/processor/vibevoice_processor.py#L400-L460. This private method accepts one or more audio samples, converts them into VAE tokens, and constructs a structured token sequence following the format:
Voice input:
Speaker N: <speech_start> <vae_token…> <speech_end>
The method returns three critical tensors: token IDs, raw audio tensors, and a speech mask indicating which tokens correspond to audio versus text.
Tokenisation and KV Cache Integration
Once created, voice prompts merge with user text through process_input_with_cached_prompt in vibevoice/processor/vibevoice_streaming_processor.py#L170-L220. This method builds a BatchEncoding containing input_ids, attention_mask, and speech_input_mask, preparing the combined representation for the transformer model.
Streaming Inference Consumption
The vibevoice/modular/modeling_vibevoice_streaming_inference.py#L622-L676 file contains the core generation loop. Here, model.generate receives the encoded inputs plus a cached_prompt parameter containing the pre-filled KV cache. The voice prompt tokens seed the acoustic diffusion, ensuring generated speech inherits the reference speaker's timbre during windowed text prefill and incremental generation.
Step-by-Step Guide to Creating a Custom Voice Prompt
To generate a production-ready custom voice prompt, follow this workflow that captures and persists the KV cache for reuse across streaming sessions.
-
Record clean reference audio – Capture 3–5 seconds of high-quality speech for your target speaker at the model's expected sampling rate (typically 16 kHz).
-
Load and process the sample – Use the processor's audio loading utilities to convert your WAV file into the required tensor format.
-
Generate voice tokens – Invoke
_create_voice_promptto build the token sequence, audio tensors, and speech masks from your reference sample. -
Execute a prefill pass – Run
model.generatewithmax_new_tokens=0to process the voice prompt through the transformer layers, capturing the resulting KV cache. -
Persist the cache – Save the output dictionary containing hidden states and cache tensors as a PyTorch checkpoint (
.ptfile). -
Deploy in streaming – Load the saved checkpoint as a
cached_promptin your streaming service configuration.
Code Implementation: Generating a Custom Voice Prompt
The following implementation demonstrates the complete pipeline from audio loading through cache persistence:
import torch
import numpy as np
from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor
from vibevoice.modular.modeling_vibevoice_streaming import VibeVoiceStreamingModel
# 1️⃣ Initialise processor & model (use the streaming checkpoint)
processor = VibeVoiceProcessor.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")
model = VibeVoiceStreamingModel.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")
# 2️⃣ Load your own speaker audio (wav → np.ndarray)
# Here `sample.wav` is a 16 kHz mono file.
wav, sr = processor.audio_processor._load_audio_from_path("sample.wav")
assert sr == processor.speech_sampling_rate
# 3️⃣ Build the voice‑prompt tokens, audio tensors & mask
voice_tokens, voice_inputs, voice_masks = processor._create_voice_prompt([wav])
# 4️⃣ Encode a dummy script (the prompt itself) – we only need the voice part.
# The processor will prepend a system prompt automatically.
dummy_script = " Voice input:\n" # minimal script – actual speech comes from voice_inputs
encoding = processor.encode(
dummy_script,
speech_inputs=voice_inputs,
speech_input_masks=voice_masks,
return_tensors="pt",
)
# 5️⃣ Run a *prefill* pass to capture KV cache
outputs = model.generate(
**encoding,
max_new_tokens=0, # no new text generation yet
return_dict_in_generate=True,
output_hidden_states=True,
)
# 6️⃣ Save the cached KV cache (prefilled prompt) for later streaming use
prefilled_path = "my_custom_prompt.pt"
torch.save(outputs, prefilled_path)
print(f"Custom voice prompt saved to {prefilled_path}")
Deploying Custom Prompts in Streaming Applications
After generating your .pt checkpoint, integrate it into the streaming inference pipeline. The reference implementation in demo/web/app.py#L160-L187 demonstrates loading pre-computed voice caches into a StreamingTTSService instance.
from demo.web.app import StreamingTTSService
service = StreamingTTSService(
model_path="microsoft/VibeVoice-Realtime-0.5B",
voice_key="my_custom", # arbitrary identifier
voice_presets={"my_custom": "my_custom_prompt.pt"},
)
# Now `service.stream("Hello world!")` will synthesize with your voice.
This pattern allows you to maintain a registry of voice presets (stored in demo/voices/streaming_model/ in the reference implementation) that load instantly without recomputing the KV cache for each request.
Key Source Files for Voice Prompt Customization
Understanding these specific files enables advanced customization beyond the standard API:
-
vibevoice/processor/vibevoice_processor.py– Contains_create_voice_promptand audio preprocessing logic for converting raw waveforms into VAE token sequences. -
vibevoice/processor/vibevoice_streaming_processor.py– Implementsprocess_input_with_cached_prompt, which merges live text inputs with pre-computed voice prompt caches. -
vibevoice/modular/modeling_vibevoice_streaming_inference.py– Houses the streaming generation loop that consumes prefixed KV caches and performs windowed diffusion-based speech synthesis. -
demo/web/app.py– Reference Flask-style service demonstrating production deployment patterns for cached voice prompts. -
demo/voices/streaming_model/*.pt– Directory containing shipped example prompts (e.g.,en-Carter_man.pt) that demonstrate the expected checkpoint format.
Summary
- Voice prompts in VibeVoice-Realtime are encoded into the KV cache before streaming begins, not processed in real-time during generation.
- Use
_create_voice_promptinvibevoice_processor.pyto convert reference audio into the required token and mask representation. - Prefill the model with
max_new_tokens=0to capture the voice-specific KV cache, then save it as a.ptcheckpoint for instant loading. - The streaming processor merges user text with cached voice prompts via
process_input_with_cached_prompt, enabling zero-latency voice switching. - Deploy custom voices by referencing saved checkpoints in your
StreamingTTSServiceconfiguration, following the pattern established indemo/web/app.py.
Frequently Asked Questions
How long should a reference audio sample be for optimal voice prompt quality?
A clean recording of 3–5 seconds provides sufficient speaker characteristics for the VAE tokenization pipeline while maintaining manageable cache file sizes. The _create_voice_prompt method processes these samples into compact token sequences that capture timbre and prosodic features without requiring extended utterances.
Can I use multiple speaker samples to create a single voice prompt?
Yes. The _create_voice_prompt method accepts a list of audio arrays as its input parameter, allowing you to concatenate multiple reference utterances from the same speaker. This improves robustness by averaging speaker characteristics across different phonetic contexts, though the method returns unified token and mask tensors representing the combined voice identity.
Why does VibeVoice-Realtime use KV cache prefill instead of processing voice prompts during generation?
The KV cache prefill strategy eliminates inference latency for voice conditioning. By pre-computing the transformer states for voice prompt tokens and persisting them as .pt files, the streaming pipeline in modeling_vibevoice_streaming_inference.py can immediately begin text processing without recomputing speaker embeddings. This architecture enables the real-time performance required for interactive applications while maintaining consistent speaker identity across long-form generation.
Where are the example voice prompt files located in the repository?
Pre-generated voice prompts ship as PyTorch checkpoints in the demo/voices/streaming_model/ directory. These files (such as en-Carter_man.pt) demonstrate the expected output format from the prefill process and can be loaded directly via the voice_presets parameter in StreamingTTSService as shown in demo/web/app.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →