How `prompt_wav_path` and `reference_wav_path` Affect Voice Cloning Quality in VoxCPM
VoxCPM selects one of three synthesis pipelines based on which audio paths you provide: reference_wav_path alone triggers pure voice cloning, prompt_wav_path with text enables style-controlled continuation, and combining both yields the highest quality by anchoring timbre and prosody simultaneously.
The OpenBMB/VoxCPM repository implements a neural codec language model for zero-shot text-to-speech and voice cloning. Understanding how prompt_wav_path and reference_wav_path interact is essential for controlling synthesis behavior and maximizing output quality, as their presence dictates which inference pipeline executes and how audio features are constructed.
Three Synthesis Modes Determined by Audio Inputs
VoxCPM does not treat these parameters as simple optional flags; their combination automatically selects the inference strategy.
Reference-only mode (reference_wav_path alone): The model encodes the reference audio into a voice-style prefix ([ref_start] … [ref_end]) that isolates the speaker’s timbre. This produces pure voice cloning without continuation from a prompt. Quality depends on the similarity between the reference and target speaker; longer, clean audio improves timbre capture.
Continuation mode (prompt_wav_path + prompt_text): The prompt audio and its transcript are encoded and placed after the target text, enabling the model to continue speaking in the same prosody and intonation. This provides style cues (speaking rate, emotion) but does not clone a new voice.
Combined mode (both paths): The reference prefix is built first, followed by the continuation suffix. This yields voice cloning while respecting the specific speaking style defined by the prompt, typically producing the highest quality output.
Source Code Implementation
The logic resides in two primary files according to the VoxCPM source code.
In src/voxcpm/core.py (lines 162-166), the public API forwards both paths to the underlying model:
def _generate(..., prompt_wav_path: str = None, prompt_text: str = None,
reference_wav_path: str = None, ...) -> Generator[...]:
File validation occurs at lines 191-210, while lines 214-219 enforce the mutual dependency between prompt_wav_path and prompt_text. Note that reference audio is only allowed for VoxCPM2 models (line 218).
Mode selection and feature construction happen in src/voxcpm/model/voxcpm2.py. The combined mode (lines 74-95) builds a reference prefix and left-pads the prompt feature. Reference-only mode (lines 21-39) creates solely the reference prefix. Continuation-only mode (lines 82-106) right-pads the prompt. The _encode_wav function (lines 86-99) loads audio via librosa, applies optional VAD-based silence trimming, and pads waveforms left or right depending on the mode.
Audio Processing Factors That Impact Quality
Several preprocessing steps in the source code directly affect cloning fidelity.
Audio length and padding: _encode_wav pads inputs to multiples of patch_len. Reference audio uses right-padding (keeping the prefix at the beginning), while prompt audio uses left-padding (aligning with the end of the generated waveform). Very short references may not capture sufficient timbral variation.
Silence trimming: When trim_silence_vad is enabled (lines 84-88 in voxcpm2.py), the model removes leading and trailing silence before encoding, focusing the embedding on speaker characteristics rather than dead air.
Noise reduction: If denoise=True, the Denoiser runs on the audio before encoding (lines 29-39 in core.py), yielding cleaner embeddings and clearer cloning results.
Text normalization: With normalize=True (lines 56-62 in core.py), the model standardizes text before tokenization, improving alignment between transcripts and audio prompts.
Implementation Examples
Python API
from voxcpm import VoxCPM
# Initialize VoxCPM2 (requires checkpoint)
model = VoxCPM.from_pretrained("openbmb/VoxCPM2")
# Reference-only: Pure voice cloning
audio = model.generate(
text="Hello, this is a cloned voice.",
reference_wav_path="examples/reference_speaker.wav",
cfg_value=2.5,
normalize=True,
)
# Prompt-only: Style continuation
audio = model.generate(
text="Now I continue the story.",
prompt_wav_path="examples/example.wav",
prompt_text="Earlier I said:",
cfg_value=2.0,
)
# Combined: Cloning + style control
audio = model.generate(
text="Finally, I wrap up.",
prompt_wav_path="examples/example.wav",
prompt_text="Earlier I said:",
reference_wav_path="examples/reference_speaker.wav",
denoise=True,
normalize=True,
)
# Save output (48 kHz)
import soundfile as sf
sf.write("output.wav", audio, model.tts_model.sample_rate)
Command Line Interface
The CLI in src/voxcpm/cli.py (lines 45-66) maps arguments to the core API:
# Reference-only cloning
voxcpm clone \
--text "This is a cloned voice." \
--reference-audio examples/reference_speaker.wav \
--output cloned.wav
# Continuation with style
voxcpm clone \
--text "Continuing the narrative." \
--prompt-audio examples/example.wav \
--prompt-text "Earlier sentence." \
--output continuation.wav
# Combined mode (recommended)
voxcpm clone \
--text "Final statement." \
--prompt-audio examples/example.wav \
--prompt-text "Earlier content." \
--reference-audio examples/reference_speaker.wav \
--denoise \
--normalize \
--output combined.wav
Summary
prompt_wav_pathprovides style cues for continuation, whilereference_wav_pathprovides timbre anchors for voice cloning.- VoxCPM automatically selects one of three synthesis modes based on which paths you provide, with the combined mode typically yielding the highest quality.
- Quality depends on audio cleanliness, length (≥ 2 seconds recommended), and proper alignment between
prompt_textandprompt_wav_path. - Enable
--denoiseand--normalizeto improve embedding quality when source recordings contain noise or inconsistent formatting. - Reference audio is only supported in VoxCPM2 models, validated in
src/voxcpm/core.pyat line 218.
Frequently Asked Questions
What is the minimum length for reference audio in VoxCPM?
The source code does not enforce a hard minimum, but the _encode_wav function pads audio to multiples of patch_len. In practice, provide at least 2 seconds of clean speech to ensure the model captures sufficient timbral variation. Very short clips may result in lower similarity to the target speaker.
Can I use reference_wav_path with VoxCPM1 models?
No. According to the validation logic in src/voxcpm/core.py (line 218), reference audio is restricted to VoxCPM2 models. Attempting to pass reference_wav_path with a VoxCPM1 checkpoint will raise an error or be ignored by the pipeline.
Why does the prompt audio get left-padded while reference audio gets right-padded?
The padding direction ensures correct temporal alignment during autoregressive generation. Reference audio uses right-padding so the voice-style prefix remains at the beginning of the sequence, influencing the entire output. Prompt audio uses left-padding so it aligns with the end of the generated waveform, acting as a continuation cue rather than a global timbre anchor (see src/voxcpm/model/voxcpm2.py, lines 86-99).
Does enabling denoising affect inference speed?
Yes. When denoise=True, the model invokes the Denoiser on the audio before encoding (lines 29-39 in src/voxcpm/core.py). This adds preprocessing time but typically improves voice cloning quality by removing background noise that could corrupt the speaker embedding.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →