# How `prompt_wav_path` and `reference_wav_path` Affect Voice Cloning Quality in VoxCPM

> Discover how prompt_wav_path and reference_wav_path influence VoxCPM voice cloning quality. Learn how to optimize audio inputs for superior synthesis and style control.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: deep-dive
- Published: 2026-04-10

---

**VoxCPM selects one of three synthesis pipelines based on which audio paths you provide: `reference_wav_path` alone triggers pure voice cloning, `prompt_wav_path` with text enables style-controlled continuation, and combining both yields the highest quality by anchoring timbre and prosody simultaneously.**

The OpenBMB/VoxCPM repository implements a neural codec language model for zero-shot text-to-speech and voice cloning. Understanding how `prompt_wav_path` and `reference_wav_path` interact is essential for controlling synthesis behavior and maximizing output quality, as their presence dictates which inference pipeline executes and how audio features are constructed.

## Three Synthesis Modes Determined by Audio Inputs

VoxCPM does not treat these parameters as simple optional flags; their combination automatically selects the inference strategy.

**Reference-only mode** (`reference_wav_path` alone): The model encodes the reference audio into a voice-style prefix (`[ref_start] … [ref_end]`) that isolates the speaker’s timbre. This produces pure voice cloning without continuation from a prompt. Quality depends on the similarity between the reference and target speaker; longer, clean audio improves timbre capture.

**Continuation mode** (`prompt_wav_path` + `prompt_text`): The prompt audio and its transcript are encoded and placed after the target text, enabling the model to continue speaking in the same prosody and intonation. This provides style cues (speaking rate, emotion) but does not clone a new voice.

**Combined mode** (both paths): The reference prefix is built first, followed by the continuation suffix. This yields voice cloning while respecting the specific speaking style defined by the prompt, typically producing the highest quality output.

## Source Code Implementation

The logic resides in two primary files according to the VoxCPM source code.

In [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py) (lines 162-166), the public API forwards both paths to the underlying model:

```python
def _generate(..., prompt_wav_path: str = None, prompt_text: str = None,
              reference_wav_path: str = None, ...) -> Generator[...]:

```

File validation occurs at lines 191-210, while lines 214-219 enforce the mutual dependency between `prompt_wav_path` and `prompt_text`. Note that reference audio is only allowed for VoxCPM2 models (line 218).

Mode selection and feature construction happen in [`src/voxcpm/model/voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm2.py). The combined mode (lines 74-95) builds a reference prefix and left-pads the prompt feature. Reference-only mode (lines 21-39) creates solely the reference prefix. Continuation-only mode (lines 82-106) right-pads the prompt. The `_encode_wav` function (lines 86-99) loads audio via librosa, applies optional VAD-based silence trimming, and pads waveforms left or right depending on the mode.

## Audio Processing Factors That Impact Quality

Several preprocessing steps in the source code directly affect cloning fidelity.

**Audio length and padding**: `_encode_wav` pads inputs to multiples of `patch_len`. Reference audio uses right-padding (keeping the prefix at the beginning), while prompt audio uses left-padding (aligning with the end of the generated waveform). Very short references may not capture sufficient timbral variation.

**Silence trimming**: When `trim_silence_vad` is enabled (lines 84-88 in [`voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/voxcpm2.py)), the model removes leading and trailing silence before encoding, focusing the embedding on speaker characteristics rather than dead air.

**Noise reduction**: If `denoise=True`, the `Denoiser` runs on the audio before encoding (lines 29-39 in [`core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/core.py)), yielding cleaner embeddings and clearer cloning results.

**Text normalization**: With `normalize=True` (lines 56-62 in [`core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/core.py)), the model standardizes text before tokenization, improving alignment between transcripts and audio prompts.

## Implementation Examples

### Python API

```python
from voxcpm import VoxCPM

# Initialize VoxCPM2 (requires checkpoint)

model = VoxCPM.from_pretrained("openbmb/VoxCPM2")

# Reference-only: Pure voice cloning

audio = model.generate(
    text="Hello, this is a cloned voice.",
    reference_wav_path="examples/reference_speaker.wav",
    cfg_value=2.5,
    normalize=True,
)

# Prompt-only: Style continuation

audio = model.generate(
    text="Now I continue the story.",
    prompt_wav_path="examples/example.wav",
    prompt_text="Earlier I said:",
    cfg_value=2.0,
)

# Combined: Cloning + style control

audio = model.generate(
    text="Finally, I wrap up.",
    prompt_wav_path="examples/example.wav",
    prompt_text="Earlier I said:",
    reference_wav_path="examples/reference_speaker.wav",
    denoise=True,
    normalize=True,
)

# Save output (48 kHz)

import soundfile as sf
sf.write("output.wav", audio, model.tts_model.sample_rate)

```

### Command Line Interface

The CLI in [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py) (lines 45-66) maps arguments to the core API:

```bash

# Reference-only cloning

voxcpm clone \
    --text "This is a cloned voice." \
    --reference-audio examples/reference_speaker.wav \
    --output cloned.wav

# Continuation with style

voxcpm clone \
    --text "Continuing the narrative." \
    --prompt-audio examples/example.wav \
    --prompt-text "Earlier sentence." \
    --output continuation.wav

# Combined mode (recommended)

voxcpm clone \
    --text "Final statement." \
    --prompt-audio examples/example.wav \
    --prompt-text "Earlier content." \
    --reference-audio examples/reference_speaker.wav \
    --denoise \
    --normalize \
    --output combined.wav

```

## Summary

- `prompt_wav_path` provides **style cues** for continuation, while `reference_wav_path` provides **timbre anchors** for voice cloning.
- VoxCPM automatically selects one of three synthesis modes based on which paths you provide, with the combined mode typically yielding the highest quality.
- Quality depends on audio cleanliness, length (≥ 2 seconds recommended), and proper alignment between `prompt_text` and `prompt_wav_path`.
- Enable `--denoise` and `--normalize` to improve embedding quality when source recordings contain noise or inconsistent formatting.
- Reference audio is only supported in VoxCPM2 models, validated in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py) at line 218.

## Frequently Asked Questions

### What is the minimum length for reference audio in VoxCPM?

The source code does not enforce a hard minimum, but the `_encode_wav` function pads audio to multiples of `patch_len`. In practice, provide at least 2 seconds of clean speech to ensure the model captures sufficient timbral variation. Very short clips may result in lower similarity to the target speaker.

### Can I use `reference_wav_path` with VoxCPM1 models?

No. According to the validation logic in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py) (line 218), reference audio is restricted to VoxCPM2 models. Attempting to pass `reference_wav_path` with a VoxCPM1 checkpoint will raise an error or be ignored by the pipeline.

### Why does the prompt audio get left-padded while reference audio gets right-padded?

The padding direction ensures correct temporal alignment during autoregressive generation. Reference audio uses right-padding so the voice-style prefix remains at the beginning of the sequence, influencing the entire output. Prompt audio uses left-padding so it aligns with the end of the generated waveform, acting as a continuation cue rather than a global timbre anchor (see [`src/voxcpm/model/voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm2.py), lines 86-99).

### Does enabling denoising affect inference speed?

Yes. When `denoise=True`, the model invokes the `Denoiser` on the audio before encoding (lines 29-39 in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py)). This adds preprocessing time but typically improves voice cloning quality by removing background noise that could corrupt the speaker embedding.