How to Maximize Voice Cloning Similarity with Reference and Prompt Audio in VoxCPM

Use VoxCPM's combined (ultimate) cloning mode by providing both reference_wav_path for pure timbre extraction and prompt_wav_path with prompt_text for acoustic continuation, while enabling trim_silence_vad and denoise preprocessing to eliminate noise that degrades the latent representations.

Voice cloning in OpenBMB/VoxCPM achieves maximum similarity when you leverage both reference audio for speaker identity and prompt audio for prosodic continuity. The VoxCPM2 architecture specifically supports a combined mode that isolates timbre from reference audio while continuing the acoustic pattern from prompt audio, a mechanism implemented in VoxCPM2Model._generate within the source repository. This guide explains the technical implementation and configuration required to achieve the highest fidelity voice cloning results.

Understanding VoxCPM's Three Voice Cloning Modes

VoxCPM supports three distinct cloning strategies, each suited for different input scenarios:

Reference-Only Mode

Input: reference_wav_path only

This mode isolates the timbre of the reference speaker and generates speech from your target text. The model encodes the reference audio using _encode_wav with padding_mode="right" and constructs a reference prefix via _make_ref_prefix in src/voxcpm/model/voxcpm2.py (lines 14-44). The special tokens [ref_start] and [ref_end] delimit the reference latent features, instructing the diffusion model to preserve this spectral fingerprint throughout generation.

Continuation-Only Mode

Input: prompt_wav_path + prompt_text

This mode continues the acoustic content of the prompt audio while following your new target text. The prompt is encoded with padding_mode="left" and its latent representation is concatenated after the target text tokens. The model conditions generation on the exact acoustic pattern of the prompt, effectively "copying and extending" the vocal characteristics.

Combined (Ultimate) Clone Mode

Input: reference_wav_path + prompt_wav_path + prompt_text

The combined mode yields the highest similarity because it processes both cues sequentially as implemented in lines 74-119 of src/voxcpm/model/voxcpm2.py:

  1. Encode reference audio (_encode_wav with padding_mode="right")
  2. Encode prompt audio (_encode_wav with padding_mode="left")
  3. Build reference prefix with special tokens [ref_start] ... [ref_end]
  4. Concatenate streams: [ref_prefix] + [target-text tokens] + [prompt-pad-tokens] for text and [ref_features] + [text-pad-features] + [prompt_features] for audio
  5. Create masks that distinguish real audio (reference/prompt) from positions requiring prediction

The reference audio provides a pure timbre cue, while the prompt audio plus transcript provides a content-level cue that preserves vocal nuance, intonation, and rhythm.

The Architecture Behind High-Fidelity Cloning

Several architectural mechanisms in VoxCPM enable precise voice replication:

Reference-Isolation Tokens

The special tokens ref_audio_start_token and ref_audio_end_token delimit the reference latent vector in the diffusion input. During the denoising process, these positions have an audio_mask value of 0, making them fixed anchors that the model cannot modify. This preserves the speaker's spectral fingerprint regardless of generation randomness.

Prompt Continuation Mechanism

The prompt audio latent features are inserted after the target-text tokens but before the diffusion prediction region. With an audio_mask value of 1, these features act as a soft constraint, allowing the model to condition generation on the exact acoustic pattern while still generating novel content for the target text.

Text-Audio Masking Strategy

Separate masks (text_mask, audio_mask) inform the diffusion transformer which positions correspond to text, reference audio, or prompt audio. This precise segregation in src/voxcpm/model/voxcpm2.py prevents cross-talk between the timbre stream (reference) and the content stream (prompt).

AudioVAE-V2 Latent Encoding

Both reference and prompt audio pass through the same AudioVAE-V2 encoder defined in src/voxcpm/modules/audiovae/audio_vae.py. This guarantees that both inputs occupy the identical latent space the diffusion model expects, eliminating domain mismatch that would otherwise reduce similarity.

Configuration Tips for Maximum Similarity

Optimize your inputs using these specific parameters:

  • Match recording conditions – Use identical microphones, distances, and room acoustics for both reference and prompt audio. The encoder is sensitive to spectral characteristics; mismatches inject noise into the timbre cue.

  • Trim silence and noise – Set trim_silence_vad=True or use the --trim-silence CLI flag. This removes non-speech frames that could encode as "silence tokens," improving the purity of the reference prefix.

  • Enable denoising – Load the denoiser via load_denoiser=True when initializing the model, then pass denoise=True to generate(). Cleaner reference and prompt audio yield more accurate latent representations.

  • Provide exact transcripts – Pass the verbatim prompt_text of the prompt audio. The model aligns latent prompt features with textual content, preventing prosodic drift.

  • Limit prompt duration – Keep prompts under 2 seconds. Long prompts increase continuation drift; the architecture optimizes for short-context conditioning.

  • Use 16 kHz WAV inputs – While the model automatically resamples to 16 kHz, providing native 16 kHz PCM WAV files avoids interpolation artifacts.

  • Set modest guidance scale – Use cfg_value=2.0 (default). Higher values over-emphasize text adherence at the expense of timbre fidelity.

Implementation: Code Examples

Python API (Ultimate Clone)

from voxcpm import VoxCPM
import soundfile as sf

# Initialize with denoiser support

model = VoxCPM.from_pretrained(
    "openbmb/VoxCPM2",
    load_denoiser=True,
)

# Audio paths (16 kHz WAV recommended)

reference_path = "samples/reference_speaker.wav"
prompt_path = "samples/prompt_continuation.wav"
prompt_text = "The exact words spoken in prompt_continuation.wav."

# Generate with maximum similarity settings

wav = model.generate(
    text="Now we continue the story with a new sentence.",
    reference_wav_path=reference_path,
    prompt_wav_path=prompt_path,
    prompt_text=prompt_text,
    trim_silence_vad=True,
    denoise=True,
    cfg_value=2.0,
)

sf.write("ultimate_clone.wav", wav, model.tts_model.sample_rate)

The key parameters reference_wav_path, prompt_wav_path, and prompt_text trigger the combined mode in VoxCPMCore.generate (defined in src/voxcpm/core.py lines 10-22).

CLI Approach

voxcpm clone \
  --text "Now we continue the story with a new sentence." \
  --reference-audio samples/reference_speaker.wav \
  --prompt-audio samples/prompt_continuation.wav \
  --prompt-text "The exact words spoken in prompt_continuation.wav." \
  --trim-silence \
  --denoise \
  --cfg 2.0 \
  --output ultimate_clone.wav

The CLI entry point in src/voxcpm/cli.py (lines 259-328) maps these flags directly to the API arguments.

Streaming Generation

chunks = model.generate_streaming(
    text="A streaming version of the same clone.",
    reference_wav_path=reference_path,
    prompt_wav_path=prompt_path,
    prompt_text=prompt_text,
    trim_silence_vad=True,
    denoise=True,
)

import numpy as np
wav = np.concatenate(list(chunks))
sf.write("streaming_clone.wav", wav, model.tts_model.sample_rate)

Note that streaming mode disables retry_badcase functionality as indicated by internal warnings in the source.

Summary

  • Use combined mode by providing both reference_wav_path and prompt_wav_path with prompt_text to leverage both timbre isolation and acoustic continuation.
  • Preprocess audio with trim_silence_vad=True and denoise=True to ensure clean latent representations in the AudioVAE-V2 encoder.
  • Reference audio sets the timbre via fixed tokens with audio_mask=0, while prompt audio guides prosody through the continuation mechanism.
  • Keep prompts short (under 2 seconds) and match recording conditions between reference and prompt for optimal similarity.
  • Set cfg_value=2.0 to balance text adherence with voice fidelity, avoiding the timbre degradation caused by excessive guidance scales.

Frequently Asked Questions

What is the difference between reference audio and prompt audio in VoxCPM?

Reference audio provides the timbre and speaker identity for the clone, while prompt audio provides the prosody, pacing, and acoustic context for continuation. The reference is processed into a fixed prefix that anchors the speaker's voice characteristics, whereas the prompt is concatenated after the target text to guide how the new speech should sound. For maximum similarity, you must provide both, as the reference alone cannot capture specific intonation patterns, and the prompt alone cannot establish a stable speaker identity.

Why does my cloned voice sound different from the reference when using only reference audio?

Using only reference_wav_path isolates timbre but provides no prosodic template for the new text. The diffusion model must infer intonation and rhythm solely from the text tokens, which often results in a "generic" delivery that lacks the specific vocal mannerisms of the speaker. Additionally, without the exact transcript alignment provided by prompt_text, the model cannot synchronize acoustic features with linguistic content, leading to subtle mismatches that reduce perceived similarity.

How long should the prompt audio be for best results?

Keep prompt audio under 2 seconds for optimal continuation quality. The VoxCPM2 architecture is optimized for short-context conditioning; longer prompts increase the risk of prosodic drift where the model fails to smoothly transition from the prompt content to the generated content. Short, clean prompts provide sufficient acoustic context without overwhelming the diffusion process with excess information that must be maintained across long sequences.

Can I use different speakers for the reference and prompt audio?

While technically possible, using different speakers for reference and prompt will degrade cloning quality. The model expects both audio inputs to originate from the same speaker, as the architecture assumes the prompt provides a continuation of the reference voice's acoustic characteristics. Mixing speakers causes the diffusion model to receive conflicting signals—the reference demands one spectral signature while the prompt demands another—resulting in muddy, inconsistent output that resembles neither speaker accurately.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →