How Much Data Is Needed for Effective VoxCPM Fine-Tuning?

As little as 5–10 minutes of clean, single-speaker audio is sufficient to achieve high-quality speaker adaptation with VoxCPM.

VoxCPM is a tokenizer-free diffusion-autoregressive TTS model developed by OpenBMB that leverages a universal latent representation across 30 languages. Because the pretrained weights already encode robust acoustic modeling, language understanding, and high-fidelity audio decoding, VoxCPM fine-tuning requires only a tiny amount of speaker-specific data to specialize the model for new voices or domain variations.

Why VoxCPM Needs So Little Data

VoxCPM’s architecture delegates the heavy lifting to pretrained components, leaving only lightweight speaker-style adjustments to be learned during fine-tuning. The model processes speech through four discrete stages operating in the latent space of an AudioVAE V2: Local Encoder (LocEnc) → Text-Supervised Language Model (TSLM) → Retrieval-Augmented Language Model (RALM) → Local Diffusion Transformer (LocDiT).

This design means:

  • The AudioVAE V2 latent encoder already extracts robust acoustic features from raw audio.
  • The MiniCPM-4 backbone provides strong linguistic grounding across languages.
  • Fine-tuning only needs to teach speaker-style embeddings and minor prosodic adjustments, which can be learned from a few short utterances.

According to the official README, this architectural efficiency is why the model achieves noticeable speaker adaptation with minimal data.

Exact Data Requirements

Minimum Recommendations

For effective VoxCPM fine-tuning, start with 5 minutes of audio (approximately 100 utterances) for basic speaker adaptation. If the target voice exhibits strong expressive or tonal variation, extending to 10 minutes (approximately 200 utterances) yields more stable style control.

The dataset should be formatted as a JSONL manifest where each line contains the audio path and transcript. Optional fields like duration and dataset_id accelerate preprocessing and enable multi-dataset training, as implemented in src/voxcpm/training/data.py.

SFT vs. LoRA Considerations

VoxCPM supports two fine-tuning modes with different data appetite:

  • LoRA (Low-Rank Adaptation) – Adds low-rank adapters (default rank = 4) to attention layers. This parameter-efficient approach requires only 5 minutes of audio and converges in fewer steps while keeping the base model weights frozen. Use this for most voice cloning tasks.
  • Full SFT (Supervised Fine-Tuning) – Rewrites all model weights. This approach benefits from 10 minutes or more of data when you need maximum quality or want to modify the entire acoustic pipeline rather than just speaker characteristics.

Preparing Your Training Dataset

JSONL Manifest Format

Create a JSONL file where each line represents one training sample. The src/voxcpm/training/data.py loader expects the following structure:

{"audio": "examples/example.wav", "text": "This is an example audio transcript for training."}
{"audio": "data/audio1.wav", "text": "Hello, I am a new speaker.", "duration": 2.3}
{"audio": "data/audio2.wav", "text": "Welcome to my voice demo.", "duration": 1.8, "dataset_id": "speaker_001"}

Including the duration field (in seconds) is optional but speeds up dataset initialization by avoiding redundant audio length checks.

Data Quality Guidelines

  • Use clean, single-speaker recordings without background noise or cross-talk.
  • Ensure transcripts are exact matches of the spoken content.
  • For multi-dataset experiments, utilize the dataset_id field to tag different speakers or domains.

Running the Fine-Tuning

LoRA Configuration (5 Minutes)

For most use cases, LoRA fine-tuning provides the best efficiency-to-quality ratio. Execute training using the provided configuration:

python scripts/train_voxcpm_finetune.py \
    --config_path conf/voxcpm_v2/voxcpm_finetune_lora.yaml

The conf/voxcpm_v2/voxcpm_finetune_lora.yaml file specifies data paths, optimizer settings, and checkpoint logging directories.

Full-Model SFT (10+ Minutes)

If you have collected approximately 10 minutes of data and require deeper acoustic modification:

python scripts/train_voxcpm_finetune.py \
    --config_path conf/voxcpm_v2/voxcpm_finetune_all.yaml

Inference with Fine-Tuned Weights

After training completes, load the base model with your adapted weights for generation:

from voxcpm import VoxCPM

# Load base model with LoRA weights

model = VoxCPM.from_pretrained(
    "openbmb/VoxCPM2",
    lora_weights_path="checkpoints/finetune_lora",
    load_denoiser=False,
)

# Generate speech

wav = model.generate(
    text="(A calm female voice)Hello, this is my newly fine-tuned voice.",
    cfg_value=2.0,
    inference_timesteps=10,
)

Summary

  • VoxCPM fine-tuning requires only 5–10 minutes of speaker-specific audio thanks to its pretrained AudioVAE V2 and MiniCPM-4 backbone.
  • Use LoRA for efficient adaptation with 5 minutes of data; use full SFT when you have 10+ minutes and need maximum acoustic control.
  • Prepare data as a JSONL manifest with audio and text fields, optionally including duration for faster loading.
  • Training entry point is scripts/train_voxcpm_finetune.py with configurations in conf/voxcpm_v2/.

Frequently Asked Questions

Can I fine-tune VoxCPM with less than 5 minutes of audio?

While the official recommendation is 5–10 minutes, you may experiment with as little as 2–3 minutes for simple voice cloning tasks, though quality and stability may degrade. The model’s strong zero-shot capabilities allow some adaptation with minimal data, but 5 minutes provides a reliable baseline for consistent speaker similarity.

What audio format should I use for VoxCPM training?

The src/voxcpm/training/data.py loader supports standard audio formats including WAV and MP3. Ensure your audio files are sampled at the rate expected by the AudioVAE V2 (typically 24kHz or 16kHz depending on the specific checkpoint). Including the duration field in your JSONL manifest helps the data loader verify compatibility without loading entire files.

How do I know if I need full SFT or LoRA fine-tuning?

Choose LoRA when you want to clone a speaker’s voice while preserving the base model’s general capabilities and minimizing storage overhead (LoRA adapters are typically only a few megabytes). Choose full SFT when you need to modify deep acoustic properties, adapt to entirely new languages not well-supported in the base model, or when you have sufficient compute and data (10+ minutes) to justify updating all 4 billion parameters.

Will fine-tuning destroy the model's multilingual capabilities?

No. Because LoRA fine-tuning adds small adapters while keeping the base MiniCPM-4 and AudioVAE V2 weights frozen, the model retains its multilingual synthesis capabilities. Even with full SFT, catastrophic forgetting is minimal due to the architecture’s robust pretrained representations, though you should verify generation quality in non-target languages if multilingual output remains a requirement.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →