# How Much Data Is Needed for Effective VoxCPM Fine-Tuning?

> Discover how little data is needed for effective VoxCPM fine-tuning. Achieve high-quality speaker adaptation with as little as 5-10 minutes of clean audio.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: performance
- Published: 2026-04-10

---

**As little as 5–10 minutes of clean, single-speaker audio is sufficient to achieve high-quality speaker adaptation with VoxCPM.**

VoxCPM is a **tokenizer-free diffusion-autoregressive TTS model** developed by OpenBMB that leverages a universal latent representation across 30 languages. Because the pretrained weights already encode robust acoustic modeling, language understanding, and high-fidelity audio decoding, **VoxCPM fine-tuning** requires only a tiny amount of speaker-specific data to specialize the model for new voices or domain variations.

## Why VoxCPM Needs So Little Data

VoxCPM’s architecture delegates the heavy lifting to pretrained components, leaving only lightweight speaker-style adjustments to be learned during fine-tuning. The model processes speech through four discrete stages operating in the latent space of an **AudioVAE V2**: Local Encoder (LocEnc) → Text-Supervised Language Model (TSLM) → Retrieval-Augmented Language Model (RALM) → Local Diffusion Transformer (LocDiT).

This design means:
- The **AudioVAE V2** latent encoder already extracts robust acoustic features from raw audio.
- The **MiniCPM-4** backbone provides strong linguistic grounding across languages.
- Fine-tuning only needs to teach **speaker-style embeddings** and minor prosodic adjustments, which can be learned from a few short utterances.

According to the [official README](https://github.com/OpenBMB/VoxCPM/blob/main/README.md#L493), this architectural efficiency is why the model achieves noticeable speaker adaptation with minimal data.

## Exact Data Requirements

### Minimum Recommendations

For effective **VoxCPM fine-tuning**, start with **5 minutes of audio** (approximately 100 utterances) for basic speaker adaptation. If the target voice exhibits strong expressive or tonal variation, extending to **10 minutes** (approximately 200 utterances) yields more stable style control.

The dataset should be formatted as a **JSONL** manifest where each line contains the audio path and transcript. Optional fields like `duration` and `dataset_id` accelerate preprocessing and enable multi-dataset training, as implemented in [`src/voxcpm/training/data.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/training/data.py).

### SFT vs. LoRA Considerations

VoxCPM supports two fine-tuning modes with different data appetite:

- **LoRA (Low-Rank Adaptation)** – Adds low-rank adapters (default rank = 4) to attention layers. This parameter-efficient approach requires **only 5 minutes** of audio and converges in fewer steps while keeping the base model weights frozen. Use this for most voice cloning tasks.
- **Full SFT (Supervised Fine-Tuning)** – Rewrites all model weights. This approach benefits from **10 minutes or more** of data when you need maximum quality or want to modify the entire acoustic pipeline rather than just speaker characteristics.

## Preparing Your Training Dataset

### JSONL Manifest Format

Create a JSONL file where each line represents one training sample. The [`src/voxcpm/training/data.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/training/data.py) loader expects the following structure:

```json
{"audio": "examples/example.wav", "text": "This is an example audio transcript for training."}
{"audio": "data/audio1.wav", "text": "Hello, I am a new speaker.", "duration": 2.3}
{"audio": "data/audio2.wav", "text": "Welcome to my voice demo.", "duration": 1.8, "dataset_id": "speaker_001"}

```

Including the `duration` field (in seconds) is optional but speeds up dataset initialization by avoiding redundant audio length checks.

### Data Quality Guidelines

- Use **clean, single-speaker recordings** without background noise or cross-talk.
- Ensure transcripts are exact matches of the spoken content.
- For multi-dataset experiments, utilize the `dataset_id` field to tag different speakers or domains.

## Running the Fine-Tuning

### LoRA Configuration (5 Minutes)

For most use cases, LoRA fine-tuning provides the best efficiency-to-quality ratio. Execute training using the provided configuration:

```bash
python scripts/train_voxcpm_finetune.py \
    --config_path conf/voxcpm_v2/voxcpm_finetune_lora.yaml

```

The [`conf/voxcpm_v2/voxcpm_finetune_lora.yaml`](https://github.com/OpenBMB/VoxCPM/blob/main/conf/voxcpm_v2/voxcpm_finetune_lora.yaml) file specifies data paths, optimizer settings, and checkpoint logging directories.

### Full-Model SFT (10+ Minutes)

If you have collected approximately 10 minutes of data and require deeper acoustic modification:

```bash
python scripts/train_voxcpm_finetune.py \
    --config_path conf/voxcpm_v2/voxcpm_finetune_all.yaml

```

## Inference with Fine-Tuned Weights

After training completes, load the base model with your adapted weights for generation:

```python
from voxcpm import VoxCPM

# Load base model with LoRA weights

model = VoxCPM.from_pretrained(
    "openbmb/VoxCPM2",
    lora_weights_path="checkpoints/finetune_lora",
    load_denoiser=False,
)

# Generate speech

wav = model.generate(
    text="(A calm female voice)Hello, this is my newly fine-tuned voice.",
    cfg_value=2.0,
    inference_timesteps=10,
)

```

## Summary

- **VoxCPM fine-tuning** requires only **5–10 minutes** of speaker-specific audio thanks to its pretrained AudioVAE V2 and MiniCPM-4 backbone.
- Use **LoRA** for efficient adaptation with 5 minutes of data; use **full SFT** when you have 10+ minutes and need maximum acoustic control.
- Prepare data as a **JSONL** manifest with `audio` and `text` fields, optionally including `duration` for faster loading.
- Training entry point is [`scripts/train_voxcpm_finetune.py`](https://github.com/OpenBMB/VoxCPM/blob/main/scripts/train_voxcpm_finetune.py) with configurations in `conf/voxcpm_v2/`.

## Frequently Asked Questions

### Can I fine-tune VoxCPM with less than 5 minutes of audio?

While the official recommendation is 5–10 minutes, you may experiment with as little as 2–3 minutes for simple voice cloning tasks, though quality and stability may degrade. The model’s strong zero-shot capabilities allow some adaptation with minimal data, but 5 minutes provides a reliable baseline for consistent speaker similarity.

### What audio format should I use for VoxCPM training?

The [`src/voxcpm/training/data.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/training/data.py) loader supports standard audio formats including WAV and MP3. Ensure your audio files are sampled at the rate expected by the AudioVAE V2 (typically 24kHz or 16kHz depending on the specific checkpoint). Including the `duration` field in your JSONL manifest helps the data loader verify compatibility without loading entire files.

### How do I know if I need full SFT or LoRA fine-tuning?

Choose **LoRA** when you want to clone a speaker’s voice while preserving the base model’s general capabilities and minimizing storage overhead (LoRA adapters are typically only a few megabytes). Choose **full SFT** when you need to modify deep acoustic properties, adapt to entirely new languages not well-supported in the base model, or when you have sufficient compute and data (10+ minutes) to justify updating all 4 billion parameters.

### Will fine-tuning destroy the model's multilingual capabilities?

No. Because **LoRA** fine-tuning adds small adapters while keeping the base MiniCPM-4 and AudioVAE V2 weights frozen, the model retains its multilingual synthesis capabilities. Even with **full SFT**, catastrophic forgetting is minimal due to the architecture’s robust pretrained representations, though you should verify generation quality in non-target languages if multilingual output remains a requirement.