How to Fine-Tune VoxCPM for Custom Speaker Adaptation Using LoRA: A Complete Guide
VoxCPM supports efficient custom speaker adaptation through Low-Rank Adaptation (LoRA), which injects trainable low-rank matrices into specific linear layers while keeping pretrained weights frozen, enabling fast fine-tuning with minimal memory overhead.
Fine-tuning VoxCPM for custom speaker adaptation using LoRA allows you to clone voices with only minutes of audio data and a consumer GPU. The OpenBMB/VoxCPM repository implements LoRA across the language model, diffusion transformer, and projection layers, offering a parameter-efficient way to learn new speaker characteristics without catastrophic forgetting.
Where LoRA Gets Injected in the VoxCPM Architecture
VoxCPM applies LoRA through targeted injection points across three core architectural components. According to the source code in src/voxcpm/modules/layers/lora.py, the LoRALinear class wraps existing linear layers with trainable low-rank matrices (A and B) while maintaining the frozen base weights.
Base Language Model and Residual Acoustic LM
The MiniCPMModel (base language model) and its residual acoustic counterpart receive LoRA injections in their attention layers. The apply_lora_to_named_linear_modules utility targets q_proj, k_proj, v_proj, and o_proj matrices within the transformer blocks:
- LM attention projections:
q_proj,k_proj,v_proj,o_proj - Residual acoustic LM: Same target modules as the base LM
- Implementation:
src/voxcpm/modules/layers/lora.pycontains theLoRALinearwrapper that injects trainable parameters without modifying the frozen base weights
Diffusion Transformer (DiT) Layers
The VoxCPMLocDiT component, which generates audio feature patches conditioned on language model hidden states, also supports LoRA adaptation. Target modules include the DiT's self-attention projections (q_proj, k_proj, v_proj, o_proj), allowing the diffusion process to adapt to speaker-specific acoustic patterns.
Projection Layers
Explicit projection mappings receive direct LoRALinear wrapping rather than pattern-based matching:
enc_to_lm_proj: Maps encoder outputs to language model spacelm_to_dit_proj: Bridges language model and diffusion transformerres_to_dit_proj: Connects residual acoustic features to diffusion space
The LoRAConfig Object
Configuration is centralized in the LoRAConfig class defined in src/voxcpm/model/voxcpm.py (lines 86-100). When passed to VoxCPMModel.from_local(..., lora_config=LoRAConfig(...)), the model's _apply_lora method executes three sequential steps:
- Wraps LM and residual LM linear layers matching
target_modules_lm - Applies LoRA to DiT layers matching
target_modules_dit - Wraps projection layers specified in
target_proj_modules
All LoRA layers share a unified scaling buffer (self.scaling) that enables on-the-fly activation without torch.compile recompilation.
Preparing Your Dataset and Configuration
Successful fine-tuning requires properly formatted data and a YAML configuration file that specifies which components to adapt and with what hyper-parameters.
Dataset Format Requirements
Create a JSONL manifest where each line contains a text-audio pair. The training script expects entries with "text" (transcription) and "audio" (path to WAV file) keys:
{"text": "Welcome to the custom voice demonstration.", "audio": "/data/speaker_001/welcome.wav"}
{"text": "This is a sample for fine-tuning.", "audio": "/data/speaker_001/sample.wav"}
Reference the example at examples/train_data_example.jsonl for the exact schema expected by the dataloader.
YAML Configuration Structure
Copy the template from conf/voxcpm_v1/voxcpm_finetune_lora.yaml and modify paths and LoRA hyper-parameters:
pretrained_path: /path/to/pretrained_voxcpm_v1
train_manifest: /path/to/train.jsonl
val_manifest: /path/to/val.jsonl
lora:
enable_lm: true
enable_dit: true
enable_proj: true
r: 8
alpha: 16
dropout: 0.0
target_modules_lm: ["q_proj", "v_proj", "k_proj", "o_proj"]
target_modules_dit: ["q_proj", "v_proj", "k_proj", "o_proj"]
target_proj_modules: ["enc_to_lm_proj", "lm_to_dit_proj", "res_to_dit_proj"]
Key parameters:
- r: Rank of the low-rank matrices (typically 4-16)
- alpha: Scaling factor (often set to 2×r)
- dropout: Regularization for LoRA layers (0.0-0.1)
- enable_lm/dit/proj: Boolean flags to selectively freeze or adapt components
Running the Fine-Tuning Process
The training entry point handles configuration parsing, LoRA injection, and checkpoint management automatically.
Command-Line Training
Execute the training script located at scripts/train_voxcpm_finetune.py:
python -m scripts.train_voxcpm_finetune \
--config_path conf/voxcpm_v1/voxcpm_finetune_lora.yaml
For multi-GPU setups or specific device selection:
CUDA_VISIBLE_DEVICES=0 python -m scripts.train_voxcpm_finetune \
--config_path my_speaker_config.yaml \
--tensorboard logs/custom_speaker
The script constructs a LoRAConfig object from the YAML (lines 44-66 in train_voxcpm_finetune.py), loads the pretrained checkpoint, and initializes training with only LoRA parameters set to trainable.
Checkpoint Structure
During training, only LoRA weights are saved, resulting in compact checkpoints:
- lora_weights.safetensors (or
.ckpt): Contains onlylora_Aandlora_Bmatrices - lora_config.json: Stores hyper-parameters and base model reference path
This approach keeps checkpoints small (typically <10 MB) and ensures the base model remains untouched. The saving logic is implemented in scripts/train_voxcpm_finetune.py (lines 71-87).
Loading and Inference with LoRA Weights
After fine-tuning, load the base model and merge the speaker-specific LoRA weights for generation.
Python API Implementation
Instantiate the base model, inject LoRA layers, and load the fine-tuned weights:
import torch
from voxcpm.model.voxcpm import VoxCPMModel, LoRAConfig
# Load the pretrained base model (frozen)
model = VoxCPMModel.from_local(
path="pretrained/voxcpm_v1",
optimize=False,
training=False,
)
# Configure LoRA exactly as during training
lora_cfg = LoRAConfig(
enable_lm=True,
enable_dit=True,
enable_proj=True,
r=8,
alpha=16,
dropout=0.0,
target_modules_lm=["q_proj", "v_proj", "k_proj", "o_proj"],
target_modules_dit=["q_proj", "v_proj", "k_proj", "o_proj"],
target_proj_modules=["enc_to_lm_proj", "lm_to_dit_proj", "res_to_dit_proj"],
)
# Inject LoRA layers and load weights
model.lora_config = lora_cfg
model._apply_lora()
loaded, skipped = model.load_lora_weights("checkpoints/latest")
print(f"Loaded {len(loaded)} LoRA layers, skipped {len(skipped)}")
# Enable LoRA scaling (alpha/r)
model.set_lora_enabled(True)
# Generate speech
text = "This is the adapted speaker voice."
audio = next(model.generate(text, inference_timesteps=20, cfg_value=2.5))
# Save output
import soundfile as sf
sf.write("adapted_output.wav", audio.squeeze(0).cpu().numpy(),
samplerate=model.sample_rate)
The load_lora_weights method (lines 265-293 in voxcpm/model/voxcpm.py) handles both SafeTensors and legacy PyTorch checkpoint formats.
Runtime Toggling
VoxCPM allows switching between base and adapted voices without reloading the model:
# Use base pretrained voice
model.set_lora_enabled(False)
audio_base = next(model.generate("Base voice test"))
# Switch to custom speaker
model.set_lora_enabled(True)
audio_adapted = next(model.generate("Custom speaker test"))
This toggling updates the scaling buffer for all LoRALinear layers instantaneously, as implemented in src/voxcpm/modules/layers/lora.py (lines 73-78).
Summary
- LoRA injection targets attention projections (
q_proj,k_proj,v_proj,o_proj) in the language model and DiT, plus explicit projection layers (enc_to_lm_proj,lm_to_dit_proj,res_to_dit_proj). - Configuration uses the
LoRAConfigclass with parametersr,alpha, anddropoutto control adaptation capacity. - Training requires only LoRA weights to train, keeping checkpoints small (<10 MB) and preventing catastrophic forgetting of base TTS capabilities.
- Inference involves loading the base model, calling
_apply_lora(), loading weights withload_lora_weights(), and enabling adaptation viaset_lora_enabled(True). - Runtime switching allows toggling between base and adapted speakers using the scaling buffer without model recompilation.
Frequently Asked Questions
What rank (r) should I use for speaker adaptation?
Start with r=8 and alpha=16. This provides sufficient capacity to capture speaker characteristics while keeping parameter counts low (approximately 0.1% of base model parameters). For voices with heavy accents or distinctive prosody, increase to r=16 or r=32, though higher ranks increase memory requirements linearly.
Can I fine-tune only the projection layers without the language model?
Yes. Set enable_lm: false and enable_dit: false in your YAML configuration while keeping enable_proj: true. This adapts only enc_to_lm_proj, lm_to_dit_proj, and res_to_dit_proj, requiring even less memory but offering less adaptation capacity for speaker-specific nuances. This mode is useful for quick timbre adjustments when the base model already handles your language well.
Why are my LoRA weights not loading after torch.compile?
The load_lora_weights method in voxcpm/model/voxcpm.py is designed to work with compiled models by handling non-persistent buffers correctly. Ensure you call model._apply_lora() before torch.compile, or set optimize=False during loading. The scaling buffer updates via set_lora_enabled are compilation-safe and do not trigger recompilation.
How do I combine multiple speaker LoRA checkpoints?
Load the base model once, then sequentially call load_lora_weights() for each speaker checkpoint. However, VoxCPM currently supports only one active LoRA configuration at a time; you must toggle between them using set_lora_enabled(True/False). For true multi-speaker merging (combining multiple LoRAs simultaneously), you would need to manually merge the lora_A and lora_B matrices before loading, as the standard API treats each checkpoint as a complete replacement.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →