# How to Fine-Tune VoxCPM for Custom Speaker Adaptation Using LoRA: A Complete Guide

> Master VoxCPM custom speaker adaptation with LoRA. This guide shows fast, memory-efficient fine-tuning by training only low-rank matrices to adapt your voice models.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: how-to-guide
- Published: 2026-04-10

---

**VoxCPM supports efficient custom speaker adaptation through Low-Rank Adaptation (LoRA), which injects trainable low-rank matrices into specific linear layers while keeping pretrained weights frozen, enabling fast fine-tuning with minimal memory overhead.**

Fine-tuning VoxCPM for custom speaker adaptation using LoRA allows you to clone voices with only minutes of audio data and a consumer GPU. The OpenBMB/VoxCPM repository implements LoRA across the language model, diffusion transformer, and projection layers, offering a parameter-efficient way to learn new speaker characteristics without catastrophic forgetting.

## Where LoRA Gets Injected in the VoxCPM Architecture

VoxCPM applies LoRA through targeted injection points across three core architectural components. According to the source code in [`src/voxcpm/modules/layers/lora.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/modules/layers/lora.py), the `LoRALinear` class wraps existing linear layers with trainable low-rank matrices (A and B) while maintaining the frozen base weights.

### Base Language Model and Residual Acoustic LM

The **MiniCPMModel** (base language model) and its residual acoustic counterpart receive LoRA injections in their attention layers. The `apply_lora_to_named_linear_modules` utility targets `q_proj`, `k_proj`, `v_proj`, and `o_proj` matrices within the transformer blocks:

- **LM attention projections**: `q_proj`, `k_proj`, `v_proj`, `o_proj`
- **Residual acoustic LM**: Same target modules as the base LM
- **Implementation**: [`src/voxcpm/modules/layers/lora.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/modules/layers/lora.py) contains the `LoRALinear` wrapper that injects trainable parameters without modifying the frozen base weights

### Diffusion Transformer (DiT) Layers

The **VoxCPMLocDiT** component, which generates audio feature patches conditioned on language model hidden states, also supports LoRA adaptation. Target modules include the DiT's self-attention projections (`q_proj`, `k_proj`, `v_proj`, `o_proj`), allowing the diffusion process to adapt to speaker-specific acoustic patterns.

### Projection Layers

Explicit projection mappings receive direct `LoRALinear` wrapping rather than pattern-based matching:
- `enc_to_lm_proj`: Maps encoder outputs to language model space
- `lm_to_dit_proj`: Bridges language model and diffusion transformer
- `res_to_dit_proj`: Connects residual acoustic features to diffusion space

### The LoRAConfig Object

Configuration is centralized in the `LoRAConfig` class defined in [`src/voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm.py) (lines 86-100). When passed to `VoxCPMModel.from_local(..., lora_config=LoRAConfig(...))`, the model's `_apply_lora` method executes three sequential steps:
1. Wraps LM and residual LM linear layers matching `target_modules_lm`
2. Applies LoRA to DiT layers matching `target_modules_dit`
3. Wraps projection layers specified in `target_proj_modules`

All LoRA layers share a unified scaling buffer (`self.scaling`) that enables on-the-fly activation without `torch.compile` recompilation.

## Preparing Your Dataset and Configuration

Successful fine-tuning requires properly formatted data and a YAML configuration file that specifies which components to adapt and with what hyper-parameters.

### Dataset Format Requirements

Create a JSONL manifest where each line contains a text-audio pair. The training script expects entries with `"text"` (transcription) and `"audio"` (path to WAV file) keys:

```json
{"text": "Welcome to the custom voice demonstration.", "audio": "/data/speaker_001/welcome.wav"}
{"text": "This is a sample for fine-tuning.", "audio": "/data/speaker_001/sample.wav"}

```

Reference the example at `examples/train_data_example.jsonl` for the exact schema expected by the dataloader.

### YAML Configuration Structure

Copy the template from [`conf/voxcpm_v1/voxcpm_finetune_lora.yaml`](https://github.com/OpenBMB/VoxCPM/blob/main/conf/voxcpm_v1/voxcpm_finetune_lora.yaml) and modify paths and LoRA hyper-parameters:

```yaml
pretrained_path: /path/to/pretrained_voxcpm_v1
train_manifest: /path/to/train.jsonl
val_manifest: /path/to/val.jsonl

lora:
  enable_lm: true
  enable_dit: true
  enable_proj: true
  r: 8
  alpha: 16
  dropout: 0.0
  target_modules_lm: ["q_proj", "v_proj", "k_proj", "o_proj"]
  target_modules_dit: ["q_proj", "v_proj", "k_proj", "o_proj"]
  target_proj_modules: ["enc_to_lm_proj", "lm_to_dit_proj", "res_to_dit_proj"]

```

**Key parameters:**
- **r**: Rank of the low-rank matrices (typically 4-16)
- **alpha**: Scaling factor (often set to 2×r)
- **dropout**: Regularization for LoRA layers (0.0-0.1)
- **enable_lm/dit/proj**: Boolean flags to selectively freeze or adapt components

## Running the Fine-Tuning Process

The training entry point handles configuration parsing, LoRA injection, and checkpoint management automatically.

### Command-Line Training

Execute the training script located at [`scripts/train_voxcpm_finetune.py`](https://github.com/OpenBMB/VoxCPM/blob/main/scripts/train_voxcpm_finetune.py):

```bash
python -m scripts.train_voxcpm_finetune \
    --config_path conf/voxcpm_v1/voxcpm_finetune_lora.yaml

```

For multi-GPU setups or specific device selection:

```bash
CUDA_VISIBLE_DEVICES=0 python -m scripts.train_voxcpm_finetune \
    --config_path my_speaker_config.yaml \
    --tensorboard logs/custom_speaker

```

The script constructs a `LoRAConfig` object from the YAML (lines 44-66 in [`train_voxcpm_finetune.py`](https://github.com/OpenBMB/VoxCPM/blob/main/train_voxcpm_finetune.py)), loads the pretrained checkpoint, and initializes training with only LoRA parameters set to trainable.

### Checkpoint Structure

During training, only LoRA weights are saved, resulting in compact checkpoints:
- **lora_weights.safetensors** (or `.ckpt`): Contains only `lora_A` and `lora_B` matrices
- **lora_config.json**: Stores hyper-parameters and base model reference path

This approach keeps checkpoints small (typically <10 MB) and ensures the base model remains untouched. The saving logic is implemented in [`scripts/train_voxcpm_finetune.py`](https://github.com/OpenBMB/VoxCPM/blob/main/scripts/train_voxcpm_finetune.py) (lines 71-87).

## Loading and Inference with LoRA Weights

After fine-tuning, load the base model and merge the speaker-specific LoRA weights for generation.

### Python API Implementation

Instantiate the base model, inject LoRA layers, and load the fine-tuned weights:

```python
import torch
from voxcpm.model.voxcpm import VoxCPMModel, LoRAConfig

# Load the pretrained base model (frozen)

model = VoxCPMModel.from_local(
    path="pretrained/voxcpm_v1",
    optimize=False,
    training=False,
)

# Configure LoRA exactly as during training

lora_cfg = LoRAConfig(
    enable_lm=True,
    enable_dit=True,
    enable_proj=True,
    r=8,
    alpha=16,
    dropout=0.0,
    target_modules_lm=["q_proj", "v_proj", "k_proj", "o_proj"],
    target_modules_dit=["q_proj", "v_proj", "k_proj", "o_proj"],
    target_proj_modules=["enc_to_lm_proj", "lm_to_dit_proj", "res_to_dit_proj"],
)

# Inject LoRA layers and load weights

model.lora_config = lora_cfg
model._apply_lora()
loaded, skipped = model.load_lora_weights("checkpoints/latest")
print(f"Loaded {len(loaded)} LoRA layers, skipped {len(skipped)}")

# Enable LoRA scaling (alpha/r)

model.set_lora_enabled(True)

# Generate speech

text = "This is the adapted speaker voice."
audio = next(model.generate(text, inference_timesteps=20, cfg_value=2.5))

# Save output

import soundfile as sf
sf.write("adapted_output.wav", audio.squeeze(0).cpu().numpy(), 
         samplerate=model.sample_rate)

```

The `load_lora_weights` method (lines 265-293 in [`voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/voxcpm/model/voxcpm.py)) handles both SafeTensors and legacy PyTorch checkpoint formats.

### Runtime Toggling

VoxCPM allows switching between base and adapted voices without reloading the model:

```python

# Use base pretrained voice

model.set_lora_enabled(False)
audio_base = next(model.generate("Base voice test"))

# Switch to custom speaker

model.set_lora_enabled(True)
audio_adapted = next(model.generate("Custom speaker test"))

```

This toggling updates the scaling buffer for all `LoRALinear` layers instantaneously, as implemented in [`src/voxcpm/modules/layers/lora.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/modules/layers/lora.py) (lines 73-78).

## Summary

- **LoRA injection** targets attention projections (`q_proj`, `k_proj`, `v_proj`, `o_proj`) in the language model and DiT, plus explicit projection layers (`enc_to_lm_proj`, `lm_to_dit_proj`, `res_to_dit_proj`).
- **Configuration** uses the `LoRAConfig` class with parameters `r`, `alpha`, and `dropout` to control adaptation capacity.
- **Training** requires only LoRA weights to train, keeping checkpoints small (<10 MB) and preventing catastrophic forgetting of base TTS capabilities.
- **Inference** involves loading the base model, calling `_apply_lora()`, loading weights with `load_lora_weights()`, and enabling adaptation via `set_lora_enabled(True)`.
- **Runtime switching** allows toggling between base and adapted speakers using the scaling buffer without model recompilation.

## Frequently Asked Questions

### What rank (r) should I use for speaker adaptation?

Start with **r=8** and **alpha=16**. This provides sufficient capacity to capture speaker characteristics while keeping parameter counts low (approximately 0.1% of base model parameters). For voices with heavy accents or distinctive prosody, increase to **r=16** or **r=32**, though higher ranks increase memory requirements linearly.

### Can I fine-tune only the projection layers without the language model?

Yes. Set `enable_lm: false` and `enable_dit: false` in your YAML configuration while keeping `enable_proj: true`. This adapts only `enc_to_lm_proj`, `lm_to_dit_proj`, and `res_to_dit_proj`, requiring even less memory but offering less adaptation capacity for speaker-specific nuances. This mode is useful for quick timbre adjustments when the base model already handles your language well.

### Why are my LoRA weights not loading after torch.compile?

The `load_lora_weights` method in [`voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/voxcpm/model/voxcpm.py) is designed to work with compiled models by handling non-persistent buffers correctly. Ensure you call `model._apply_lora()` before `torch.compile`, or set `optimize=False` during loading. The scaling buffer updates via `set_lora_enabled` are compilation-safe and do not trigger recompilation.

### How do I combine multiple speaker LoRA checkpoints?

Load the base model once, then sequentially call `load_lora_weights()` for each speaker checkpoint. However, VoxCPM currently supports only one active LoRA configuration at a time; you must toggle between them using `set_lora_enabled(True/False)`. For true multi-speaker merging (combining multiple LoRAs simultaneously), you would need to manually merge the `lora_A` and `lora_B` matrices before loading, as the standard API treats each checkpoint as a complete replacement.