# How to Fine-Tune VibeVoice-ASR with LoRA: A Complete Parameter-Efficient Guide

> Fine-tune VibeVoice-ASR efficiently with LoRA. Adapt your speech model for specific domains using low-rank matrices on modest GPU hardware. Read our complete parameter-efficient guide.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: tutorial
- Published: 2026-03-28

---

**You can fine-tune VibeVoice-ASR efficiently using LoRA (Low-Rank Adaptation) by freezing the pretrained speech encoder and LLM decoder weights while injecting trainable low-rank matrices into specific projection layers, enabling domain-specific adaptation on modest GPU hardware like a single RTX 3090.**

VibeVoice-ASR is Microsoft's open-source multimodal automatic speech recognition system that pairs a speech encoder with a large language model decoder. This guide walks through the complete process of fine-tuning VibeVoice-ASR with LoRA using the reference implementation in the `microsoft/VibeVoice` repository, covering everything from data preparation to inference with saved adapters.

## VibeVoice-ASR Architecture for LoRA Adaptation

The fine-tuning pipeline centers on **VibeVoiceASRForConditionalGeneration**, defined in [`vibevoice/modular/modeling_vibevoice_asr.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_asr.py), which combines a speech tokenizer encoder with an LLM decoder backbone. When applying LoRA, the implementation selectively targets the decoder's linear projection layers while keeping the speech encoder frozen.

Key components in the fine-tuning stack include:

- **VibeVoiceASRProcessor** ([`vibevoice/processor/vibevoice_asr_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_asr_processor.py)) – Handles audio preprocessing, feature extraction, and chat-style prompt construction using the model's conversation template.
- **LoRA Configuration** (`get_lora_config()` in [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py)) – Defines the rank (`r`), scaling factor (`alpha`), dropout rate, and specifically targets the query, key, value, and output projection matrices (`q_proj`, `k_proj`, `v_proj`, `o_proj`) plus MLP layers.
- **VibeVoiceASRDataset** ([`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py)) – Loads audio-transcript pairs from disk, filters by duration, and injects optional domain context into the prompts.
- **VibeVoiceASRDataCollator** ([`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py)) – Handles dynamic padding of variable-length audio sequences and token IDs for the Hugging Face Trainer.

## Configuring LoRA Parameters

The `get_lora_config()` function in [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py) constructs a `peft.LoraConfig` object that controls which parameters receive gradient updates. By default, the configuration targets the language model's attention projections and feed-forward layers while excluding the speech encoder and embedding layers from training.

Critical parameters include:

- **lora_r** – The rank of the low-rank decomposition (commonly set to 16).
- **lora_alpha** – The scaling factor for the LoRA layers (typically 32, creating a 2:1 alpha-to-r ratio).
- **lora_dropout** – Regularization dropout applied to the LoRA layers (commonly 0.05).
- **target_modules** – Explicit list including `q_proj`, `k_proj`, `v_proj`, `o_proj`, and MLP projections.

This selective targeting ensures that fewer than 5% of total parameters become trainable, dramatically reducing memory overhead compared to full fine-tuning.

## Preparing Your Training Data

The `VibeVoiceASRDataset` class expects a directory of audio files (`.mp3` or `.wav`) paired with JSON label files containing structured transcript information. Each JSON entry must specify the audio path, duration, speaker segments, and optional domain context.

Example dataset entry:

```json
{
  "audio_path": "0.mp3",
  "audio_duration": 12.34,
  "segments": [
    {"speaker": 0, "text": "Hello, welcome to VibeVoice.", "start": 0.0, "end": 3.2},
    {"speaker": 1, "text": "Thanks for the intro.", "start": 3.5, "end": 5.0}
  ],
  "customized_context": ["Domain: Podcast", "Topic: AI assistants"]
}

```

Place these JSON files alongside their referenced audio files in your data directory. The processor automatically applies the chat template to combine the context, speaker information, and transcription into the format expected by the multimodal model.

## The Fine-Tuning Workflow

The `train()` function in [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py) orchestrates the complete training pipeline through the following steps:

1. **Argument Parsing** – Uses `HfArgumentParser` to process `ModelArguments`, `DataArguments`, `LoraArguments`, and standard `TrainingArguments` from the 🤗 Transformers library.

2. **Model Loading and Freezing** – `setup_model_for_training()` loads the pretrained checkpoint from `microsoft/VibeVoice-ASR`, freezes the speech tokenizer parameters, and prepares the LLM decoder for adapter injection.

3. **LoRA Adapter Injection** – Invokes `peft.get_peft_model()` to wrap the base model with the LoRA configuration, exposing only the low-rank matrices for gradient computation.

4. **Dataset Construction** – Instantiates `VibeVoiceASRDataset` with the specified audio directory, applying the processor to convert raw audio and JSON labels into tokenized model inputs.

5. **Collator Initialization** – Creates `VibeVoiceASRDataCollator` to handle batching and padding of mixed audio-text sequences.

6. **Trainer Execution** – The 🤗 Transformers `Trainer` manages the optimization loop, gradient accumulation, mixed-precision training (`bfloat16`), and distributed checkpointing.

7. **Adapter Saving** – Exports the trained LoRA weights and processor configuration to the specified `output_dir`, preserving the frozen base model weights separately.

Launch training with the following command:

```bash
python -m finetuning-asr.lora_finetune \
    --model_path microsoft/VibeVoice-ASR \
    --data_dir ./finetuning-asr/toy_dataset \
    --output_dir ./lora_checkpoints \
    --per_device_train_batch_size 2 \
    --gradient_accumulation_steps 4 \
    --num_train_epochs 3 \
    --learning_rate 3e-4 \
    --lora_r 16 \
    --lora_alpha 32 \
    --lora_dropout 0.05

```

## Running Inference with LoRA Adapters

After training, use [`finetuning-asr/inference_lora.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/inference_lora.py) to apply the fine-tuned adapter to new audio inputs. The script loads the base VibeVoice-ASR model, merges the LoRA weights from your checkpoint, and executes `model.generate()` on the processed audio features.

Run inference with:

```bash
python -m finetuning-asr.inference_lora \
    --model_path microsoft/VibeVoice-ASR \
    --lora_path ./lora_checkpoints \
    --audio_path ./demo/asr_demo/demo1-chat.mp3 \
    --output_file result.txt

```

The inference script handles adapter merging automatically, allowing you to switch between different domain-specific LoRA checkpoints without reloading the full base model weights.

## Summary

- **LoRA enables efficient adaptation** of VibeVoice-ASR by training less than 5% of total parameters while freezing the speech encoder and LLM decoder.
- **Target modules** for adaptation are configured in `get_lora_config()` and include the decoder's `q_proj`, `k_proj`, `v_proj`, `o_proj`, and MLP layers.
- **Data format** requires paired audio files and JSON transcripts with speaker segments and optional domain context.
- **Training entry point** is [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py), which integrates with the 🤗 Transformers Trainer for distributed, mixed-precision training.
- **Inference** uses [`finetuning-asr/inference_lora.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/inference_lora.py) to load base weights and apply domain-specific LoRA adapters for transcription.

## Frequently Asked Questions

### What hardware is required to fine-tune VibeVoice-ASR with LoRA?

According to the implementation in `microsoft/VibeVoice`, LoRA fine-tuning requires significantly less memory than full fine-tuning because only the low-rank adapter matrices receive gradients. You can train on a single RTX 3090 (24GB VRAM) or a single A100 using `bfloat16` mixed precision, depending on batch size and sequence length.

### Which model layers are frozen during LoRA fine-tuning?

The `setup_model_for_training()` function in [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py) freezes the entire speech tokenizer/encoder and the base LLM decoder weights. Only the injected LoRA adapters attached to the decoder's projection layers (`q_proj`, `k_proj`, `v_proj`, `o_proj`) and MLP blocks remain trainable, drastically reducing the optimizer state memory footprint.

### How do I format my own training data for VibeVoice-ASR?

Create JSON files matching the schema expected by `VibeVoiceASRDataset` in [`finetuning-asr/lora_finetune.py`](https://github.com/microsoft/VibeVoice/blob/main/finetuning-asr/lora_finetune.py). Each JSON must contain `audio_path` (relative filename), `audio_duration` (seconds), and `segments` (array of speaker/text/start/end objects). Optional `customized_context` arrays allow you to inject domain metadata that the processor formats into the chat template prompt.

### Can I merge the LoRA weights back into the base model for deployment?

Yes. The [`inference_lora.py`](https://github.com/microsoft/VibeVoice/blob/main/inference_lora.py) script demonstrates loading LoRA weights alongside the base checkpoint. For permanent merging, you can use the PEFT library's `merge_and_unload()` method on the model object after loading with `peft.get_peft_model()`, though the repository provides separate inference scripts that apply adapters dynamically without merging.