How to Fine-Tune VibeVoice-ASR with LoRA: A Complete Parameter-Efficient Guide

You can fine-tune VibeVoice-ASR efficiently using LoRA (Low-Rank Adaptation) by freezing the pretrained speech encoder and LLM decoder weights while injecting trainable low-rank matrices into specific projection layers, enabling domain-specific adaptation on modest GPU hardware like a single RTX 3090.

VibeVoice-ASR is Microsoft's open-source multimodal automatic speech recognition system that pairs a speech encoder with a large language model decoder. This guide walks through the complete process of fine-tuning VibeVoice-ASR with LoRA using the reference implementation in the microsoft/VibeVoice repository, covering everything from data preparation to inference with saved adapters.

VibeVoice-ASR Architecture for LoRA Adaptation

The fine-tuning pipeline centers on VibeVoiceASRForConditionalGeneration, defined in vibevoice/modular/modeling_vibevoice_asr.py, which combines a speech tokenizer encoder with an LLM decoder backbone. When applying LoRA, the implementation selectively targets the decoder's linear projection layers while keeping the speech encoder frozen.

Key components in the fine-tuning stack include:

  • VibeVoiceASRProcessor (vibevoice/processor/vibevoice_asr_processor.py) – Handles audio preprocessing, feature extraction, and chat-style prompt construction using the model's conversation template.
  • LoRA Configuration (get_lora_config() in finetuning-asr/lora_finetune.py) – Defines the rank (r), scaling factor (alpha), dropout rate, and specifically targets the query, key, value, and output projection matrices (q_proj, k_proj, v_proj, o_proj) plus MLP layers.
  • VibeVoiceASRDataset (finetuning-asr/lora_finetune.py) – Loads audio-transcript pairs from disk, filters by duration, and injects optional domain context into the prompts.
  • VibeVoiceASRDataCollator (finetuning-asr/lora_finetune.py) – Handles dynamic padding of variable-length audio sequences and token IDs for the Hugging Face Trainer.

Configuring LoRA Parameters

The get_lora_config() function in finetuning-asr/lora_finetune.py constructs a peft.LoraConfig object that controls which parameters receive gradient updates. By default, the configuration targets the language model's attention projections and feed-forward layers while excluding the speech encoder and embedding layers from training.

Critical parameters include:

  • lora_r – The rank of the low-rank decomposition (commonly set to 16).
  • lora_alpha – The scaling factor for the LoRA layers (typically 32, creating a 2:1 alpha-to-r ratio).
  • lora_dropout – Regularization dropout applied to the LoRA layers (commonly 0.05).
  • target_modules – Explicit list including q_proj, k_proj, v_proj, o_proj, and MLP projections.

This selective targeting ensures that fewer than 5% of total parameters become trainable, dramatically reducing memory overhead compared to full fine-tuning.

Preparing Your Training Data

The VibeVoiceASRDataset class expects a directory of audio files (.mp3 or .wav) paired with JSON label files containing structured transcript information. Each JSON entry must specify the audio path, duration, speaker segments, and optional domain context.

Example dataset entry:

{
  "audio_path": "0.mp3",
  "audio_duration": 12.34,
  "segments": [
    {"speaker": 0, "text": "Hello, welcome to VibeVoice.", "start": 0.0, "end": 3.2},
    {"speaker": 1, "text": "Thanks for the intro.", "start": 3.5, "end": 5.0}
  ],
  "customized_context": ["Domain: Podcast", "Topic: AI assistants"]
}

Place these JSON files alongside their referenced audio files in your data directory. The processor automatically applies the chat template to combine the context, speaker information, and transcription into the format expected by the multimodal model.

The Fine-Tuning Workflow

The train() function in finetuning-asr/lora_finetune.py orchestrates the complete training pipeline through the following steps:

  1. Argument Parsing – Uses HfArgumentParser to process ModelArguments, DataArguments, LoraArguments, and standard TrainingArguments from the 🤗 Transformers library.

  2. Model Loading and Freezing – setup_model_for_training() loads the pretrained checkpoint from microsoft/VibeVoice-ASR, freezes the speech tokenizer parameters, and prepares the LLM decoder for adapter injection.

  3. LoRA Adapter Injection – Invokes peft.get_peft_model() to wrap the base model with the LoRA configuration, exposing only the low-rank matrices for gradient computation.

  4. Dataset Construction – Instantiates VibeVoiceASRDataset with the specified audio directory, applying the processor to convert raw audio and JSON labels into tokenized model inputs.

  5. Collator Initialization – Creates VibeVoiceASRDataCollator to handle batching and padding of mixed audio-text sequences.

  6. Trainer Execution – The 🤗 Transformers Trainer manages the optimization loop, gradient accumulation, mixed-precision training (bfloat16), and distributed checkpointing.

  7. Adapter Saving – Exports the trained LoRA weights and processor configuration to the specified output_dir, preserving the frozen base model weights separately.

Launch training with the following command:

python -m finetuning-asr.lora_finetune \
    --model_path microsoft/VibeVoice-ASR \
    --data_dir ./finetuning-asr/toy_dataset \
    --output_dir ./lora_checkpoints \
    --per_device_train_batch_size 2 \
    --gradient_accumulation_steps 4 \
    --num_train_epochs 3 \
    --learning_rate 3e-4 \
    --lora_r 16 \
    --lora_alpha 32 \
    --lora_dropout 0.05

Running Inference with LoRA Adapters

After training, use finetuning-asr/inference_lora.py to apply the fine-tuned adapter to new audio inputs. The script loads the base VibeVoice-ASR model, merges the LoRA weights from your checkpoint, and executes model.generate() on the processed audio features.

Run inference with:

python -m finetuning-asr.inference_lora \
    --model_path microsoft/VibeVoice-ASR \
    --lora_path ./lora_checkpoints \
    --audio_path ./demo/asr_demo/demo1-chat.mp3 \
    --output_file result.txt

The inference script handles adapter merging automatically, allowing you to switch between different domain-specific LoRA checkpoints without reloading the full base model weights.

Summary

  • LoRA enables efficient adaptation of VibeVoice-ASR by training less than 5% of total parameters while freezing the speech encoder and LLM decoder.
  • Target modules for adaptation are configured in get_lora_config() and include the decoder's q_proj, k_proj, v_proj, o_proj, and MLP layers.
  • Data format requires paired audio files and JSON transcripts with speaker segments and optional domain context.
  • Training entry point is finetuning-asr/lora_finetune.py, which integrates with the 🤗 Transformers Trainer for distributed, mixed-precision training.
  • Inference uses finetuning-asr/inference_lora.py to load base weights and apply domain-specific LoRA adapters for transcription.

Frequently Asked Questions

What hardware is required to fine-tune VibeVoice-ASR with LoRA?

According to the implementation in microsoft/VibeVoice, LoRA fine-tuning requires significantly less memory than full fine-tuning because only the low-rank adapter matrices receive gradients. You can train on a single RTX 3090 (24GB VRAM) or a single A100 using bfloat16 mixed precision, depending on batch size and sequence length.

Which model layers are frozen during LoRA fine-tuning?

The setup_model_for_training() function in finetuning-asr/lora_finetune.py freezes the entire speech tokenizer/encoder and the base LLM decoder weights. Only the injected LoRA adapters attached to the decoder's projection layers (q_proj, k_proj, v_proj, o_proj) and MLP blocks remain trainable, drastically reducing the optimizer state memory footprint.

How do I format my own training data for VibeVoice-ASR?

Create JSON files matching the schema expected by VibeVoiceASRDataset in finetuning-asr/lora_finetune.py. Each JSON must contain audio_path (relative filename), audio_duration (seconds), and segments (array of speaker/text/start/end objects). Optional customized_context arrays allow you to inject domain metadata that the processor formats into the chat template prompt.

Can I merge the LoRA weights back into the base model for deployment?

Yes. The inference_lora.py script demonstrates loading LoRA weights alongside the base checkpoint. For permanent merging, you can use the PEFT library's merge_and_unload() method on the model object after loading with peft.get_peft_model(), though the repository provides separate inference scripts that apply adapters dynamically without merging.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →