How to Fine-Tune a Speech-to-Speech Model: A Complete Guide for the Hugging Face Pipeline
Fine-tuning a speech-to-speech model involves independently training the STT, LLM, or TTS components using standard Hugging Face workflows, then swapping the fine-tuned checkpoints into the modular pipeline via CLI arguments.
The huggingface/speech-to-speech repository provides a modular, low-latency pipeline that converts spoken input into text, processes it through a language model, and synthesizes the response back into speech. Because each stage operates as a plug-and-play handler, you can fine-tune any learnable component independently and integrate it back into the system without modifying the core architecture. This guide walks you through fine-tuning the Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) components using the actual source code implementation.
Understanding the Modular Architecture
The pipeline consists of four interchangeable components orchestrated by the core class in src/speech_to_speech/s2s_pipeline.py. This file wires together VAD (Voice Activity Detection), STT, LLM, and TTS stages via queues and threads defined in src/speech_to_speech/pipeline/queue_types.py and src/speech_to_speech/pipeline/messages.py.
The default component stack includes:
- VAD: Silero VAD v5 for speech boundary detection
- STT: Parakeet TDT or Whisper models for transcription
- LLM: OpenAI-compatible Responses API or local Transformers models for response generation
- TTS: Qwen3-TTS (GGML on Linux, mlx-audio on macOS) for speech synthesis
Each component accepts custom model paths through dedicated argument classes located in src/speech_to_speech/arguments_classes/, allowing you to drop in fine-tuned checkpoints using simple CLI flags.
Fine-Tuning the Speech-to-Text (STT) Component
Most STT backends in the repository are standard Hugging Face models compatible with the transformers library. You can fine-tune Whisper, Faster-Whisper, or Paraformer models using the standard training workflow, then point the pipeline to your checkpoint.
Training Whisper on Custom Data
Install the repository with the appropriate extras and prepare your dataset:
pip install "speech-to-speech[faster-whisper]"
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync
# Prepare dataset (example using Common Voice)
python -m datasets load_dataset common_voice --lang en --split train
Execute the fine-tuning using the Hugging Face Trainer:
python -m transformers.trainer \
--model_name_or_path openai/whisper-large-v2 \
--dataset_name ./dataset \
--output_dir ./whisper_finetuned \
--per_device_train_batch_size 8 \
--learning_rate 5e-5 \
--num_train_epochs 3 \
--max_steps 2000
Alternatively, use a Python script for more control over the preprocessing:
from datasets import load_dataset
from transformers import WhisperProcessor, WhisperForConditionalGeneration, Trainer, TrainingArguments
dataset = load_dataset("common_voice", "en", split="train")
processor = WhisperProcessor.from_pretrained("openai/whisper-large-v2")
model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-large-v2")
def prepare_batch(batch):
audio = batch["audio"]
inputs = processor(audio["array"], sampling_rate=audio["sampling_rate"], return_tensors="pt")
with processor.as_target_processor():
labels = processor(batch["sentence"], return_tensors="pt").input_ids
batch["input_features"] = inputs.input_features.squeeze()
batch["labels"] = labels.squeeze()
return batch
train_dataset = dataset.map(prepare_batch, remove_columns=dataset.column_names)
training_args = TrainingArguments(
output_dir="./whisper_finetuned",
per_device_train_batch_size=8,
learning_rate=5e-5,
num_train_epochs=3,
fp16=True,
)
trainer = Trainer(model=model, args=training_args, train_dataset=train_dataset)
trainer.train()
Integrating the Fine-Tuned STT Model
The STT handler reads the --stt_model_name argument through src/speech_to_speech/arguments_classes/whisper_stt_arguments.py. After training completes, launch the pipeline with your fine-tuned checkpoint:
speech-to-speech \
--stt whisper \
--stt_model_name ./whisper_finetuned \
--mode realtime
Fine-Tuning the Large Language Model (LLM)
The LLM component supports both remote OpenAI-compatible endpoints and locally hosted models via Transformers or mlx-lm. For local fine-tuning, use standard Hugging Face training scripts with PEFT for efficient adaptation.
Local Fine-Tuning with LoRA
Install the optional dependencies for local LLM inference:
pip install "speech-to-speech[mlx-lm]"
uv sync
Fine-tune a Qwen-3 model using LoRA:
python -m peft.trainer \
--model_name_or_path Qwen/Qwen3-4B-Instruct-2507 \
--dataset_name ./dialogues \
--output_dir ./qwen_finetuned \
--lora_rank 8 \
--learning_rate 2e-4 \
--num_train_epochs 2
The LLM argument handler in src/speech_to_speech/arguments_classes/module_arguments.py processes the --model_name flag. Insert your fine-tuned model using:
speech-to-speech \
--llm_backend transformers \
--model_name ./qwen_finetuned \
--mode realtime
Using Fine-Tuned Endpoints
If you prefer a hosted backend, configure the pipeline to point to your fine-tuned endpoint by setting the arguments defined in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py:
speech-to-speech \
--responses_api_base_url https://api.your-finetuned-endpoint.com \
--responses_api_api_key $YOUR_API_KEY \
--mode realtime
Fine-Tuning the Text-to-Speech (TTS) Component
The default TTS backend uses Qwen3-TTS. Fine-tuning typically requires a paired audio-text dataset and the official Qwen3-TTS training scripts (external to this repository). After training, export your checkpoint and reference it via the --qwen3_tts_model_name argument handled in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py.
Install the necessary backend dependencies:
pip install "speech-to-speech[torch]"
uv sync
Execute the pipeline with your fine-tuned TTS model:
speech-to-speech \
--tts qwen3 \
--qwen3_tts_model_name ./qwen3_tts_finetuned \
--mode realtime
Running the Complete Fine-Tuned Pipeline
After fine-tuning any combination of components, run the fully integrated pipeline with a single command. The CLI automatically instantiates the appropriate handlers and connects them via the queue system defined in the pipeline modules:
speech-to-speech \
--stt whisper \
--stt_model_name ./whisper_finetuned \
--llm_backend transformers \
--model_name ./qwen_finetuned \
--tts qwen3 \
--qwen3_tts_model_name ./qwen3_tts_finetuned \
--mode realtime \
--enable_live_transcription
For programmatic access, use the Python API directly:
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
args = ModuleArguments(
mode="realtime",
stt="whisper",
stt_model_name="./whisper_finetuned",
llm_backend="transformers",
model_name="./qwen_finetuned",
tts="qwen3",
qwen3_tts_model_name="./qwen3_tts_finetuned",
enable_live_transcription=True,
)
pipeline = SpeechToSpeechPipeline(args)
pipeline.run() # blocks; press Ctrl-C to stop
Advanced Configuration Options
The modular architecture supports several advanced features to optimize your fine-tuned deployment:
- Speculative Turns: Reduce latency by generating LLM output while the user is still speaking. Enable via
--enable_speculative_turns(implementation insrc/speech_to_speech/pipeline/speculative_turns.py). - Multi-Language Support: Pass
--language autoand--enable_lang_promptto ensure your fine-tuned STT and TTS models handle the target language correctly. - Custom VAD: Swap the default Silero VAD for a different model by implementing a new handler and registering it via
src/speech_to_speech/arguments_classes/vad_arguments.py. - Production Deployment: Use the provided
docker-compose.ymlto containerize the pipeline with your fine-tuned models for scalable deployment.
Summary
- The
speech-to-speechpipeline consists of four modular components (VAD, STT, LLM, TTS) orchestrated throughs2s_pipeline.pyusing thread-safe queues. - Fine-tune STT models using standard Hugging Face Trainer workflows, then specify the checkpoint with
--stt_model_name. - Fine-tune LLM models using PEFT/LoRA or external APIs, integrating via
--model_nameor--responses_api_base_url. - Fine-tune TTS models using the Qwen3-TTS training framework, loading checkpoints with
--qwen3_tts_model_name. - All component arguments are validated through dedicated classes in
src/speech_to_speech/arguments_classes/. - The pipeline supports speculative turns and live transcription out-of-the-box with fine-tuned models.
Frequently Asked Questions
Can I fine-tune only one component of the speech-to-speech pipeline?
Yes. Because the architecture uses plug-and-play handlers connected via queues in queue_types.py, you can fine-tune just the STT, LLM, or TTS component independently. Simply specify the fine-tuned checkpoint path for the component you modified while keeping the default models for the others.
What hardware requirements are needed for fine-tuning?
For STT fine-tuning, you need a GPU with sufficient VRAM to hold the Whisper model (typically 8GB+ for Whisper Large). For LLM fine-tuning, using LoRA through PEFT reduces memory requirements significantly, allowing fine-tuning of 4B parameter models on consumer GPUs (8-12GB VRAM). TTS fine-tuning requirements depend on the specific Qwen3-TTS implementation.
How do I deploy my fine-tuned models in production?
Use the Docker Compose configuration provided in docker-compose.yml to containerize the pipeline. Mount your fine-tuned model directories as volumes and reference them via the appropriate CLI arguments. The WebSocket server in src/speech_to_speech/api/openai_realtime/server.py exposes an OpenAI Realtime-compatible API for production integration.
Does fine-tuning the LLM affect the latency of the pipeline?
Fine-tuning itself does not inherently increase latency if you maintain the same model architecture and quantization settings. The pipeline's speculative turns feature in speculative_turns.py can actually reduce perceived latency by generating responses before the user finishes speaking, regardless of whether the model is fine-tuned.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →