How to Fine-Tune a Pre-Trained Speech-to-Speech Model: A Complete Guide
Fine-tuning a pre-trained speech-to-speech model involves training the underlying STT, LLM, or TTS components outside the pipeline using standard Hugging Face scripts, then loading the custom checkpoints via CLI flags like --whisper_stt_model_name or --qwen3_tts_model_name in the speech-to-speech pipeline.
The huggingface/speech-to-speech repository implements a modular, low-latency voice-assistant pipeline that wires together four interchangeable components. To fine-tune a pre-trained speech-to-speech model, you must adapt the individual sub-models—Whisper for speech-to-text, Gemma or Qwen for the language model, or Qwen-3-TTS for synthesis—using external training scripts, then integrate the resulting checkpoints back into the pipeline without modifying the core orchestration logic.
Understanding the Modular Architecture
The pipeline orchestrates four distinct stages that run concurrently in separate threads, communicating via thread-safe queues initialized by initialize_queues_and_events. The high-level s2s_pipeline.py script drives the entire system using a ThreadManager to coordinate handlers.
The four components are:
- Voice Activity Detection (VAD) – Detects speaker start/stop events (typically used as-is without fine-tuning).
- Speech-to-Text (STT) – Transcribes audio; defaults to
distil-whisper/distil-large-v3. - Large Language Model (LLM) – Generates responses from transcripts.
- Text-to-Speech (TTS) – Synthesizes the LLM output into audio.
Each component loads its model via standard Hugging Face Auto* classes inside dedicated handlers (e.g., WhisperSTTHandler in src/speech_to_speech/STT/whisper_stt_handler.py and Qwen3TTSHandler in src/speech_to_speech/TTS/qwen3_tts_handler.py).
Fine-Tuning Workflow
The repository does not ship training loops. Instead, you fine-tune models externally and point the pipeline to your checkpoints using the arguments classes defined in src/speech_to_speech/arguments_classes/.
1. Select the Component to Adapt
Determine which stage requires domain-specific improvement:
- STT: Fine-tune a Whisper model on domain-specific audio-transcript pairs.
- LLM: Fine-tune a Transformers language model (e.g., Gemma, Qwen) on dialogue data.
- TTS: Fine-tune Qwen-3-TTS on a voice-clone dataset.
2. Train Outside the Pipeline
Follow standard Hugging Face training scripts for your selected model family. The pipeline expects standard checkpoint formats compatible with AutoModelForSpeechSeq2Seq.from_pretrained or AutoTokenizer.from_pretrained.
3. Push the Checkpoint to the Hub
Upload your fine-tuned weights to Hugging Face or keep them locally:
huggingface-cli login
# Upload via huggingface_hub or git clone https://huggingface.co/your-username/model-name
4. Configure the Pipeline
Supply your checkpoint to the appropriate CLI flag. The rename_args helper in s2s_pipeline.py automatically rewrites prefixed arguments into handler-specific configurations:
| Component | CLI Flag | Example Value |
|---|---|---|
| STT | --whisper_stt_model_name |
my-org/whisper-finetuned-english |
| LLM | --model_name |
my-org/gemma-7b-finetuned-conversational |
| TTS | --qwen3_tts_model_name |
my-org/qwen3-tts-finetuned-voice-a |
5. Launch the Pipeline
Execute the pipeline with your custom models:
speech-to-speech \
--stt whisper \
--whisper_stt_model_name my-org/whisper-finetuned-english \
--llm_backend responses-api \
--model_name my-org/gemma-finetuned-conv \
--tts qwen3 \
--qwen3_tts_model_name my-org/qwen3-tts-finetuned-voice-a \
--mode realtime
6. Validate the Integration
Test the fine-tuned pipeline using the provided demo scripts (scripts/listen_and_play_realtime.py or scripts/listen_and_play.py) to confirm improved performance on your target domain.
Code Examples
Fine-Tuning Whisper for STT
Train a domain-specific Whisper model using the Transformers Trainer:
from transformers import WhisperProcessor, WhisperForConditionalGeneration, Trainer, TrainingArguments
import datasets
# Load dataset of (audio, transcript) pairs
train_dataset = datasets.load_dataset("path/to/your/dataset", split="train")
processor = WhisperProcessor.from_pretrained("distil-whisper/distil-large-v3")
model = WhisperForConditionalGeneration.from_pretrained("distil-whisper/distil-large-v3")
def prepare_batch(batch):
audio = batch["audio"]["array"]
input_features = processor(audio, sampling_rate=16_000, return_tensors="pt").input_features
labels = processor.tokenizer(batch["text"], return_tensors="pt").input_ids
batch["input_features"] = input_features.squeeze()
batch["labels"] = labels.squeeze()
return batch
tokenized = train_dataset.map(prepare_batch, remove_columns=train_dataset.column_names)
training_args = TrainingArguments(
output_dir="./whisper_finetuned",
per_device_train_batch_size=8,
num_train_epochs=3,
learning_rate=5e-5,
fp16=True,
logging_steps=10,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized,
)
trainer.train()
model.save_pretrained("my-org/whisper-finetuned-english")
processor.save_pretrained("my-org/whisper-finetuned-english")
Launching with Fine-Tuned Checkpoints (Python)
Programmatically launch the pipeline with custom checkpoints:
from speech_to_speech.s2s_pipeline import main as run_pipeline
import sys
# Simulate command-line arguments
sys.argv = [
"s2s_pipeline.py",
"--stt", "whisper",
"--whisper_stt_model_name", "my-org/whisper-finetuned-english",
"--llm_backend", "responses-api",
"--model_name", "my-org/gemma-finetuned-conv",
"--tts", "qwen3",
"--qwen3_tts_model_name", "my-org/qwen3-tts-finetuned-voice-a",
"--mode", "realtime",
]
run_pipeline()
Testing with the Realtime Client
Validate your fine-tuned pipeline using the OpenAI Realtime protocol:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8765/v1",
websocket_base_url="ws://localhost:8765/v1",
api_key="not-needed",
)
with client.realtime.connect(model="local") as conn:
conn.send({
"type": "session.update",
"session": {
"type": "realtime",
"instructions": "You are a helpful assistant.",
"audio": {"input": {"turn_detection": {"type": "server_vad", "interrupt_response": True}}},
},
})
# Speak into the microphone; the pipeline returns audio generated by the fine-tuned TTS
Key Files and Implementation Details
Understanding these source files ensures successful integration of fine-tuned models:
src/speech_to_speech/s2s_pipeline.py– Top-level orchestrator that parses arguments, builds handlers viaThreadManager, and starts the pipeline.src/speech_to_speech/arguments_classes/– Data-classes exposing all CLI flags; each component reads its own prefixed arguments.src/speech_to_speech/STT/whisper_stt_handler.py– Loads Whisper checkpoints usingAutoProcessorandAutoModelForSpeechSeq2Seq.src/speech_to_speech/TTS/qwen3_tts_handler.py– Loads Qwen-3-TTS viaFasterQwen3TTS.from_pretrained.src/speech_to_speech/LLM/language_model.py– Handles local Transformers ormlx-lmLLMs; API backends useResponsesApiModelHandler.
Summary
- Fine-tuning occurs externally – Train Whisper, LLM, or TTS models using standard Hugging Face scripts outside the pipeline.
- Zero code changes required – The modular design accepts custom checkpoints via CLI flags without modifying
s2s_pipeline.py. - Use specific flags – Reference your models with
--whisper_stt_model_name,--model_name, or--qwen3_tts_model_name. - Automatic loading – Handlers automatically instantiate models using
from_pretrainedmethods. - Validation tools – Use
scripts/listen_and_play_realtime.pyto verify your fine-tuned pipeline performance.
Frequently Asked Questions
Can I fine-tune the entire speech-to-speech pipeline end-to-end?
No, the huggingface/speech-to-speech repository does not implement end-to-end training loops. You must fine-tune each component (STT, LLM, TTS) separately using their respective training frameworks, then integrate the checkpoints into the inference pipeline.
Which component should I fine-tune first for better accuracy?
Start with the STT component if your domain contains specialized vocabulary or accents, as transcription errors propagate to the LLM. If your use case requires specific voice characteristics or speaking styles, prioritize fine-tuning the TTS component. Fine-tune the LLM only when you need specialized reasoning or conversational capabilities beyond the base model.
How do I know if my fine-tuned model is compatible with the pipeline?
Your checkpoint must be loadable by the standard Hugging Face Auto classes (e.g., AutoModelForSpeechSeq2Seq for Whisper, AutoTokenizer for LLMs). The pipeline expects the same format as the original pre-trained models (e.g., distil-whisper/distil-large-v3 or Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice). Ensure you save both the model weights and processor/tokenizer files.
Do I need to modify the source code to use my custom model?
No. As implemented in s2s_pipeline.py, the rename_args helper automatically routes your CLI arguments to the appropriate handlers. Simply pass your model identifier to the relevant flag (e.g., --whisper_stt_model_name my-org/model), and the handler will load it via from_pretrained without requiring modifications to src/speech_to_speech/STT/whisper_stt_handler.py or other handler files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →