How to Fine-Tune a Pre-Trained Speech-to-Speech Model: A Complete Guide

Fine-tuning a pre-trained speech-to-speech model involves training the underlying STT, LLM, or TTS components outside the pipeline using standard Hugging Face scripts, then loading the custom checkpoints via CLI flags like --whisper_stt_model_name or --qwen3_tts_model_name in the speech-to-speech pipeline.

The huggingface/speech-to-speech repository implements a modular, low-latency voice-assistant pipeline that wires together four interchangeable components. To fine-tune a pre-trained speech-to-speech model, you must adapt the individual sub-models—Whisper for speech-to-text, Gemma or Qwen for the language model, or Qwen-3-TTS for synthesis—using external training scripts, then integrate the resulting checkpoints back into the pipeline without modifying the core orchestration logic.

Understanding the Modular Architecture

The pipeline orchestrates four distinct stages that run concurrently in separate threads, communicating via thread-safe queues initialized by initialize_queues_and_events. The high-level s2s_pipeline.py script drives the entire system using a ThreadManager to coordinate handlers.

The four components are:

  • Voice Activity Detection (VAD) – Detects speaker start/stop events (typically used as-is without fine-tuning).
  • Speech-to-Text (STT) – Transcribes audio; defaults to distil-whisper/distil-large-v3.
  • Large Language Model (LLM) – Generates responses from transcripts.
  • Text-to-Speech (TTS) – Synthesizes the LLM output into audio.

Each component loads its model via standard Hugging Face Auto* classes inside dedicated handlers (e.g., WhisperSTTHandler in src/speech_to_speech/STT/whisper_stt_handler.py and Qwen3TTSHandler in src/speech_to_speech/TTS/qwen3_tts_handler.py).

Fine-Tuning Workflow

The repository does not ship training loops. Instead, you fine-tune models externally and point the pipeline to your checkpoints using the arguments classes defined in src/speech_to_speech/arguments_classes/.

1. Select the Component to Adapt

Determine which stage requires domain-specific improvement:

  • STT: Fine-tune a Whisper model on domain-specific audio-transcript pairs.
  • LLM: Fine-tune a Transformers language model (e.g., Gemma, Qwen) on dialogue data.
  • TTS: Fine-tune Qwen-3-TTS on a voice-clone dataset.

2. Train Outside the Pipeline

Follow standard Hugging Face training scripts for your selected model family. The pipeline expects standard checkpoint formats compatible with AutoModelForSpeechSeq2Seq.from_pretrained or AutoTokenizer.from_pretrained.

3. Push the Checkpoint to the Hub

Upload your fine-tuned weights to Hugging Face or keep them locally:

huggingface-cli login

# Upload via huggingface_hub or git clone https://huggingface.co/your-username/model-name

4. Configure the Pipeline

Supply your checkpoint to the appropriate CLI flag. The rename_args helper in s2s_pipeline.py automatically rewrites prefixed arguments into handler-specific configurations:

Component CLI Flag Example Value
STT --whisper_stt_model_name my-org/whisper-finetuned-english
LLM --model_name my-org/gemma-7b-finetuned-conversational
TTS --qwen3_tts_model_name my-org/qwen3-tts-finetuned-voice-a

5. Launch the Pipeline

Execute the pipeline with your custom models:

speech-to-speech \
    --stt whisper \
    --whisper_stt_model_name my-org/whisper-finetuned-english \
    --llm_backend responses-api \
    --model_name my-org/gemma-finetuned-conv \
    --tts qwen3 \
    --qwen3_tts_model_name my-org/qwen3-tts-finetuned-voice-a \
    --mode realtime

6. Validate the Integration

Test the fine-tuned pipeline using the provided demo scripts (scripts/listen_and_play_realtime.py or scripts/listen_and_play.py) to confirm improved performance on your target domain.

Code Examples

Fine-Tuning Whisper for STT

Train a domain-specific Whisper model using the Transformers Trainer:

from transformers import WhisperProcessor, WhisperForConditionalGeneration, Trainer, TrainingArguments
import datasets

# Load dataset of (audio, transcript) pairs

train_dataset = datasets.load_dataset("path/to/your/dataset", split="train")

processor = WhisperProcessor.from_pretrained("distil-whisper/distil-large-v3")
model = WhisperForConditionalGeneration.from_pretrained("distil-whisper/distil-large-v3")

def prepare_batch(batch):
    audio = batch["audio"]["array"]
    input_features = processor(audio, sampling_rate=16_000, return_tensors="pt").input_features
    labels = processor.tokenizer(batch["text"], return_tensors="pt").input_ids
    batch["input_features"] = input_features.squeeze()
    batch["labels"] = labels.squeeze()
    return batch

tokenized = train_dataset.map(prepare_batch, remove_columns=train_dataset.column_names)

training_args = TrainingArguments(
    output_dir="./whisper_finetuned",
    per_device_train_batch_size=8,
    num_train_epochs=3,
    learning_rate=5e-5,
    fp16=True,
    logging_steps=10,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized,
)

trainer.train()
model.save_pretrained("my-org/whisper-finetuned-english")
processor.save_pretrained("my-org/whisper-finetuned-english")

Launching with Fine-Tuned Checkpoints (Python)

Programmatically launch the pipeline with custom checkpoints:

from speech_to_speech.s2s_pipeline import main as run_pipeline
import sys

# Simulate command-line arguments

sys.argv = [
    "s2s_pipeline.py",
    "--stt", "whisper",
    "--whisper_stt_model_name", "my-org/whisper-finetuned-english",
    "--llm_backend", "responses-api",
    "--model_name", "my-org/gemma-finetuned-conv",
    "--tts", "qwen3",
    "--qwen3_tts_model_name", "my-org/qwen3-tts-finetuned-voice-a",
    "--mode", "realtime",
]

run_pipeline()

Testing with the Realtime Client

Validate your fine-tuned pipeline using the OpenAI Realtime protocol:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed",
)

with client.realtime.connect(model="local") as conn:
    conn.send({
        "type": "session.update",
        "session": {
            "type": "realtime",
            "instructions": "You are a helpful assistant.",
            "audio": {"input": {"turn_detection": {"type": "server_vad", "interrupt_response": True}}},
        },
    })
    # Speak into the microphone; the pipeline returns audio generated by the fine-tuned TTS

Key Files and Implementation Details

Understanding these source files ensures successful integration of fine-tuned models:

Summary

  • Fine-tuning occurs externally – Train Whisper, LLM, or TTS models using standard Hugging Face scripts outside the pipeline.
  • Zero code changes required – The modular design accepts custom checkpoints via CLI flags without modifying s2s_pipeline.py.
  • Use specific flags – Reference your models with --whisper_stt_model_name, --model_name, or --qwen3_tts_model_name.
  • Automatic loading – Handlers automatically instantiate models using from_pretrained methods.
  • Validation tools – Use scripts/listen_and_play_realtime.py to verify your fine-tuned pipeline performance.

Frequently Asked Questions

Can I fine-tune the entire speech-to-speech pipeline end-to-end?

No, the huggingface/speech-to-speech repository does not implement end-to-end training loops. You must fine-tune each component (STT, LLM, TTS) separately using their respective training frameworks, then integrate the checkpoints into the inference pipeline.

Which component should I fine-tune first for better accuracy?

Start with the STT component if your domain contains specialized vocabulary or accents, as transcription errors propagate to the LLM. If your use case requires specific voice characteristics or speaking styles, prioritize fine-tuning the TTS component. Fine-tune the LLM only when you need specialized reasoning or conversational capabilities beyond the base model.

How do I know if my fine-tuned model is compatible with the pipeline?

Your checkpoint must be loadable by the standard Hugging Face Auto classes (e.g., AutoModelForSpeechSeq2Seq for Whisper, AutoTokenizer for LLMs). The pipeline expects the same format as the original pre-trained models (e.g., distil-whisper/distil-large-v3 or Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice). Ensure you save both the model weights and processor/tokenizer files.

Do I need to modify the source code to use my custom model?

No. As implemented in s2s_pipeline.py, the rename_args helper automatically routes your CLI arguments to the appropriate handlers. Simply pass your model identifier to the relevant flag (e.g., --whisper_stt_model_name my-org/model), and the handler will load it via from_pretrained without requiring modifications to src/speech_to_speech/STT/whisper_stt_handler.py or other handler files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →