# How to Fine-Tune a Pre-Trained Speech-to-Speech Model: A Complete Guide

> Learn how to fine-tune a pre-trained speech-to-speech model using Hugging Face. Customize STT, LLM, or TTS components with custom checkpoints for powerful results.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-07

---

**Fine-tuning a pre-trained speech-to-speech model involves training the underlying STT, LLM, or TTS components outside the pipeline using standard Hugging Face scripts, then loading the custom checkpoints via CLI flags like `--whisper_stt_model_name` or `--qwen3_tts_model_name` in the `speech-to-speech` pipeline.**

The `huggingface/speech-to-speech` repository implements a **modular, low-latency voice-assistant pipeline** that wires together four interchangeable components. To fine-tune a pre-trained speech-to-speech model, you must adapt the individual sub-models—Whisper for speech-to-text, Gemma or Qwen for the language model, or Qwen-3-TTS for synthesis—using external training scripts, then integrate the resulting checkpoints back into the pipeline without modifying the core orchestration logic.

## Understanding the Modular Architecture

The pipeline orchestrates four distinct stages that run concurrently in separate threads, communicating via thread-safe queues initialized by `initialize_queues_and_events`. The high-level [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) script drives the entire system using a `ThreadManager` to coordinate handlers.

The four components are:

- **Voice Activity Detection (VAD)** – Detects speaker start/stop events (typically used as-is without fine-tuning).
- **Speech-to-Text (STT)** – Transcribes audio; defaults to `distil-whisper/distil-large-v3`.
- **Large Language Model (LLM)** – Generates responses from transcripts.
- **Text-to-Speech (TTS)** – Synthesizes the LLM output into audio.

Each component loads its model via standard Hugging Face `Auto*` classes inside dedicated handlers (e.g., `WhisperSTTHandler` in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) and `Qwen3TTSHandler` in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)).

## Fine-Tuning Workflow

The repository **does not ship training loops**. Instead, you fine-tune models externally and point the pipeline to your checkpoints using the arguments classes defined in `src/speech_to_speech/arguments_classes/`.

### 1. Select the Component to Adapt

Determine which stage requires domain-specific improvement:

- **STT**: Fine-tune a Whisper model on domain-specific audio-transcript pairs.
- **LLM**: Fine-tune a Transformers language model (e.g., Gemma, Qwen) on dialogue data.
- **TTS**: Fine-tune Qwen-3-TTS on a voice-clone dataset.

### 2. Train Outside the Pipeline

Follow standard Hugging Face training scripts for your selected model family. The pipeline expects standard checkpoint formats compatible with `AutoModelForSpeechSeq2Seq.from_pretrained` or `AutoTokenizer.from_pretrained`.

### 3. Push the Checkpoint to the Hub

Upload your fine-tuned weights to Hugging Face or keep them locally:

```bash
huggingface-cli login

# Upload via huggingface_hub or git clone https://huggingface.co/your-username/model-name

```

### 4. Configure the Pipeline

Supply your checkpoint to the appropriate CLI flag. The `rename_args` helper in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) automatically rewrites prefixed arguments into handler-specific configurations:

| Component | CLI Flag | Example Value |
|-----------|----------|---------------|
| STT | `--whisper_stt_model_name` | `my-org/whisper-finetuned-english` |
| LLM | `--model_name` | `my-org/gemma-7b-finetuned-conversational` |
| TTS | `--qwen3_tts_model_name` | `my-org/qwen3-tts-finetuned-voice-a` |

### 5. Launch the Pipeline

Execute the pipeline with your custom models:

```bash
speech-to-speech \
    --stt whisper \
    --whisper_stt_model_name my-org/whisper-finetuned-english \
    --llm_backend responses-api \
    --model_name my-org/gemma-finetuned-conv \
    --tts qwen3 \
    --qwen3_tts_model_name my-org/qwen3-tts-finetuned-voice-a \
    --mode realtime

```

### 6. Validate the Integration

Test the fine-tuned pipeline using the provided demo scripts ([`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) or [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py)) to confirm improved performance on your target domain.

## Code Examples

### Fine-Tuning Whisper for STT

Train a domain-specific Whisper model using the Transformers `Trainer`:

```python
from transformers import WhisperProcessor, WhisperForConditionalGeneration, Trainer, TrainingArguments
import datasets

# Load dataset of (audio, transcript) pairs

train_dataset = datasets.load_dataset("path/to/your/dataset", split="train")

processor = WhisperProcessor.from_pretrained("distil-whisper/distil-large-v3")
model = WhisperForConditionalGeneration.from_pretrained("distil-whisper/distil-large-v3")

def prepare_batch(batch):
    audio = batch["audio"]["array"]
    input_features = processor(audio, sampling_rate=16_000, return_tensors="pt").input_features
    labels = processor.tokenizer(batch["text"], return_tensors="pt").input_ids
    batch["input_features"] = input_features.squeeze()
    batch["labels"] = labels.squeeze()
    return batch

tokenized = train_dataset.map(prepare_batch, remove_columns=train_dataset.column_names)

training_args = TrainingArguments(
    output_dir="./whisper_finetuned",
    per_device_train_batch_size=8,
    num_train_epochs=3,
    learning_rate=5e-5,
    fp16=True,
    logging_steps=10,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized,
)

trainer.train()
model.save_pretrained("my-org/whisper-finetuned-english")
processor.save_pretrained("my-org/whisper-finetuned-english")

```

### Launching with Fine-Tuned Checkpoints (Python)

Programmatically launch the pipeline with custom checkpoints:

```python
from speech_to_speech.s2s_pipeline import main as run_pipeline
import sys

# Simulate command-line arguments

sys.argv = [
    "s2s_pipeline.py",
    "--stt", "whisper",
    "--whisper_stt_model_name", "my-org/whisper-finetuned-english",
    "--llm_backend", "responses-api",
    "--model_name", "my-org/gemma-finetuned-conv",
    "--tts", "qwen3",
    "--qwen3_tts_model_name", "my-org/qwen3-tts-finetuned-voice-a",
    "--mode", "realtime",
]

run_pipeline()

```

### Testing with the Realtime Client

Validate your fine-tuned pipeline using the OpenAI Realtime protocol:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed",
)

with client.realtime.connect(model="local") as conn:
    conn.send({
        "type": "session.update",
        "session": {
            "type": "realtime",
            "instructions": "You are a helpful assistant.",
            "audio": {"input": {"turn_detection": {"type": "server_vad", "interrupt_response": True}}},
        },
    })
    # Speak into the microphone; the pipeline returns audio generated by the fine-tuned TTS

```

## Key Files and Implementation Details

Understanding these source files ensures successful integration of fine-tuned models:

- **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)** – Top-level orchestrator that parses arguments, builds handlers via `ThreadManager`, and starts the pipeline.
- **`src/speech_to_speech/arguments_classes/`** – Data-classes exposing all CLI flags; each component reads its own prefixed arguments.
- **[`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py)** – Loads Whisper checkpoints using `AutoProcessor` and `AutoModelForSpeechSeq2Seq`.
- **[`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)** – Loads Qwen-3-TTS via `FasterQwen3TTS.from_pretrained`.
- **[`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py)** – Handles local Transformers or `mlx-lm` LLMs; API backends use `ResponsesApiModelHandler`.

## Summary

- **Fine-tuning occurs externally** – Train Whisper, LLM, or TTS models using standard Hugging Face scripts outside the pipeline.
- **Zero code changes required** – The modular design accepts custom checkpoints via CLI flags without modifying [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).
- **Use specific flags** – Reference your models with `--whisper_stt_model_name`, `--model_name`, or `--qwen3_tts_model_name`.
- **Automatic loading** – Handlers automatically instantiate models using `from_pretrained` methods.
- **Validation tools** – Use [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) to verify your fine-tuned pipeline performance.

## Frequently Asked Questions

### Can I fine-tune the entire speech-to-speech pipeline end-to-end?

No, the `huggingface/speech-to-speech` repository does not implement end-to-end training loops. You must fine-tune each component (STT, LLM, TTS) separately using their respective training frameworks, then integrate the checkpoints into the inference pipeline.

### Which component should I fine-tune first for better accuracy?

Start with the **STT component** if your domain contains specialized vocabulary or accents, as transcription errors propagate to the LLM. If your use case requires specific voice characteristics or speaking styles, prioritize fine-tuning the **TTS component**. Fine-tune the LLM only when you need specialized reasoning or conversational capabilities beyond the base model.

### How do I know if my fine-tuned model is compatible with the pipeline?

Your checkpoint must be loadable by the standard Hugging Face `Auto` classes (e.g., `AutoModelForSpeechSeq2Seq` for Whisper, `AutoTokenizer` for LLMs). The pipeline expects the same format as the original pre-trained models (e.g., `distil-whisper/distil-large-v3` or `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice`). Ensure you save both the model weights and processor/tokenizer files.

### Do I need to modify the source code to use my custom model?

No. As implemented in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), the `rename_args` helper automatically routes your CLI arguments to the appropriate handlers. Simply pass your model identifier to the relevant flag (e.g., `--whisper_stt_model_name my-org/model`), and the handler will load it via `from_pretrained` without requiring modifications to [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) or other handler files.