# How to Fine-Tune a Speech-to-Speech Model: A Complete Guide for the Hugging Face Pipeline

> Learn to fine-tune a speech-to-speech model using the Hugging Face pipeline. Independently train STT LLM or TTS components and integrate them for custom speech applications.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**Fine-tuning a speech-to-speech model involves independently training the STT, LLM, or TTS components using standard Hugging Face workflows, then swapping the fine-tuned checkpoints into the modular pipeline via CLI arguments.**

The `huggingface/speech-to-speech` repository provides a modular, low-latency pipeline that converts spoken input into text, processes it through a language model, and synthesizes the response back into speech. Because each stage operates as a plug-and-play handler, you can fine-tune any learnable component independently and integrate it back into the system without modifying the core architecture. This guide walks you through fine-tuning the Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) components using the actual source code implementation.

## Understanding the Modular Architecture

The pipeline consists of four interchangeable components orchestrated by the core class in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). This file wires together VAD (Voice Activity Detection), STT, LLM, and TTS stages via queues and threads defined in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py) and [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py).

The default component stack includes:

- **VAD**: Silero VAD v5 for speech boundary detection
- **STT**: Parakeet TDT or Whisper models for transcription
- **LLM**: OpenAI-compatible Responses API or local Transformers models for response generation
- **TTS**: Qwen3-TTS (GGML on Linux, mlx-audio on macOS) for speech synthesis

Each component accepts custom model paths through dedicated argument classes located in `src/speech_to_speech/arguments_classes/`, allowing you to drop in fine-tuned checkpoints using simple CLI flags.

## Fine-Tuning the Speech-to-Text (STT) Component

Most STT backends in the repository are standard Hugging Face models compatible with the `transformers` library. You can fine-tune Whisper, Faster-Whisper, or Paraformer models using the standard training workflow, then point the pipeline to your checkpoint.

### Training Whisper on Custom Data

Install the repository with the appropriate extras and prepare your dataset:

```bash
pip install "speech-to-speech[faster-whisper]"
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync

# Prepare dataset (example using Common Voice)

python -m datasets load_dataset common_voice --lang en --split train

```

Execute the fine-tuning using the Hugging Face Trainer:

```bash
python -m transformers.trainer \
  --model_name_or_path openai/whisper-large-v2 \
  --dataset_name ./dataset \
  --output_dir ./whisper_finetuned \
  --per_device_train_batch_size 8 \
  --learning_rate 5e-5 \
  --num_train_epochs 3 \
  --max_steps 2000

```

Alternatively, use a Python script for more control over the preprocessing:

```python
from datasets import load_dataset
from transformers import WhisperProcessor, WhisperForConditionalGeneration, Trainer, TrainingArguments

dataset = load_dataset("common_voice", "en", split="train")
processor = WhisperProcessor.from_pretrained("openai/whisper-large-v2")
model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-large-v2")

def prepare_batch(batch):
    audio = batch["audio"]
    inputs = processor(audio["array"], sampling_rate=audio["sampling_rate"], return_tensors="pt")
    with processor.as_target_processor():
        labels = processor(batch["sentence"], return_tensors="pt").input_ids
    batch["input_features"] = inputs.input_features.squeeze()
    batch["labels"] = labels.squeeze()
    return batch

train_dataset = dataset.map(prepare_batch, remove_columns=dataset.column_names)

training_args = TrainingArguments(
    output_dir="./whisper_finetuned",
    per_device_train_batch_size=8,
    learning_rate=5e-5,
    num_train_epochs=3,
    fp16=True,
)

trainer = Trainer(model=model, args=training_args, train_dataset=train_dataset)
trainer.train()

```

### Integrating the Fine-Tuned STT Model

The **STT handler** reads the `--stt_model_name` argument through [`src/speech_to_speech/arguments_classes/whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/whisper_stt_arguments.py). After training completes, launch the pipeline with your fine-tuned checkpoint:

```bash
speech-to-speech \
  --stt whisper \
  --stt_model_name ./whisper_finetuned \
  --mode realtime

```

## Fine-Tuning the Large Language Model (LLM)

The LLM component supports both remote OpenAI-compatible endpoints and locally hosted models via Transformers or `mlx-lm`. For local fine-tuning, use standard Hugging Face training scripts with PEFT for efficient adaptation.

### Local Fine-Tuning with LoRA

Install the optional dependencies for local LLM inference:

```bash
pip install "speech-to-speech[mlx-lm]"
uv sync

```

Fine-tune a Qwen-3 model using LoRA:

```bash
python -m peft.trainer \
  --model_name_or_path Qwen/Qwen3-4B-Instruct-2507 \
  --dataset_name ./dialogues \
  --output_dir ./qwen_finetuned \
  --lora_rank 8 \
  --learning_rate 2e-4 \
  --num_train_epochs 2

```

The **LLM argument handler** in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) processes the `--model_name` flag. Insert your fine-tuned model using:

```bash
speech-to-speech \
  --llm_backend transformers \
  --model_name ./qwen_finetuned \
  --mode realtime

```

### Using Fine-Tuned Endpoints

If you prefer a hosted backend, configure the pipeline to point to your fine-tuned endpoint by setting the arguments defined in [`src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py):

```bash
speech-to-speech \
  --responses_api_base_url https://api.your-finetuned-endpoint.com \
  --responses_api_api_key $YOUR_API_KEY \
  --mode realtime

```

## Fine-Tuning the Text-to-Speech (TTS) Component

The default TTS backend uses Qwen3-TTS. Fine-tuning typically requires a paired audio-text dataset and the official Qwen3-TTS training scripts (external to this repository). After training, export your checkpoint and reference it via the `--qwen3_tts_model_name` argument handled in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py).

Install the necessary backend dependencies:

```bash
pip install "speech-to-speech[torch]"
uv sync

```

Execute the pipeline with your fine-tuned TTS model:

```bash
speech-to-speech \
  --tts qwen3 \
  --qwen3_tts_model_name ./qwen3_tts_finetuned \
  --mode realtime

```

## Running the Complete Fine-Tuned Pipeline

After fine-tuning any combination of components, run the fully integrated pipeline with a single command. The CLI automatically instantiates the appropriate handlers and connects them via the queue system defined in the pipeline modules:

```bash
speech-to-speech \
  --stt whisper \
  --stt_model_name ./whisper_finetuned \
  --llm_backend transformers \
  --model_name ./qwen_finetuned \
  --tts qwen3 \
  --qwen3_tts_model_name ./qwen3_tts_finetuned \
  --mode realtime \
  --enable_live_transcription

```

For programmatic access, use the Python API directly:

```python
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments

args = ModuleArguments(
    mode="realtime",
    stt="whisper",
    stt_model_name="./whisper_finetuned",
    llm_backend="transformers",
    model_name="./qwen_finetuned",
    tts="qwen3",
    qwen3_tts_model_name="./qwen3_tts_finetuned",
    enable_live_transcription=True,
)

pipeline = SpeechToSpeechPipeline(args)
pipeline.run()  # blocks; press Ctrl-C to stop

```

## Advanced Configuration Options

The modular architecture supports several advanced features to optimize your fine-tuned deployment:

- **Speculative Turns**: Reduce latency by generating LLM output while the user is still speaking. Enable via `--enable_speculative_turns` (implementation in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py)).
- **Multi-Language Support**: Pass `--language auto` and `--enable_lang_prompt` to ensure your fine-tuned STT and TTS models handle the target language correctly.
- **Custom VAD**: Swap the default Silero VAD for a different model by implementing a new handler and registering it via [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py).
- **Production Deployment**: Use the provided [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) to containerize the pipeline with your fine-tuned models for scalable deployment.

## Summary

- The `speech-to-speech` pipeline consists of four modular components (VAD, STT, LLM, TTS) orchestrated through [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) using thread-safe queues.
- **Fine-tune STT models** using standard Hugging Face Trainer workflows, then specify the checkpoint with `--stt_model_name`.
- **Fine-tune LLM models** using PEFT/LoRA or external APIs, integrating via `--model_name` or `--responses_api_base_url`.
- **Fine-tune TTS models** using the Qwen3-TTS training framework, loading checkpoints with `--qwen3_tts_model_name`.
- All component arguments are validated through dedicated classes in `src/speech_to_speech/arguments_classes/`.
- The pipeline supports speculative turns and live transcription out-of-the-box with fine-tuned models.

## Frequently Asked Questions

### Can I fine-tune only one component of the speech-to-speech pipeline?

Yes. Because the architecture uses plug-and-play handlers connected via queues in [`queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/queue_types.py), you can fine-tune just the STT, LLM, or TTS component independently. Simply specify the fine-tuned checkpoint path for the component you modified while keeping the default models for the others.

### What hardware requirements are needed for fine-tuning?

For STT fine-tuning, you need a GPU with sufficient VRAM to hold the Whisper model (typically 8GB+ for Whisper Large). For LLM fine-tuning, using LoRA through PEFT reduces memory requirements significantly, allowing fine-tuning of 4B parameter models on consumer GPUs (8-12GB VRAM). TTS fine-tuning requirements depend on the specific Qwen3-TTS implementation.

### How do I deploy my fine-tuned models in production?

Use the Docker Compose configuration provided in [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) to containerize the pipeline. Mount your fine-tuned model directories as volumes and reference them via the appropriate CLI arguments. The WebSocket server in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) exposes an OpenAI Realtime-compatible API for production integration.

### Does fine-tuning the LLM affect the latency of the pipeline?

Fine-tuning itself does not inherently increase latency if you maintain the same model architecture and quantization settings. The pipeline's speculative turns feature in [`speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/speculative_turns.py) can actually reduce perceived latency by generating responses before the user finishes speaking, regardless of whether the model is fine-tuned.