How to Fine-Tune a Speech-to-Speech Model: A Complete Guide for the Hugging Face Pipeline

Fine-tuning a speech-to-speech model involves independently training the STT, LLM, or TTS components using standard Hugging Face workflows, then swapping the fine-tuned checkpoints into the modular pipeline via CLI arguments.

The huggingface/speech-to-speech repository provides a modular, low-latency pipeline that converts spoken input into text, processes it through a language model, and synthesizes the response back into speech. Because each stage operates as a plug-and-play handler, you can fine-tune any learnable component independently and integrate it back into the system without modifying the core architecture. This guide walks you through fine-tuning the Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) components using the actual source code implementation.

Understanding the Modular Architecture

The pipeline consists of four interchangeable components orchestrated by the core class in src/speech_to_speech/s2s_pipeline.py. This file wires together VAD (Voice Activity Detection), STT, LLM, and TTS stages via queues and threads defined in src/speech_to_speech/pipeline/queue_types.py and src/speech_to_speech/pipeline/messages.py.

The default component stack includes:

  • VAD: Silero VAD v5 for speech boundary detection
  • STT: Parakeet TDT or Whisper models for transcription
  • LLM: OpenAI-compatible Responses API or local Transformers models for response generation
  • TTS: Qwen3-TTS (GGML on Linux, mlx-audio on macOS) for speech synthesis

Each component accepts custom model paths through dedicated argument classes located in src/speech_to_speech/arguments_classes/, allowing you to drop in fine-tuned checkpoints using simple CLI flags.

Fine-Tuning the Speech-to-Text (STT) Component

Most STT backends in the repository are standard Hugging Face models compatible with the transformers library. You can fine-tune Whisper, Faster-Whisper, or Paraformer models using the standard training workflow, then point the pipeline to your checkpoint.

Training Whisper on Custom Data

Install the repository with the appropriate extras and prepare your dataset:

pip install "speech-to-speech[faster-whisper]"
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync

# Prepare dataset (example using Common Voice)

python -m datasets load_dataset common_voice --lang en --split train

Execute the fine-tuning using the Hugging Face Trainer:

python -m transformers.trainer \
  --model_name_or_path openai/whisper-large-v2 \
  --dataset_name ./dataset \
  --output_dir ./whisper_finetuned \
  --per_device_train_batch_size 8 \
  --learning_rate 5e-5 \
  --num_train_epochs 3 \
  --max_steps 2000

Alternatively, use a Python script for more control over the preprocessing:

from datasets import load_dataset
from transformers import WhisperProcessor, WhisperForConditionalGeneration, Trainer, TrainingArguments

dataset = load_dataset("common_voice", "en", split="train")
processor = WhisperProcessor.from_pretrained("openai/whisper-large-v2")
model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-large-v2")

def prepare_batch(batch):
    audio = batch["audio"]
    inputs = processor(audio["array"], sampling_rate=audio["sampling_rate"], return_tensors="pt")
    with processor.as_target_processor():
        labels = processor(batch["sentence"], return_tensors="pt").input_ids
    batch["input_features"] = inputs.input_features.squeeze()
    batch["labels"] = labels.squeeze()
    return batch

train_dataset = dataset.map(prepare_batch, remove_columns=dataset.column_names)

training_args = TrainingArguments(
    output_dir="./whisper_finetuned",
    per_device_train_batch_size=8,
    learning_rate=5e-5,
    num_train_epochs=3,
    fp16=True,
)

trainer = Trainer(model=model, args=training_args, train_dataset=train_dataset)
trainer.train()

Integrating the Fine-Tuned STT Model

The STT handler reads the --stt_model_name argument through src/speech_to_speech/arguments_classes/whisper_stt_arguments.py. After training completes, launch the pipeline with your fine-tuned checkpoint:

speech-to-speech \
  --stt whisper \
  --stt_model_name ./whisper_finetuned \
  --mode realtime

Fine-Tuning the Large Language Model (LLM)

The LLM component supports both remote OpenAI-compatible endpoints and locally hosted models via Transformers or mlx-lm. For local fine-tuning, use standard Hugging Face training scripts with PEFT for efficient adaptation.

Local Fine-Tuning with LoRA

Install the optional dependencies for local LLM inference:

pip install "speech-to-speech[mlx-lm]"
uv sync

Fine-tune a Qwen-3 model using LoRA:

python -m peft.trainer \
  --model_name_or_path Qwen/Qwen3-4B-Instruct-2507 \
  --dataset_name ./dialogues \
  --output_dir ./qwen_finetuned \
  --lora_rank 8 \
  --learning_rate 2e-4 \
  --num_train_epochs 2

The LLM argument handler in src/speech_to_speech/arguments_classes/module_arguments.py processes the --model_name flag. Insert your fine-tuned model using:

speech-to-speech \
  --llm_backend transformers \
  --model_name ./qwen_finetuned \
  --mode realtime

Using Fine-Tuned Endpoints

If you prefer a hosted backend, configure the pipeline to point to your fine-tuned endpoint by setting the arguments defined in src/speech_to_speech/arguments_classes/responses_api_language_model_arguments.py:

speech-to-speech \
  --responses_api_base_url https://api.your-finetuned-endpoint.com \
  --responses_api_api_key $YOUR_API_KEY \
  --mode realtime

Fine-Tuning the Text-to-Speech (TTS) Component

The default TTS backend uses Qwen3-TTS. Fine-tuning typically requires a paired audio-text dataset and the official Qwen3-TTS training scripts (external to this repository). After training, export your checkpoint and reference it via the --qwen3_tts_model_name argument handled in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py.

Install the necessary backend dependencies:

pip install "speech-to-speech[torch]"
uv sync

Execute the pipeline with your fine-tuned TTS model:

speech-to-speech \
  --tts qwen3 \
  --qwen3_tts_model_name ./qwen3_tts_finetuned \
  --mode realtime

Running the Complete Fine-Tuned Pipeline

After fine-tuning any combination of components, run the fully integrated pipeline with a single command. The CLI automatically instantiates the appropriate handlers and connects them via the queue system defined in the pipeline modules:

speech-to-speech \
  --stt whisper \
  --stt_model_name ./whisper_finetuned \
  --llm_backend transformers \
  --model_name ./qwen_finetuned \
  --tts qwen3 \
  --qwen3_tts_model_name ./qwen3_tts_finetuned \
  --mode realtime \
  --enable_live_transcription

For programmatic access, use the Python API directly:

from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments

args = ModuleArguments(
    mode="realtime",
    stt="whisper",
    stt_model_name="./whisper_finetuned",
    llm_backend="transformers",
    model_name="./qwen_finetuned",
    tts="qwen3",
    qwen3_tts_model_name="./qwen3_tts_finetuned",
    enable_live_transcription=True,
)

pipeline = SpeechToSpeechPipeline(args)
pipeline.run()  # blocks; press Ctrl-C to stop

Advanced Configuration Options

The modular architecture supports several advanced features to optimize your fine-tuned deployment:

  • Speculative Turns: Reduce latency by generating LLM output while the user is still speaking. Enable via --enable_speculative_turns (implementation in src/speech_to_speech/pipeline/speculative_turns.py).
  • Multi-Language Support: Pass --language auto and --enable_lang_prompt to ensure your fine-tuned STT and TTS models handle the target language correctly.
  • Custom VAD: Swap the default Silero VAD for a different model by implementing a new handler and registering it via src/speech_to_speech/arguments_classes/vad_arguments.py.
  • Production Deployment: Use the provided docker-compose.yml to containerize the pipeline with your fine-tuned models for scalable deployment.

Summary

  • The speech-to-speech pipeline consists of four modular components (VAD, STT, LLM, TTS) orchestrated through s2s_pipeline.py using thread-safe queues.
  • Fine-tune STT models using standard Hugging Face Trainer workflows, then specify the checkpoint with --stt_model_name.
  • Fine-tune LLM models using PEFT/LoRA or external APIs, integrating via --model_name or --responses_api_base_url.
  • Fine-tune TTS models using the Qwen3-TTS training framework, loading checkpoints with --qwen3_tts_model_name.
  • All component arguments are validated through dedicated classes in src/speech_to_speech/arguments_classes/.
  • The pipeline supports speculative turns and live transcription out-of-the-box with fine-tuned models.

Frequently Asked Questions

Can I fine-tune only one component of the speech-to-speech pipeline?

Yes. Because the architecture uses plug-and-play handlers connected via queues in queue_types.py, you can fine-tune just the STT, LLM, or TTS component independently. Simply specify the fine-tuned checkpoint path for the component you modified while keeping the default models for the others.

What hardware requirements are needed for fine-tuning?

For STT fine-tuning, you need a GPU with sufficient VRAM to hold the Whisper model (typically 8GB+ for Whisper Large). For LLM fine-tuning, using LoRA through PEFT reduces memory requirements significantly, allowing fine-tuning of 4B parameter models on consumer GPUs (8-12GB VRAM). TTS fine-tuning requirements depend on the specific Qwen3-TTS implementation.

How do I deploy my fine-tuned models in production?

Use the Docker Compose configuration provided in docker-compose.yml to containerize the pipeline. Mount your fine-tuned model directories as volumes and reference them via the appropriate CLI arguments. The WebSocket server in src/speech_to_speech/api/openai_realtime/server.py exposes an OpenAI Realtime-compatible API for production integration.

Does fine-tuning the LLM affect the latency of the pipeline?

Fine-tuning itself does not inherently increase latency if you maintain the same model architecture and quantization settings. The pipeline's speculative turns feature in speculative_turns.py can actually reduce perceived latency by generating responses before the user finishes speaking, regardless of whether the model is fine-tuned.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →