# Can I Use Custom Datasets for Training Speech-to-Speech Models? A Deep Dive into the Hugging Face Repository

> Cannot train speech-to-speech models on custom datasets with huggingface/speech-to-speech. This repository is for inference only, lacking a training pipeline. Learn more about its limitations.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-08-02

---

**No—you cannot train or fine-tune models directly on custom datasets using the huggingface/speech-to-speech repository, as it is strictly an inference framework with no training pipeline.**

The `speech-to-speech` repository provides real-time speech-to-speech translation and voice conversion by orchestrating pre-trained models. Its architecture is designed around loading existing checkpoints via `from_pretrained` APIs, not iterating over training data or performing optimizer steps. If you need to adapt models to your own audio-text pairs, you must fine-tune them elsewhere and load the resulting checkpoints here.

## Why the Repository Doesn't Support Training

A thorough examination of the source code confirms there is **no training infrastructure** anywhere in the codebase. The repository lacks:

- Any `torch.utils.data.DataLoader` loops
- Optimizer initialization or `backward()` calls
- `Trainer` or `Accelerate` integration
- Loss computation for model updates

The only data-related code handles **voice activity detection configuration** ([`VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/VAD/vad_iterator.py)) and **evaluation datasets** in test suites—not training or fine-tuning.

In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the `SpeechToSpeechPipeline` class orchestrates components through a simple inference chain: STT → (optional LLM) → TTS. It never exposes hooks for gradient updates or dataset iteration.

## The Correct Workflow: Train Elsewhere, Load Here

Since custom dataset training happens **outside** this repository, follow this two-step process:

### Step 1: Fine-Tune the Underlying Model

Use the original training scripts from each model's source repository:

| Model | Training Resource | Typical Fine-Tuning Data |
|-------|-------------------|--------------------------|
| **Whisper** | `transformers` examples | Audio-transcript pairs |
| **Qwen-3-TTS** | Model-specific library | Text-audio pairs |
| **Facebook MMS** | `fairseq` recipes | Multilingual speech |

### Step 2: Load Your Fine-Tuned Checkpoint

All handlers in this repository support loading any Hugging Face-compatible checkpoint through `from_pretrained`:

```python
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler

# Load your custom fine-tuned Whisper model

stt = WhisperSTTHandler.from_pretrained(
    model_name="my-org/whisper-finetuned-custom-dataset",
    language="en",
    device="cuda"
)

```

This pattern applies identically to TTS handlers in `src/speech_to_speech/TTS/`.

## Complete Inference Example with Custom Checkpoints

After fine-tuning on your dataset externally, wire everything together:

```python
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline

# 1. Load fine-tuned STT (trained on custom dataset elsewhere)

stt = WhisperSTTHandler.from_pretrained(
    model_name="organization/my-custom-whisper",
    language="en",
    device="cuda"
)

# 2. Load fine-tuned TTS (trained on custom dataset elsewhere)

tts = Qwen3TTSHandler.from_pretrained(
    model_name="organization/my-custom-qwen3-tts",
    device="cuda"
)

# 3. Assemble inference pipeline—no training occurs here

pipeline = SpeechToSpeechPipeline(
    stt_handler=stt,
    tts_handler=tts
)

# 4. Run on new audio

result = pipeline.run(input_wav_path="input.wav")

```

The `from_pretrained` method in [`src/speech_to_speech/STT/base_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/base_stt_handler.py) and its TTS equivalent delegate directly to Hugging Face's `AutoModel` classes, ensuring compatibility with any properly exported checkpoint.

## Key Source Files Confirming Inference-Only Design

| File Path | Role in Architecture |
|-----------|----------------------|
| [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) | Orchestrates STT→LLM→TTS chain with no training hooks |
| [`src/speech_to_speech/STT/base_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/base_stt_handler.py) | Abstract base defining `from_pretrained` loading only |
| [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) | Implements Whisper via `AutoProcessor`/`AutoModelForSpeechSeq2Seq` |
| [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) | Loads Qwen-3-TTS through `FasterQwen3TTS.from_pretrained` |
| [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py) | Speech segmentation pre-processing, no training logic |

These files collectively demonstrate the repository's exclusive focus on **model loading and execution**, not **parameter optimization**.

## What About On-the-Fly Adaptation?

The repository offers no mechanism for:

- **Few-shot learning** from new audio samples
- **Online adaptation** to speaker characteristics
- **Continual fine-tuning** during inference

Any adaptation must be performed in separate training runs using external tools, then redeployed via updated checkpoint paths.

## Summary

- The `speech-to-speech` repository is **inference-only** with no training pipeline for custom datasets
- Fine-tune models using their native frameworks (Transformers, Fairseq, etc.), then load checkpoints via `from_pretrained`
- All handlers support standard Hugging Face checkpoint loading—just point to your fine-tuned model ID or local path
- The architecture in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) is strictly real-time inference with no backward pass or data loader

## Frequently Asked Questions

### Can I add a training script to the speech-to-speech repository?

You technically could fork and extend it, but this would violate the repository's modular design philosophy. The intended approach is to train components independently (Whisper, Qwen-3-TTS, etc.) and use this repository strictly for orchestration. Training scripts for individual models are already mature in their respective source repositories.

### How do I know if my fine-tuned checkpoint is compatible?

Any checkpoint that works with `transformers.AutoModel.from_pretrained()` or the model-specific equivalent will load correctly. Test with: `WhisperSTTHandler.from_pretrained("your-model-id", device="cpu")`. If this succeeds without errors, the pipeline will accept it.

### Does the repository support domain adaptation without full fine-tuning?

No built-in support exists for parameter-efficient methods like LoRA adapters or prompt tuning. You would need to merge adapters into full weights externally, then load the merged checkpoint. The handlers currently instantiate complete model objects without adapter hooks.

### What evaluation capabilities exist for custom datasets?

The test suite uses Hugging Face `datasets` for **benchmark evaluation**, not training. You can run inference over evaluation sets and compute metrics externally, but no built-in metric computation or loss logging exists in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) or related files.