# How to Prepare Data for Speech-to-Speech Training: A Complete Guide

> Learn how to prepare data for speech-to-speech training. Provide 16 kHz mono audio with UTF-8 transcripts organized as a Hugging Face Dataset for optimal model performance.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**To prepare data for speech-to-speech training, provide 16 kHz mono audio in int16 PCM format paired with clean UTF-8 transcripts, organized as a Hugging Face Dataset with optional language tags for multilingual models.**

The `huggingface/speech-to-speech` repository implements a modular four-stage pipeline (VAD → STT → LLM → TTS) designed primarily for real-time inference. While the repository does not ship dedicated training scripts, the data contracts defined in each component's handler—located in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)—provide the exact specifications needed to build compatible training datasets for fine-tuning individual models or creating synthetic conversational data.

## Pipeline Data Specifications

Each stage of the speech-to-speech pipeline enforces specific data formats through its respective handler classes.

### Voice Activity Detection (VAD)

The VAD implementation utilizes Silero VAD v5 and expects raw audio at **16 kHz sample rate, mono channel, int16 PCM format**. According to [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py), input audio must be properly resampled and converted to `int16` to avoid detection errors. Pre-processing should include padding with silence to prevent cutting off speech segments.

### Speech-to-Text (STT)

For STT training, you need audio clips paired with **exact transcripts**. The `WhisperSTTHandler` class in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) loads models using `AutoProcessor.from_pretrained` and `AutoModelForSpeechSeq2Seq.from_pretrained`, expecting audio that matches the backend's requirements—typically 16 kHz, 16-bit PCM for Whisper models. The [`whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_arguments.py) file defines the `--stt_model_name` CLI argument used to specify the model checkpoint.

### Language Model (LLM)

The LLM stage requires **UTF-8 text prompts** and optional metadata such as language tags. As implemented in [`src/speech_to_speech/LLM/voice_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/voice_prompt.py) and [`src/speech_to_speech/LLM/text_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/text_prompt.py), the pipeline injects system prompts (voice prompts) and user text prompts without special audio formatting requirements.

### Text-to-Speech (TTS)

TTS training requires **text utterances paired with high-quality reference audio**, optionally including speaker or language tags. The `Qwen3TTSHandler` in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) loads models via `FasterQwen3TTS.from_pretrained` and uses reference text stored in the handler configuration. The [`qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_arguments.py) file defines arguments like `--qwen3_tts_model_name` for model selection.

## Step-by-Step Training Data Preparation

Follow these steps to create a training corpus compatible with the speech-to-speech pipeline components.

### Collect Paired Audio-Text Data

For **STT fine-tuning**, each audio file must have a single transcription string from human annotators. For **TTS fine-tuning**, each text utterance requires a corresponding high-quality audio recording of the target voice.

### Standardize Audio Format

All audio must be uniform 16-bit PCM WAV files. Use the following preprocessing function to ensure compatibility:

```python
import librosa
import soundfile as sf

def load_resample(path, sr=16000):
    wav, _ = librosa.load(path, sr=sr, mono=True)
    wav_int16 = (wav * 32767).astype('int16')
    return wav_int16

```

This converts audio to the 16 kHz mono int16 format expected by the VAD and STT handlers.

### Create a Hugging Face Dataset Manifest

Organize your data using the Hugging Face `datasets` library to ensure compatibility with the pipeline's lazy loading and audio API:

```python
from datasets import Dataset, Audio

records = [
    {"audio": {"path": "audio/001.wav"}, "transcript": "Hello world"},
    # Additional records...

]
ds = Dataset.from_dict(records)
ds = ds.cast_column("audio", Audio(sampling_rate=16000))
ds.save_to_disk("my_s2s_dataset")

```

The `Audio` feature automatically exposes `audio["array"]` and `audio["sampling_rate"]`, which the STT and TTS handlers expect.

### Add Language Tags for Multilingual Models

When training multilingual models like Qwen3-TTS or Whisper, include a `language` field in your dataset:

```python
{"audio": {"path": "audio/fr.wav"}, "transcript": "Bonjour le monde", "language": "fr"}

```

This metadata is crucial for models that require language conditioning during training.

### Validate Against Pipeline Handlers

Before beginning training, validate your dataset by loading a sample through the actual handler to catch format mismatches:

```python
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler

handler = WhisperSTTHandler(stt_model_name="distil-whisper/distil-large-v3")
item = ds[0]
wav = item["audio"]["array"]
transcript = handler.transcribe(wav)

assert transcript.strip().lower() == item["transcript"].strip().lower()

```

This verification ensures your audio processing pipeline matches the expectations of [`whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_handler.py).

## Generating Synthetic Training Data

For LLM fine-tuning, you can generate synthetic conversational data using the repository's provided script. Run the full pipeline and capture exchanges:

```bash
python scripts/synthetic_conversation_realtime_client.py \
    --stt whisper-mlx \
    --tts qwen3 \
    --llm_backend responses-api \
    --model_name "gpt-4o-mini"

```

This script records audio-text pairs in a JSON log suitable for training data creation. The logging logic referenced in [`src/speech_to_speech/api/openai_realtime/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/README.md) leverages the same argument classes used for standard inference, ensuring data consistency.

## Summary

- **Audio format**: All pipeline components expect 16 kHz, mono, int16 PCM audio files.
- **Data pairing**: STT requires audio with exact transcripts; TTS requires text with high-quality reference audio.
- **Dataset structure**: Use Hugging Face `datasets` with the `Audio` feature to ensure handler compatibility.
- **Validation**: Test samples through actual handlers like `WhisperSTTHandler` before training.
- **Synthetic data**: Use [`scripts/synthetic_conversation_realtime_client.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/synthetic_conversation_realtime_client.py) to generate conversational training data for LLM fine-tuning.

## Frequently Asked Questions

### What audio format does the speech-to-speech pipeline require?

The pipeline requires **16 kHz sample rate, mono channel, int16 PCM** audio format. This specification is enforced by the VAD iterator in [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py) and expected by Whisper STT handlers. Audio should be saved as WAV files and resampled using libraries like librosa to ensure compatibility.

### Can I use the repository to train models from scratch?

The repository is designed for real-time inference and does not include dedicated training scripts. However, you can prepare training data following the data contracts in each handler (STT, TTS, LLM) and use standard training frameworks like Hugging Face `transformers` or `trl` to fine-tune the underlying models.

### How do I validate my dataset before training?

Load a single item from your dataset through the corresponding handler class. For example, instantiate `WhisperSTTHandler` with your chosen model name, pass the audio array from your dataset, and verify that the generated transcript matches your ground truth. This catches sample rate mismatches or formatting errors early.

### What is the purpose of the synthetic conversation script?

The [`synthetic_conversation_realtime_client.py`](https://github.com/huggingface/speech-to-speech/blob/main/synthetic_conversation_realtime_client.py) script runs the full VAD → STT → LLM → TTS pipeline to generate synthetic conversational data. It logs the audio-text exchanges in JSON format, which you can repurpose as training data for fine-tuning the LLM component or for creating parallel corpora for other stages.