How to Prepare Data for Speech-to-Speech Training: A Complete Guide

To prepare data for speech-to-speech training, provide 16 kHz mono audio in int16 PCM format paired with clean UTF-8 transcripts, organized as a Hugging Face Dataset with optional language tags for multilingual models.

The huggingface/speech-to-speech repository implements a modular four-stage pipeline (VAD → STT → LLM → TTS) designed primarily for real-time inference. While the repository does not ship dedicated training scripts, the data contracts defined in each component's handler—located in src/speech_to_speech/s2s_pipeline.py—provide the exact specifications needed to build compatible training datasets for fine-tuning individual models or creating synthetic conversational data.

Pipeline Data Specifications

Each stage of the speech-to-speech pipeline enforces specific data formats through its respective handler classes.

Voice Activity Detection (VAD)

The VAD implementation utilizes Silero VAD v5 and expects raw audio at 16 kHz sample rate, mono channel, int16 PCM format. According to src/speech_to_speech/VAD/vad_iterator.py, input audio must be properly resampled and converted to int16 to avoid detection errors. Pre-processing should include padding with silence to prevent cutting off speech segments.

Speech-to-Text (STT)

For STT training, you need audio clips paired with exact transcripts. The WhisperSTTHandler class in src/speech_to_speech/STT/whisper_stt_handler.py loads models using AutoProcessor.from_pretrained and AutoModelForSpeechSeq2Seq.from_pretrained, expecting audio that matches the backend's requirements—typically 16 kHz, 16-bit PCM for Whisper models. The whisper_stt_arguments.py file defines the --stt_model_name CLI argument used to specify the model checkpoint.

Language Model (LLM)

The LLM stage requires UTF-8 text prompts and optional metadata such as language tags. As implemented in src/speech_to_speech/LLM/voice_prompt.py and src/speech_to_speech/LLM/text_prompt.py, the pipeline injects system prompts (voice prompts) and user text prompts without special audio formatting requirements.

Text-to-Speech (TTS)

TTS training requires text utterances paired with high-quality reference audio, optionally including speaker or language tags. The Qwen3TTSHandler in src/speech_to_speech/TTS/qwen3_tts_handler.py loads models via FasterQwen3TTS.from_pretrained and uses reference text stored in the handler configuration. The qwen3_tts_arguments.py file defines arguments like --qwen3_tts_model_name for model selection.

Step-by-Step Training Data Preparation

Follow these steps to create a training corpus compatible with the speech-to-speech pipeline components.

Collect Paired Audio-Text Data

For STT fine-tuning, each audio file must have a single transcription string from human annotators. For TTS fine-tuning, each text utterance requires a corresponding high-quality audio recording of the target voice.

Standardize Audio Format

All audio must be uniform 16-bit PCM WAV files. Use the following preprocessing function to ensure compatibility:

import librosa
import soundfile as sf

def load_resample(path, sr=16000):
    wav, _ = librosa.load(path, sr=sr, mono=True)
    wav_int16 = (wav * 32767).astype('int16')
    return wav_int16

This converts audio to the 16 kHz mono int16 format expected by the VAD and STT handlers.

Create a Hugging Face Dataset Manifest

Organize your data using the Hugging Face datasets library to ensure compatibility with the pipeline's lazy loading and audio API:

from datasets import Dataset, Audio

records = [
    {"audio": {"path": "audio/001.wav"}, "transcript": "Hello world"},
    # Additional records...

]
ds = Dataset.from_dict(records)
ds = ds.cast_column("audio", Audio(sampling_rate=16000))
ds.save_to_disk("my_s2s_dataset")

The Audio feature automatically exposes audio["array"] and audio["sampling_rate"], which the STT and TTS handlers expect.

Add Language Tags for Multilingual Models

When training multilingual models like Qwen3-TTS or Whisper, include a language field in your dataset:

{"audio": {"path": "audio/fr.wav"}, "transcript": "Bonjour le monde", "language": "fr"}

This metadata is crucial for models that require language conditioning during training.

Validate Against Pipeline Handlers

Before beginning training, validate your dataset by loading a sample through the actual handler to catch format mismatches:

from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler

handler = WhisperSTTHandler(stt_model_name="distil-whisper/distil-large-v3")
item = ds[0]
wav = item["audio"]["array"]
transcript = handler.transcribe(wav)

assert transcript.strip().lower() == item["transcript"].strip().lower()

This verification ensures your audio processing pipeline matches the expectations of whisper_stt_handler.py.

Generating Synthetic Training Data

For LLM fine-tuning, you can generate synthetic conversational data using the repository's provided script. Run the full pipeline and capture exchanges:

python scripts/synthetic_conversation_realtime_client.py \
    --stt whisper-mlx \
    --tts qwen3 \
    --llm_backend responses-api \
    --model_name "gpt-4o-mini"

This script records audio-text pairs in a JSON log suitable for training data creation. The logging logic referenced in src/speech_to_speech/api/openai_realtime/README.md leverages the same argument classes used for standard inference, ensuring data consistency.

Summary

  • Audio format: All pipeline components expect 16 kHz, mono, int16 PCM audio files.
  • Data pairing: STT requires audio with exact transcripts; TTS requires text with high-quality reference audio.
  • Dataset structure: Use Hugging Face datasets with the Audio feature to ensure handler compatibility.
  • Validation: Test samples through actual handlers like WhisperSTTHandler before training.
  • Synthetic data: Use scripts/synthetic_conversation_realtime_client.py to generate conversational training data for LLM fine-tuning.

Frequently Asked Questions

What audio format does the speech-to-speech pipeline require?

The pipeline requires 16 kHz sample rate, mono channel, int16 PCM audio format. This specification is enforced by the VAD iterator in src/speech_to_speech/VAD/vad_iterator.py and expected by Whisper STT handlers. Audio should be saved as WAV files and resampled using libraries like librosa to ensure compatibility.

Can I use the repository to train models from scratch?

The repository is designed for real-time inference and does not include dedicated training scripts. However, you can prepare training data following the data contracts in each handler (STT, TTS, LLM) and use standard training frameworks like Hugging Face transformers or trl to fine-tune the underlying models.

How do I validate my dataset before training?

Load a single item from your dataset through the corresponding handler class. For example, instantiate WhisperSTTHandler with your chosen model name, pass the audio array from your dataset, and verify that the generated transcript matches your ground truth. This catches sample rate mismatches or formatting errors early.

What is the purpose of the synthetic conversation script?

The synthetic_conversation_realtime_client.py script runs the full VAD → STT → LLM → TTS pipeline to generate synthetic conversational data. It logs the audio-text exchanges in JSON format, which you can repurpose as training data for fine-tuning the LLM component or for creating parallel corpora for other stages.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →