How to Convert Speech from One Language to Another Using the Hugging Face Speech-to-Speech Library
You can convert speech from one language to another by configuring the S2SPipeline with a Whisper STT handler for source detection, a language model handler for translation, and a Qwen-3 TTS handler for synthesis, passing the target language_code through runtime_config to propagate language metadata across all pipeline stages.
The Hugging Face speech-to-speech repository provides an open-source framework for real-time speech translation. By orchestrating three specialized handlers through the S2SPipeline class defined in src/speech_to_speech/s2s_pipeline.py, you can convert spoken audio from a detected source language to synthesized speech in a specified target language with minimal latency.
The Three-Stage Translation Architecture
The library implements a full-duplex pipeline that processes audio through three sequential stages, each communicating via asynchronous queues and typed message objects from src/speech_to_speech/pipeline/messages.py.
Speech-to-Text (STT): The WhisperSTTHandler in src/speech_to_speech/STT/whisper_stt_handler.py transcribes incoming audio into a Transcription message. When configured with language="auto", Whisper detects the source language and stores the ISO code (e.g., "en", "fr") in handler.last_language, which attaches to the message's language_code field.
Language Model Translation: The LanguageModelHandler (utilizing src/speech_to_speech/LLM/language_model.py) receives the transcription and generates a translation. It emits GenerateResponseRequest and LLMResponseChunk messages carrying the target language_code specified in your configuration, effectively translating content while updating language metadata for downstream handlers.
Text-to-Speech (Synthesis): The Qwen3TTSHandler in src/speech_to_speech/TTS/qwen3_tts_handler.py consumes the translated text along with the target language_code, synthesizing audio that matches the desired output language and packaging it into an AudioOutput message.
Configuring Source and Target Languages
Language flow is controlled through the runtime_config dictionary passed to S2SPipeline. The pipeline distinguishes between automatic source detection and explicit target specification.
Source Language Detection
When instantiating WhisperSTTArguments, set language="auto" to enable automatic language detection. The handler stores the detected code in handler.last_language, which propagates to the Transcription message's language_code attribute. For explicit control, pass a specific ISO code such as language="de" to skip detection and force German recognition.
Target Language Specification
Set the language_code key in your runtime_config to the desired target ISO code (e.g., "es" for Spanish). This value flows to the LanguageModelHandler, instructing it to generate LLMResponseChunk objects in that language, and subsequently to the TTS handler for voice synthesis.
Implementation Examples
High-Level Pipeline Configuration
The following example demonstrates configuring the pipeline for English-to-Spanish translation using the argument classes and S2SPipeline:
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelArguments
# Configure runtime with target language
runtime_config = {
"language_code": "es", # Target: Spanish
"stt": WhisperSTTArguments(language="auto"), # Auto-detect source
"tts": Qwen3TTSArguments(),
"lm": LanguageModelArguments(),
}
pipeline = S2SPipeline(runtime_config)
pipeline.run() # Starts listening and speaking loop
Programmatic Audio Processing
For batch processing or custom audio sources, feed raw bytes directly without using the microphone loop:
# Initialize pipeline (English input → French output)
config = {
"language_code": "fr",
"stt": WhisperSTTArguments(language="en"),
"tts": Qwen3TTSArguments(),
"lm": LanguageModelArguments(),
}
pipeline = S2SPipeline(config)
# Process audio file
with open("input.wav", "rb") as f:
audio_chunk = f.read()
pipeline.feed_audio(audio_chunk)
output_audio = pipeline.get_output_audio()
Fixed Source Language Configuration
To bypass automatic detection and reduce latency when the source language is known:
# Explicit German to English translation
runtime_config = {
"language_code": "en",
"stt": WhisperSTTArguments(language="de"), # Fixed source
"tts": Qwen3TTSArguments(),
"lm": LanguageModelArguments(),
}
pipeline = S2SPipeline(runtime_config)
pipeline.run()
Key Source Files and Components
Understanding these core files helps debug and extend the translation pipeline:
src/speech_to_speech/s2s_pipeline.py: Contains theS2SPipelineclass that initializes handler queues and manages the event loop.src/speech_to_speech/pipeline/messages.py: Defines message types (Transcription,TTSInput,AudioOutput) that carrylanguage_codemetadata between stages.src/speech_to_speech/pipeline/events.py: Implements typed events likeTranscriptionCompletedEventthat trigger stage transitions.src/speech_to_speech/STT/whisper_stt_handler.py: Handles speech recognition and source language detection.src/speech_to_speech/LLM/language_model.py: Provides translation capabilities via theLanguageModelHandler.src/speech_to_speech/TTS/qwen3_tts_handler.py: Synthesizes speech for the target language.scripts/listen_and_play.py: Reference implementation demonstrating microphone input to speaker output.
Summary
- The
S2SPipelineorchestrates three stages (STT, LM, TTS) to convert speech between languages. - Source language is detected automatically by Whisper or set explicitly via
WhisperSTTArguments. - Target language is specified through
runtime_config["language_code"]and propagated through message objects. - Handlers normalize language aliases (e.g.,
"English"to"en") and fallback tolast_languageif unsupported codes are encountered. - The pipeline supports both real-time streaming and programmatic batch processing.
Frequently Asked Questions
How does the pipeline detect the source language?
The WhisperSTTHandler uses OpenAI's Whisper model with language="auto" to detect the spoken language from audio features. It stores the ISO 639-1 code in the language_code field of the Transcription message, which subsequent handlers receive for processing.
Can I use different models for STT or TTS?
Yes. While the examples use WhisperSTTHandler and Qwen3TTSHandler, the pipeline architecture accepts any handler implementing the base interfaces. Replace WhisperSTTArguments with your custom STT arguments class or substitute the TTS handler by modifying the tts key in runtime_config to point to alternative implementations.
What happens if I specify an unsupported target language?
The handlers normalize language aliases and validate against supported codes. If an unsupported code is encountered, the system falls back to the last_language attribute stored in the handler state, ensuring the pipeline continues processing rather than failing.
Is this processing suitable for real-time conversations?
Yes. The S2SPipeline uses asynchronous queues and streaming handlers to minimize latency between speech input and audio output. The scripts/listen_and_play.py example demonstrates this by processing microphone input and playing synthesized responses with minimal delay, suitable for interactive dialogue.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →