How to Convert Speech from One Language to Another with Hugging Face Speech-to-Speech
The speech-to-speech library converts spoken input to spoken output in a different language through a three-stage pipeline: Speech-to-Text (STT) detects the source language, a Language Model translates the text, and Text-to-Speech (TTS) synthesizes the target language audio—all coordinated by S2SPipeline in s2s_pipeline.py.
This guide walks through the complete process of building a multilingual speech translator using the huggingface/speech-to-speech repository. The library provides a modular, full-duplex pipeline that chains together STT, translation, and TTS components with automatic language code propagation.
Understanding the Multilingual Pipeline Architecture
The speech-to-speech translation flow consists of three core stages that communicate via typed messages carrying language_code fields:
| Stage | Handler Location | Function |
|---|---|---|
| Speech-to-Text | src/speech_to_speech/STT/whisper_stt_handler.py |
Captures audio, runs Whisper, emits Transcription with detected source language_code |
| Language Model / Translation | src/speech_to_speech/LLM/language_model.py via LanguageModelHandler |
Receives transcription, generates translation, returns LLMResponseChunk with target language_code |
| Text-to-Speech | src/speech_to_speech/TTS/qwen3_tts_handler.py |
Synthesizes audio using the target language_code from the LM stage |
The S2SPipeline class in src/speech_to_speech/s2s_pipeline.py orchestrates these stages by creating async queues and wiring handlers together. It automatically propagates language metadata through each message type defined in src/speech_to_speech/pipeline/messages.py.
Language Handling Mechanics
Source Language Detection
When WhisperSTTHandler receives --language auto, Whisper detects the spoken language and stores it in handler.last_language. This value attaches to the Transcription message as language_code (e.g., "en", "fr", "de").
Target Language Selection
The target language flows through the pipeline via runtime_config.language_code. When LanguageModelHandler builds a GenerateResponseRequest, the language_code field signals which language the LLM should generate. Handlers normalize aliases (e.g., "English" → "en"; "nb" → "no") and fall back to the last valid language if unsupported.
Running Speech-to-Speech Translation
Command-Line Approach with listen_and_play.py
The repository includes scripts/listen_and_play.py as a complete working example:
# scripts/listen_and_play.py
import argparse
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelArguments
parser = argparse.ArgumentParser(description="Speech-to-speech translation")
parser.add_argument("--source-language", default="auto", help="Source language (auto for detection)")
parser.add_argument("--target-language", required=True, help="Target language code, e.g. 'fr' or 'de'")
args = parser.parse_args()
runtime_config = {
"language_code": args.target_language, # Target language propagated to LM and TTS
"stt": WhisperSTTArguments(language=args.source_language),
"tts": Qwen3TTSArguments(),
"lm": LanguageModelArguments(),
}
pipeline = S2SPipeline(runtime_config)
pipeline.run() # Starts interactive microphone-to-speaker loop
The pipeline executes as follows:
- WhisperSTTHandler detects source language and emits
Transcription(language_code="en") - LanguageModelHandler prompts the LLM to translate, returns
LLMResponseChunk(language_code="fr") - Qwen3TTSHandler synthesizes French speech using the target
language_code
Programmatic API Without CLI
For integration into larger applications, instantiate S2SPipeline directly:
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelArguments
# Configure for Spanish output
config = {
"language_code": "es",
"stt": WhisperSTTArguments(language="auto"),
"tts": Qwen3TTSArguments(),
"lm": LanguageModelArguments(),
}
pipeline = S2SPipeline(config)
# Feed audio from file, microphone, or stream
with open("input_english.wav", "rb") as f:
audio_bytes = f.read()
pipeline.feed_audio(audio_bytes)
# Retrieve translated audio
translated_audio = pipeline.get_output_audio()
Fixed Source Language (Skip Auto-Detection)
For known input languages, disable detection to reduce latency:
config["stt"] = WhisperSTTArguments(language="de") # German input
config["language_code"] = "en" # English output
pipeline = S2SPipeline(config)
Key Source Files for Language Conversion
| File | Path | Purpose |
|---|---|---|
| Pipeline orchestrator | src/speech_to_speech/s2s_pipeline.py |
S2SPipeline class creates queues and runs the async event loop |
| Message types | src/speech_to_speech/pipeline/messages.py |
Transcription, TTSInput, AudioOutput with language_code fields |
| Event types | src/speech_to_speech/pipeline/events.py |
TranscriptionCompletedEvent and other pipeline events |
| Whisper STT | src/speech_to_speech/STT/whisper_stt_handler.py |
Speech recognition with language detection |
| Qwen-3 TTS | src/speech_to_speech/TTS/qwen3_tts_handler.py |
Speech synthesis respecting target language_code |
| LLM translation | src/speech_to_speech/LLM/language_model.py |
Translation via LanguageModelHandler |
| Demo script | scripts/listen_and_play.py |
Complete end-to-end example |
Swappable Backend Components
The modular design allows mixing STT, LM, and TTS backends. Replace argument classes to use alternatives:
- STT:
WhisperSTTArguments, or custom handlers insrc/speech_to_speech/STT/ - LM: Any OpenAI-compatible API via
LanguageModelArguments, or local models - TTS:
Qwen3TTSArguments, or other handlers insrc/speech_to_speech/TTS/
All handlers respect the language_code field in their input messages, ensuring consistent language propagation regardless of backend choice.
Summary
- The speech-to-speech library implements translation through a three-stage pipeline coordinated by
S2SPipelineinsrc/speech_to_speech/s2s_pipeline.py - Source language is detected by Whisper or specified explicitly; stored in
Transcription.language_code - Target language is set via
runtime_config.language_codeand propagated through LM and TTS stages - Message types in
src/speech_to_speech/pipeline/messages.pycarrylanguage_codefields ensuring metadata flows through the pipeline scripts/listen_and_play.pyprovides a complete CLI example for interactive translation- Backend components are swappable—mix Whisper STT, any OpenAI-compatible LLM, and Qwen-3 TTS or alternatives
Frequently Asked Questions
What language codes does the speech-to-speech library support?
The library uses standard ISO 639-1 codes ("en", "fr", "de", "es", "zh", etc.). Handlers in whisper_stt_handler.py and qwen3_tts_handler.py normalize aliases automatically, mapping "English" to "en" and "nb" (Norwegian Bokmål) to "no". Unsupported codes trigger fallback to the last valid language.
Can I use a different translation model instead of the default LLM?
Yes. The LanguageModelArguments class accepts any OpenAI-compatible API endpoint. Configure the base_url and model parameters to switch to local models (via vLLM, llama.cpp) or alternative APIs. The LanguageModelHandler builds translation prompts and returns LLMResponseChunk with the target language_code regardless of backend.
How does the pipeline handle language detection failures?
If Whisper cannot determine the source language, whisper_stt_handler.py falls back to handler.last_language (previously detected) or a configured default. The language_code field is never empty—handlers ensure a valid code propagates to prevent TTS synthesis errors.
Is real-time streaming translation supported?
Yes. The S2SPipeline uses asynchronous queues between stages, enabling chunked processing. scripts/listen_and_play.py demonstrates real-time microphone input with immediate speaker output. For lower latency, disable auto-detection with a fixed source language and use smaller Whisper models.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →