Configuring Multi-Language Detection and Switching in HuggingFace Speech-to-Speech
To enable automatic language detection and dynamic switching in the speech-to-speech pipeline, set language="auto" in WhisperSTTHandlerArguments and "auto" (or a specific ISO-639-1 code) in your TTS handler arguments—the pipeline propagates detected language codes from STT through LLM to TTS for seamless multilingual conversations.
The HuggingFace speech-to-speech library provides built-in support for multi-language detection and dynamic language switching across its processing pipeline. By configuring the STT component to auto-detect spoken language and ensuring downstream handlers respect the detected language_code, you can build applications that automatically respond in the same language as the user's input—even when languages change mid-conversation.
Understanding Language Propagation in the Pipeline
The speech-to-speech architecture uses a message-passing system where language metadata travels alongside audio and text data. This propagation happens through three distinct stages: detection in STT, optional transformation in LLM, and synthesis in TTS.
STT Language Detection with Whisper
The WhisperSTTHandlerArguments dataclass in src/speech_to_speech/arguments_classes/whisper_stt_arguments.py controls how the Whisper model handles incoming speech:
@dataclass
class WhisperSTTHandlerArguments:
stt_model_name: str = "distil-whisper/distil-large-v3"
language: str = "en" # Set to "auto" for detection
When language="auto", Whisper runs internal language detection on each utterance and returns an ISO-639-1 language code (e.g., "en", "fr", "es"). The STT handler packages this as language_code in the transcription result.
The detection happens per utterance, so switching languages mid-conversation requires no additional configuration—Whisper automatically identifies the new language on the next speech segment.
Message Types and Language Code Propagation
The pipeline's message types in src/speech_to_speech/pipeline/messages.py and src/speech_to_speech/pipeline/events.py define the language_code field that carries language metadata:
RecognizeEvent— carries STT output including detected languageUserSpeechEvent— represents user input with language annotation- Pipeline tuples — structured as
(text, language_code, tools)
When the STT handler emits a transcription, the router forwards the complete payload preserving this metadata. No manual intervention is required to pass language information between pipeline stages.
Configuring LLM and TTS for Multi-Language Output
LLM Language Control
According to src/speech_to_speech/LLM/README.md, LLM handlers can use the incoming language_code to prepend control instructions that guide the model's response language. For example, detecting "fr" might prepend "Please answer in French." to the prompt.
This approach ensures the language model generates text in the same language as the detected speech, maintaining conversational coherence without requiring separate model instances per language.
TTS Language Handling
All TTS handlers accept language configuration through their argument classes. The Qwen3TTSHandler in src/speech_to_speech/TTS/qwen3_tts_handler.py demonstrates this pattern:
# From qwen3_tts_handler.py - language normalization and passing
def _normalize_language(self, lang_code: str) -> str:
# Maps ISO-639-1 codes to model-specific format
...
def handle(self, event):
lang = event.language_code if self.language == "auto" else self.language
# Pass to synthesis model
When qwen3_tts_language="auto", the handler extracts language_code from incoming events. Setting a concrete value like "es" forces Spanish synthesis regardless of detection results.
Complete Implementation Example
This minimal configuration enables end-to-end automatic language detection and response:
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments
# STT: Enable automatic language detection
stt_args = WhisperSTTHandlerArguments(
stt_model_name="distil-whisper/distil-large-v3",
language="auto", # Key: triggers Whisper detection
)
# TTS: Accept detected language from pipeline
tts_args = Qwen3TTSArguments(
qwen3_tts_language="auto", # Uses incoming language_code
)
# Build and run pipeline
pipeline = SpeechToSpeechPipeline(
stt_handler_kwargs=stt_args,
tts_handler_kwargs=tts_args,
)
pipeline.run() # Processes microphone input with auto language switching
Execution flow:
- User speaks in French → Whisper detects
language_code="fr" - Pipeline propagates
"fr"through to TTS handler - Qwen-3-TTS synthesizes response in French
- User switches to Spanish → Whisper detects
language_code="es"on next utterance - TTS automatically switches to Spanish synthesis
Advanced Configuration Patterns
Force Fixed TTS Language
To override detection and always synthesize in a specific language:
tts_args = Qwen3TTSArguments(
qwen3_tts_language="de", # Always German, ignores STT detection
)
Runtime Language Switching
Call handler-specific methods to change language mid-session:
# Force TTS to new language while STT continues auto-detection
pipeline.set_tts_language("it") # Italian synthesis
The STT handler continues detecting actual spoken language, but responses synthesize in Italian until reconfigured.
Per-Handler Language Overrides
Different handlers in the same pipeline can use different language strategies:
| Component | Configuration | Behavior |
|---|---|---|
| STT | language="auto" |
Detect actual spoken language |
| LLM | language="auto" (default) |
Follow STT detection with instruction prepending |
| TTS | language="fr" |
Force French output regardless |
Key Source Files for Multi-Language Configuration
| File Path | Purpose |
|---|---|
src/speech_to_speech/arguments_classes/whisper_stt_arguments.py |
STT language detection settings |
src/speech_to_speech/TTS/qwen3_tts_handler.py |
TTS language acceptance and normalization |
src/speech_to_speech/pipeline/messages.py |
Message type definitions with language_code |
src/speech_to_speech/LLM/README.md |
LLM language control documentation |
Summary
- Enable detection: Set
language="auto"inWhisperSTTHandlerArgumentsto activate Whisper's per-utterance language detection - Propagate codes: The pipeline automatically forwards
language_codefrom STT throughRecognizeEventand related message types - Configure TTS: Use
"auto"to follow detection or specify ISO-639-1 codes for forced output languages - Switch dynamically: Detection runs on every utterance; TTS can be reconfigured at runtime via handler methods
- Control LLM output: The LLM handler prepends language instructions based on incoming codes to ensure response language matching
Frequently Asked Questions
How does automatic language detection work in the speech-to-speech pipeline?
Whisper STT runs language identification on each audio segment when language="auto" is configured, returning ISO-639-1 codes like "en" or "fr". The pipeline's message system carries this language_code through to downstream components without requiring manual handling.
Can I force a specific output language while keeping input detection?
Yes. Configure STT with language="auto" for detection, but set your TTS handler arguments to a concrete language code like qwen3_tts_language="es". The system detects the input language for logging or LLM context, but always synthesizes in the specified TTS language.
What happens if someone switches languages mid-conversation?
Whisper detects the new language on the next utterance and the pipeline propagates the updated language_code. If TTS is configured with "auto", synthesis switches automatically. No pipeline restart or code changes are required.
Which TTS handlers support automatic language switching?
The Qwen3TTSHandler explicitly implements this pattern with its qwen3_tts_language argument. Other TTS handlers in the repository follow similar conventions—check their respective argument classes for language or language_code configuration options.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →