How to Implement Multi-Language Voice Conversations with Auto-Switching in Hugging Face Speech-to-Speech
You can achieve automatic language switching in real-time voice conversations by propagating ISO-639-1 language codes through an asynchronous VAD → STT → LLM → TTS pipeline, where the STT handler detects language using lingua-py and TTS handlers dynamically swap models when the detected language changes.
The huggingface/speech-to-speech repository provides a fully asynchronous, real-time pipeline for speech-to-speech translation and conversation. By implementing multi-language voice conversations with auto-switching, you enable seamless multilingual interactions where the system automatically detects the speaker's language and updates synthesis models without service restarts. Every pipeline message carries a language_code field, allowing each component to react to language changes independently.
Language Detection at the STT Stage
The ParakeetTDTSTTHandler in /src/speech_to_speech/STT/parakeet_tdt_handler.py (lines 79-104, 389-404) handles multilingual speech recognition and automatic language identification.
Building the Detector
The handler initializes a lingua-py detector via _build_lingua_detector during setup. It maps 25 supported European languages to lingua’s internal codes using the _LINGUA_CODE_MAP dictionary.
Detection Logic
For each completed transcription, the handler calls _detect_language_from_text, which skips utterances shorter than 20 characters to avoid false positives. For valid inputs, it returns a normalized ISO-639-1 language code that is packaged into the Transcription message.
Message Structure
The Transcription class defined in /src/speech_to_speech/pipeline/messages.py (lines 65-71) includes a language_code field that carries the detected language forward to downstream components.
Propagating Language Codes Through the Pipeline
Language codes flow through the pipeline via strongly-typed message objects, ensuring type safety and clear data contracts.
STT to LLM
The Transcription message passes its language_code to the LLM handler. The LLM can maintain the current language or explicitly request a switch via system prompts.
LLM to TTS
The LLMResponseChunk class in /src/speech_to_speech/pipeline/messages.py (lines 76-82) optionally contains a language_code field. The LMOutputProcessor generates TTSInput objects that inherit this language code, ensuring the TTS stage receives explicit language instructions.
Runtime TTS Model Switching
TTS handlers monitor the incoming language_code and reload models automatically when the language changes.
Facebook MMS Handler
The FacebookMMSTTSHandler in /src/speech_to_speech/TTS/facebookmms_handler.py (lines 65-71, 67-78) implements dynamic loading via load_model. It looks up the target language in WHISPER_LANGUAGE_TO_FACEBOOK_LANGUAGE, loads the corresponding facebook/mms-tts-* model, and updates self.language to reflect the active language.
Qwen3-TTS Handler
The Qwen3TTSHandler in /src/speech_to_speech/TTS/qwen3_tts_handler.py (lines 23-30, 61-70) normalizes language codes using _normalize_language. When the code differs from the currently loaded model, it reloads the appropriate checkpoint or switches voices to match the requested language.
Implementation Examples
Running a Demo with Auto-Language Detection
This example starts the real-time server with built-in language detection enabled by default through the Parakeet TDT handler:
# demo_auto_lang.py
from speech_to_speech.s2s_pipeline import parse_arguments, s2s_pipeline
# Use the default argument classes; the Parakeet TDT STT handler enables lingua-py.
# No extra flags are required – language detection is built-in.
args = parse_arguments() # reads CLI args or a JSON config
pipeline = s2s_pipeline(args) # creates the full pipeline object
# Start the realtime server (WebSocket) – the server prints the detected language.
pipeline.run() # blocks until Ctrl-C
Forcing a Language Switch from the Client
You can explicitly trigger a language change by sending text instructions that the LLM processes into a specific language_code:
# client_switch_lang.py
import websockets, json, asyncio
async def main():
uri = "ws://localhost:8000/realtime"
async with websockets.connect(uri) as ws:
# Send a user utterance (audio omitted for brevity)
await ws.send(json.dumps({"type": "audio_start"}))
# Later, ask the assistant to respond in Spanish
await ws.send(json.dumps({
"type": "input_text",
"content": "Por favor, habla en español a partir de ahora."
}))
# The LLM will generate a response with language_code="es",
# the TTS handler will load the Spanish model automatically.
while True:
msg = await ws.recv()
print(msg)
asyncio.run(main())
Inspecting Language Codes in Custom Handlers
Extend the pipeline with custom logic that inspects or modifies the language code before synthesis:
class MyCustomHandler(BaseHandler[TTSIn, TTSOut]):
def process(self, tts_input: TTSIn):
lang = tts_input.language_code or "en"
print(f"Generating audio for language: {lang}")
# Load / switch model here if needed …
yield from super().process(tts_input) # delegate to existing logic
Summary
- Automatic detection uses
ParakeetTDTSTTHandlerwith lingua-py to identify 25 European languages from transcribed text, skipping utterances under 20 characters. - Code propagation occurs through the
language_codefield inTranscription,LLMResponseChunk, andTTSInputmessages defined in/src/speech_to_speech/pipeline/messages.py. - Dynamic model loading happens in
FacebookMMSTTSHandlerandQwen3TTSHandler, which map ISO-639-1 codes to specific model checkpoints and reload when languages change. - Seamless switching requires no pipeline restarts; each component reacts independently to language code changes carried in pipeline messages.
Frequently Asked Questions
How does the system handle mixed-language utterances?
The ParakeetTDTSTTHandler analyzes the complete transcription text using lingua-py to determine the dominant language. Single utterances containing mixed language are classified by the majority language detected. For robust multi-language support, ensure utterances exceed the 20-character minimum threshold defined in _detect_language_from_text to avoid misclassification of short phrases.
Can I disable automatic language switching and force a specific language?
Yes. You can bypass the lingua-py detection by modifying the ParakeetTDTSTTHandler to return a hardcoded language code, or by configuring the LLM system prompt to always emit responses with a specific language_code in the LLMResponseChunk. The TTS handlers will respect the explicit code and maintain the forced language without reloading models.
What happens if a language is detected that the TTS model does not support?
The TTS handler raises a lookup error or falls back to the default language depending on implementation. In FacebookMMSTTSHandler, unsupported codes will fail lookup in WHISPER_LANGUAGE_TO_FACEBOOK_LANGUAGE, while Qwen3TTSHandler normalizes the code via _normalize_language and may map unsupported languages to the nearest available variant or default voice. Always validate language support against your specific TTS model's capabilities.
Is there a performance penalty when switching languages?
Yes, there is a brief latency spike during model loading. Both FacebookMMSTTSHandler and Qwen3TTSHandler block briefly to load new checkpoints when the language_code changes. For production deployments, consider pre-loading common language models or implementing a caching strategy to minimize the impact of frequent switches between multiple languages.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →