# How to Implement Multi-Language Voice Conversations with Auto-Switching in Hugging Face Speech-to-Speech

> Implement multi-language voice conversations with auto-switching using Hugging Face. Detect languages with lingua py and dynamically swap TTS models in real-time.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**You can achieve automatic language switching in real-time voice conversations by propagating ISO-639-1 language codes through an asynchronous VAD → STT → LLM → TTS pipeline, where the STT handler detects language using lingua-py and TTS handlers dynamically swap models when the detected language changes.**

The huggingface/speech-to-speech repository provides a fully asynchronous, real-time pipeline for speech-to-speech translation and conversation. By implementing multi-language voice conversations with auto-switching, you enable seamless multilingual interactions where the system automatically detects the speaker's language and updates synthesis models without service restarts. Every pipeline message carries a `language_code` field, allowing each component to react to language changes independently.

## Language Detection at the STT Stage

The `ParakeetTDTSTTHandler` in [`/src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main//src/speech_to_speech/STT/parakeet_tdt_handler.py) (lines 79-104, 389-404) handles multilingual speech recognition and automatic language identification.

**Building the Detector**
The handler initializes a lingua-py detector via `_build_lingua_detector` during setup. It maps 25 supported European languages to lingua’s internal codes using the `_LINGUA_CODE_MAP` dictionary.

**Detection Logic**
For each completed transcription, the handler calls `_detect_language_from_text`, which skips utterances shorter than 20 characters to avoid false positives. For valid inputs, it returns a normalized ISO-639-1 language code that is packaged into the `Transcription` message.

**Message Structure**
The `Transcription` class defined in [`/src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main//src/speech_to_speech/pipeline/messages.py) (lines 65-71) includes a `language_code` field that carries the detected language forward to downstream components.

## Propagating Language Codes Through the Pipeline

Language codes flow through the pipeline via strongly-typed message objects, ensuring type safety and clear data contracts.

**STT to LLM**
The `Transcription` message passes its `language_code` to the LLM handler. The LLM can maintain the current language or explicitly request a switch via system prompts.

**LLM to TTS**
The `LLMResponseChunk` class in [`/src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main//src/speech_to_speech/pipeline/messages.py) (lines 76-82) optionally contains a `language_code` field. The `LMOutputProcessor` generates `TTSInput` objects that inherit this language code, ensuring the TTS stage receives explicit language instructions.

## Runtime TTS Model Switching

TTS handlers monitor the incoming `language_code` and reload models automatically when the language changes.

**Facebook MMS Handler**
The `FacebookMMSTTSHandler` in [`/src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main//src/speech_to_speech/TTS/facebookmms_handler.py) (lines 65-71, 67-78) implements dynamic loading via `load_model`. It looks up the target language in `WHISPER_LANGUAGE_TO_FACEBOOK_LANGUAGE`, loads the corresponding `facebook/mms-tts-*` model, and updates `self.language` to reflect the active language.

**Qwen3-TTS Handler**
The `Qwen3TTSHandler` in [`/src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main//src/speech_to_speech/TTS/qwen3_tts_handler.py) (lines 23-30, 61-70) normalizes language codes using `_normalize_language`. When the code differs from the currently loaded model, it reloads the appropriate checkpoint or switches voices to match the requested language.

## Implementation Examples

### Running a Demo with Auto-Language Detection

This example starts the real-time server with built-in language detection enabled by default through the Parakeet TDT handler:

```python

# demo_auto_lang.py

from speech_to_speech.s2s_pipeline import parse_arguments, s2s_pipeline

# Use the default argument classes; the Parakeet TDT STT handler enables lingua-py.

# No extra flags are required – language detection is built-in.

args = parse_arguments()          # reads CLI args or a JSON config

pipeline = s2s_pipeline(args)    # creates the full pipeline object

# Start the realtime server (WebSocket) – the server prints the detected language.

pipeline.run()                    # blocks until Ctrl-C

```

### Forcing a Language Switch from the Client

You can explicitly trigger a language change by sending text instructions that the LLM processes into a specific `language_code`:

```python

# client_switch_lang.py

import websockets, json, asyncio

async def main():
    uri = "ws://localhost:8000/realtime"
    async with websockets.connect(uri) as ws:
        # Send a user utterance (audio omitted for brevity)

        await ws.send(json.dumps({"type": "audio_start"}))

        # Later, ask the assistant to respond in Spanish

        await ws.send(json.dumps({
            "type": "input_text",
            "content": "Por favor, habla en español a partir de ahora."
        }))

        # The LLM will generate a response with language_code="es",

        # the TTS handler will load the Spanish model automatically.

        while True:
            msg = await ws.recv()
            print(msg)

asyncio.run(main())

```

### Inspecting Language Codes in Custom Handlers

Extend the pipeline with custom logic that inspects or modifies the language code before synthesis:

```python
class MyCustomHandler(BaseHandler[TTSIn, TTSOut]):
    def process(self, tts_input: TTSIn):
        lang = tts_input.language_code or "en"
        print(f"Generating audio for language: {lang}")
        # Load / switch model here if needed …

        yield from super().process(tts_input)   # delegate to existing logic

```

## Summary

- **Automatic detection** uses `ParakeetTDTSTTHandler` with lingua-py to identify 25 European languages from transcribed text, skipping utterances under 20 characters.
- **Code propagation** occurs through the `language_code` field in `Transcription`, `LLMResponseChunk`, and `TTSInput` messages defined in [`/src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main//src/speech_to_speech/pipeline/messages.py).
- **Dynamic model loading** happens in `FacebookMMSTTSHandler` and `Qwen3TTSHandler`, which map ISO-639-1 codes to specific model checkpoints and reload when languages change.
- **Seamless switching** requires no pipeline restarts; each component reacts independently to language code changes carried in pipeline messages.

## Frequently Asked Questions

### How does the system handle mixed-language utterances?

The `ParakeetTDTSTTHandler` analyzes the complete transcription text using lingua-py to determine the dominant language. Single utterances containing mixed language are classified by the majority language detected. For robust multi-language support, ensure utterances exceed the 20-character minimum threshold defined in `_detect_language_from_text` to avoid misclassification of short phrases.

### Can I disable automatic language switching and force a specific language?

Yes. You can bypass the lingua-py detection by modifying the `ParakeetTDTSTTHandler` to return a hardcoded language code, or by configuring the LLM system prompt to always emit responses with a specific `language_code` in the `LLMResponseChunk`. The TTS handlers will respect the explicit code and maintain the forced language without reloading models.

### What happens if a language is detected that the TTS model does not support?

The TTS handler raises a lookup error or falls back to the default language depending on implementation. In `FacebookMMSTTSHandler`, unsupported codes will fail lookup in `WHISPER_LANGUAGE_TO_FACEBOOK_LANGUAGE`, while `Qwen3TTSHandler` normalizes the code via `_normalize_language` and may map unsupported languages to the nearest available variant or default voice. Always validate language support against your specific TTS model's capabilities.

### Is there a performance penalty when switching languages?

Yes, there is a brief latency spike during model loading. Both `FacebookMMSTTSHandler` and `Qwen3TTSHandler` block briefly to load new checkpoints when the `language_code` changes. For production deployments, consider pre-loading common language models or implementing a caching strategy to minimize the impact of frequent switches between multiple languages.