# Multi-Language Support and Automatic Language Detection in Hugging Face Speech-to-Speech

> Explore multi-language support and automatic language detection in Hugging Face Speech-to-Speech. Learn how Whisper and Parakeet TDT identify languages.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**The Hugging Face speech-to-speech framework enables automatic language detection through model-embedded tokens in Whisper handlers and the lingua-py library in Parakeet TDT, while TTS handlers use alias mappings to convert ISO codes to model-specific identifiers.**

The huggingface/speech-to-speech repository delivers end-to-end voice processing with comprehensive multi-language support across both Speech-to-Text (STT) and Text-to-Speech (TTS) pipelines. The system handles **automatic language detection** to power seamless multilingual conversations, intelligently routing detected language codes from transcription through synthesis without manual configuration. Whether processing 25 European languages with Parakeet TDT or leveraging Whisper's embedded language tokens, the framework maintains language context throughout the entire pipeline.

## Multi-Language Support in STT Handlers

Each STT handler exposes a `SUPPORTED_LANGUAGES` constant that defines its transcription capabilities. When you set `language="auto"` or omit the language parameter, the handler activates its specific detection mechanism to identify the spoken language on-the-fly.

### Parakeet TDT Handler

The **Parakeet TDT handler** supports 25 European languages including English, German, French, and Lithuanian, defined in the `SUPPORTED_LANGUAGES` constant within [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) (lines 40-68). For automatic detection, it utilizes the **lingua-py** library rather than the model itself.

At module initialization, the handler builds a pre-loaded detector via `_build_lingua_detector()` to eliminate warm-up latency during the first request. The `_detect_language_from_text()` method (lines 93-104) invokes this detector only on utterances exceeding 20 characters to avoid false positives on short phrases. The implementation also remaps certain codes—for example, converting Norwegian (`no`) to `nb` to match the handler's supported list.

### Whisper STT Handler

The **Whisper STT handler** supports 12 languages (English, French, Spanish, Chinese, Japanese, Korean, Hindi, German, Portuguese, Polish, Italian, and Dutch) as defined in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) (lines 19-32). Rather than using an external library, Whisper performs **automatic language detection** through model-embedded language tokens.

The handler extracts the language code directly from the first generated token sequence:

```python
language_code = self.processor.tokenizer.decode(pred_ids[0, 1])[2:-2]

```

If the detected code is not present in `SUPPORTED_LANGUAGES`, the system falls back to the `last_language` attribute and re-runs the request with the previously known language. When operating in auto mode, the detected code is appended with `-auto` to signal downstream components that the language was inferred rather than specified (line 38).

### MLX-Audio Whisper Handler

The **MLX-Audio Whisper handler** for Apple Silicon follows the same pattern as the standard Whisper implementation, relying on the model's internal language token for detection. It maintains an identical `SUPPORTED_LANGUAGES` structure with `DEFAULT_LANGUAGE = "en"`, implemented in [`src/speech_to_speech/STT/mlx_audio_whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/mlx_audio_whisper_handler.py).

## How Automatic Language Detection Works

The framework implements two distinct architectural approaches for language identification, depending on the STT backend in use.

### External Linguistic Detection with Lingua-Py

The Parakeet TDT handler implements a two-stage detection process using the lingua-py library. During startup, the system initializes `_lingua_detector` with pre-loaded language models to prevent latency during the first transcription request.

When processing text, the `_detect_language_from_text()` method applies a length threshold to ensure accuracy:

```python
if len(text) > 20:
    detected = _lingua_detector.detect_language_of(text)
    code = detected.iso_code_639_1.name.lower()
    # Remap codes to match SUPPORTED_LANGUAGES

    return {v: k for k, v in _LINGUA_CODE_MAP.items()}.get(code, code)

```

This approach isolates language detection from the speech recognition model, allowing the Parakeet TDT model to focus on transcription while the linguistic detector handles language classification.

### Model-Embedded Token Extraction

Whisper-based handlers leverage the model's inherent multilingual training. Whisper models generate a language token (e.g., `<|en|>`, `<|fr|>`) as the first output token during the generation process. The handler extracts this token by decoding the second token ID in the prediction sequence (index 1) and stripping the delimiters:

```python

# pred_ids[0, 1] contains the language token

language_code = self.processor.tokenizer.decode(pred_ids[0, 1])[2:-2]

```

This method provides immediate language identification without additional inference overhead, as the token is produced as part of the standard transcription generation process.

## Multi-Language Support in TTS Handlers

Unlike STT handlers, TTS handlers do not perform **automatic language detection** from audio. Instead, they receive language codes from the STT stage or user input and translate standard ISO-639-1 codes into model-specific identifiers through alias mappings.

### Language Alias Mappings

Each TTS handler maintains a dictionary that maps standard language codes to the strings required by its underlying model:

- **Qwen-3 TTS**: Uses `QWEN3_LANGUAGE_ALIASES` in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) (lines 55-76) to map codes like `en` to `english`, `zh-cn` to `chinese`, and `es` to `spanish`.

- **Kokoro TTS**: Implements `WHISPER_LANGUAGE_TO_KOKORO_LANG` in [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py) to convert Whisper language tokens to Kokoro's speaker language codes.

- **Facebook MMS TTS**: Utilizes `WHISPER_LANGUAGE_TO_FACEBOOK_LANGUAGE` in [`src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/facebookmms_handler.py) for similar conversions.

These mappings ensure that when the STT handler detects `fr` (French), the TTS handler can translate this to the specific string format required by the voice synthesis model.

## Practical Implementation Examples

### Enabling Automatic Detection with Parakeet TDT

To process audio with automatic language detection using the Parakeet handler:

```python
from speech_to_speech import SpeechToSpeechPipeline

pipeline = SpeechToSpeechPipeline(
    stt_handler="parakeet_tdt",
    language="auto",  # Enables lingua-py detection

)

transcriptions = pipeline.process(audio_np_array)
for t in transcriptions:
    print(f"Text: {t.text}, Language: {t.language_code}")

```

This configuration triggers the linguistic detector in [`parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/parakeet_tdt_handler.py), which analyzes transcriptions longer than 20 characters to determine the language.

### Whisper with Fallback Behavior

For Whisper-based transcription with automatic detection and fallback:

```python
pipeline = SpeechToSpeechPipeline(
    stt_handler="whisper",
    language="auto",
)

for output in pipeline.process(audio_chunk):
    print(f"{output.language_code}: {output.text}")

```

If Whisper emits an unsupported language token, the handler automatically retries the request using `self.last_language` to maintain conversation continuity.

### TTS with Language Mapping

To synthesize speech in a specific language using Qwen-3 TTS:

```python
pipeline = SpeechToSpeechPipeline(
    tts_handler="qwen3_tts",
    language="es",  # Spanish ISO-639-1 code

)

pipeline.speak("Hola, ¿cómo estás?")

# Internally maps "es" → "spanish" via QWEN3_LANGUAGE_ALIASES

```

The handler automatically converts the ISO code to the model's internal identifier before generation.

## Summary

- **Parakeet TDT** supports 25 European languages and uses the **lingua-py** library for external linguistic detection, applying a 20-character minimum threshold to avoid false positives.
- **Whisper handlers** extract language tokens directly from the model's generation output (`pred_ids[0, 1]`), providing zero-overhead detection with fallback to the last known language if unsupported codes are detected.
- **TTS handlers** rely on alias dictionaries (`QWEN3_LANGUAGE_ALIASES`, `WHISPER_LANGUAGE_TO_KOKORO_LANG`) to translate standard ISO-639-1 codes into model-specific language identifiers.
- All handlers expose explicit `SUPPORTED_LANGUAGES` lists, and the `language="auto"` parameter activates the appropriate detection mechanism for seamless multilingual operation.

## Frequently Asked Questions

### Which STT handler supports the most languages?

The **Parakeet TDT handler** supports 25 European languages, making it the most extensive option in the repository for European language coverage. The Whisper handlers support 12 languages focused on high-resource languages including Chinese, Japanese, Korean, and Hindi alongside European languages.

### How does the system handle short utterances in automatic language detection?

The Parakeet TDT handler specifically ignores utterances shorter than 20 characters when using **automatic language detection** to prevent inaccurate classifications on brief phrases like "yes" or "hello." This threshold is implemented in the `_detect_language_from_text()` method in [`parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/parakeet_tdt_handler.py). Whisper does not apply this restriction, as its language token is generated as part of the standard transcription process regardless of utterance length.

### Can I use automatic language detection with TTS handlers?

No, TTS handlers do not perform **automatic language detection** from audio or text content. They expect a language code to be supplied either by the STT handler upstream or via the `language` parameter. The TTS handler then uses its specific alias mapping (such as `QWEN3_LANGUAGE_ALIASES` in [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py)) to convert the ISO code into the format required by the underlying voice model.

### What happens when Whisper detects an unsupported language?

When the Whisper STT handler extracts a language token that is not present in its `SUPPORTED_LANGUAGES` list, it implements a fallback mechanism. The handler re-runs the transcription request using `self.last_language` (the language from the previous successful transcription) to ensure the conversation continues without interruption. This behavior is defined in the language validation logic within [`whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_handler.py).