How to Set Up Multi-Language Support with Automatic Language Detection in Speech-to-Speech

Install lingua-py, use the Parakeet TDT handler, and set language="auto" (or omit the language parameter) to enable automatic language detection across 25 supported European languages.

The Hugging Face speech-to-speech repository provides real-time speech-to-speech translation pipelines with built-in multi-language support. When you need the system to automatically identify the spoken language without manual configuration, the Parakeet TDT STT handler integrates the Lingua-py library for automatic language detection (ALD) across 25 European languages.

Prerequisites: Installing the Language Detection Dependency

Automatic language detection requires the optional lingua-py package. The handler implements a lazy-loading mechanism via the LINGUA_AVAILABLE flag, but you must install the dependency before starting the pipeline.

pip install lingua-py

Without this package, the handler will skip detection and fall back to the default language or the last detected language.

How Automatic Language Detection Works

The automatic language detection system operates entirely within the ParakeetTDTSTTHandler class in src/speech_to_speech/STT/parakeet_tdt_handler.py. The handler detects language in real-time after each transcription using the following logic:

  • Detection trigger: When language=None or language="auto" is passed to the handler, the system calls _detect_language_from_text after generating the transcription text.
  • Language library: The handler uses Lingua-py, a high-accuracy natural language detection library, to infer ISO-639-1 language codes from the transcribed text.
  • Supported languages: The system supports 25 European languages defined in the SUPPORTED_LANGUAGES list, ensuring the detector returns only valid, supported codes.
  • Pre-loading: To avoid latency on first use, the handler pre-loads language models at startup using LanguageDetectorBuilder.from_languages(*_lingua_languages).with_preloaded_language_models().

Configuration Methods

You can enable multi-language support with automatic detection via the Python API or the command-line interface.

Python API Configuration

When constructing the pipeline in src/speech_to_speech/s2s_pipeline.py, pass language=None or omit the parameter to trigger automatic detection:

from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline

# language=None enables automatic detection

pipeline = SpeechToSpeechPipeline(
    stt_handler="parakeet",
    language=None,  # or "auto"

    # other arguments...

)

The SpeechToSpeechPipeline internally instantiates the ParakeetTDTSTTHandler and passes the language argument defined in src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py.

Command-Line Interface

Use the listen_and_play_realtime.py script with the --language auto flag:

python -m speech_to_speech.scripts.listen_and_play_realtime \
    --stt-handler parakeet \
    --language auto

If you want to lock to a specific language and disable automatic detection, provide an ISO-639-1 code (e.g., --language de or --language fr).

Implementation Details: The Detection Pipeline

The detection logic follows a specific priority order to determine the final language code for each utterance.

Language Detection Method

The _detect_language_from_text method (lines 79-104 in parakeet_tdt_handler.py) implements the core detection logic:

def _detect_language_from_text(self, text: str) -> Optional[str]:
    if not LINGUA_AVAILABLE:
        return None
    if not text or len(text.strip()) < 20:
        return None  # Skip very short utterances

    
    detected = _lingua_detector.detect_language_of(text)
    if detected is None:
        return None
    
    code = detected.iso_code_639_1.name.lower()
    # Map Lingua code back to STS code table

    return {v: k for k, v in _LINGUA_CODE_MAP.items()}.get(code, code)

The handler skips detection for utterances shorter than 20 characters to ensure sufficient text for accurate classification.

Language Resolution Priority

After transcription in _process_mlx or _process_nano_parakeet, the handler resolves the language code in this order:

  1. User-provided language: If self.start_language is set and not equal to "auto", the handler uses this value and skips detection.
  2. Auto-detected language: If the user requested auto-detection, the handler calls _detect_language_from_text.
  3. Last known language: If detection fails (returns None), the system falls back to self.last_language, persisting the last successful detection across utterances.
if self.start_language and self.start_language != "auto":
    language_code = self.start_language
else:
    detected_lang = self._detect_language_from_text(pred_text)
    language_code = detected_lang or self.last_language

Handling Language Fallbacks and Edge Cases

The handler maintains state across utterances through the self.last_language attribute. This ensures that if the detector fails to identify a short or ambiguous phrase, the system continues using the previously detected language rather than defaulting to a potentially incorrect code.

To force a specific language and disable automatic detection, specify the language code explicitly:


# Force German language, skip auto-detection

pipeline = SpeechToSpeechPipeline(
    stt_handler="parakeet",
    language="de"
)

Summary

  • Install lingua-py to enable automatic language detection capabilities.
  • Use the Parakeet TDT handler (--stt-handler parakeet) as it is the only handler shipping with built-in automatic language detection.
  • Set language="auto" or language=None to trigger detection, or omit the parameter entirely.
  • Detection requires 20+ characters of transcribed text; shorter utterances fall back to the last detected language.
  • Configure via Python using SpeechToSpeechPipeline or via CLI using listen_and_play_realtime.py.
  • Supported languages include 25 European languages mapped to ISO-639-1 codes in the handler's SUPPORTED_LANGUAGES list.

Frequently Asked Questions

What is the minimum text length required for automatic language detection?

The _detect_language_from_text method in parakeet_tdt_handler.py skips detection for any utterance shorter than 20 characters. This threshold ensures Lingua-py has sufficient content to accurately classify the language. If the text is too short, the handler returns None and falls back to self.last_language.

Can I use automatic language detection with other STT handlers?

No. According to the source code, only the Parakeet TDT handler (ParakeetTDTSTTHandler) implements automatic language detection. Other handlers in the repository do not include the _detect_language_from_text method or the Lingua-py integration. You must specify stt_handler="parakeet" to use this feature.

How does the system handle language detection failures?

If Lingua-py returns None or the text is too short, the handler falls back to self.last_language, which stores the most recently successful detection. This state persists across utterances, ensuring continuity in multi-turn conversations even when individual phrases are too brief for reliable detection.

What languages are supported for automatic detection?

The handler supports 25 European languages defined in the SUPPORTED_LANGUAGES list within parakeet_tdt_handler.py. The detection system uses ISO-639-1 codes (e.g., en, fr, de, es) and maps them through the Lingua-py library's iso_code_639_1 enumeration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →