How to Configure Multi-Language Support with Automatic Language Detection in Speech-to-Speech
Enable automatic language detection by installing lingua-py and using the Parakeet TDT STT handler with language="auto" or language=None, which activates the _detect_language_from_text method to identify 25 supported European languages from transcribed text.
The huggingface/speech-to-speech repository provides real-time speech-to-speech translation capabilities with built-in support for configuring multi-language support with automatic language detection. While multiple STT handlers exist, only the Parakeet TDT implementation offers native automatic language detection (ALD) through the Lingua-py library. This guide explains how to enable and configure this feature using both the Python API and command-line interface.
How Automatic Language Detection Works
The Parakeet TDT handler (ParakeetTDTSTTHandler) is currently the only STT handler in the repository that ships with built-in automatic language detection capabilities. Located in src/speech_to_speech/STT/parakeet_tdt_handler.py, this handler integrates the Lingua-py library to infer ISO-639-1 language codes from transcribed text when users do not specify a target language.
The Detection Mechanism
The handler initializes a lazy-loaded language detector via the LINGUA_AVAILABLE flag. Lines 72-89 pre-load language models at startup to eliminate latency on first use:
# src/speech_to_speech/STT/parakeet_tdt_handler.py
if LINGUA_AVAILABLE:
_lingua_iso_to_code = { … }
_lingua_languages = [ … ]
def _build_lingua_detector():
return LanguageDetectorBuilder.from_languages(*_lingua_languages)\
.with_preloaded_language_models()\
.build()
_lingua_detector = _build_lingua_detector()
Detection Logic Implementation
The _detect_language_from_text method (lines 79-104) handles the actual detection logic. It skips very short utterances (fewer than 20 characters), runs the detector, maps the Lingua code back to the STS code table, and returns None if detection fails:
def _detect_language_from_text(self, text: str) -> Optional[str]:
if not LINGUA_AVAILABLE: …
if not text or len(text.strip()) < 20: …
detected = _lingua_detector.detect_language_of(text)
if detected is None: return None
code = detected.iso_code_639_1.name.lower()
return {v: k for k, v in _LINGUA_CODE_MAP.items()}.get(code, code)
Installation Requirements
Installing the Lingua-py Dependency
Automatic language detection requires the optional Lingua-py package. Without this dependency, the handler cannot perform language inference and will skip detection entirely.
pip install lingua-py
Configuration Methods
You can activate multi-language support through either the Python API or the command-line interface. Both methods require setting the language parameter to "auto" or None to trigger the detection pipeline.
Python API Configuration
In src/speech_to_speech/s2s_pipeline.py, the SpeechToSpeechPipeline class accepts a language argument that it passes to the handler. Set language=None to enable automatic detection:
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
# language=None defaults to auto-detection
pipeline = SpeechToSpeechPipeline(
stt_handler="parakeet",
language=None, # activates automatic language detection
# other arguments …
)
The pipeline internally creates a ParakeetTDTSTTHandler instance and passes the language argument defined in src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py.
Command Line Configuration
Use the scripts/listen_and_play_realtime.py entry point with the --language auto flag:
python -m speech_to_speech.scripts.listen_and_play_realtime \
--stt-handler parakeet \
--language auto # auto-detect per utterance
Language Selection Priority
After transcription in _process_mlx and _process_nano_parakeet, the handler determines the final language code using this priority order:
- User-provided language: If
self.start_languageis set and not equal to"auto", the handler uses this value exclusively. - Detected language: If auto-detection is active, the handler calls
_detect_language_from_texton the transcribed text. - Fallback language: If detection fails or returns
None, the system usesself.last_language(the most recently successful detection).
if self.start_language and self.start_language != "auto":
language_code = self.start_language
else:
detected_lang = self._detect_language_from_text(pred_text)
language_code = detected_lang or self.last_language
Supported Languages and Constraints
The Parakeet TDT handler supports 25 European languages defined in the SUPPORTED_LANGUAGES constant. The Lingua-py detector is configured specifically for this subset, ensuring returned ISO-639-1 codes (e.g., en, fr, de, es) are always compatible with the handler's capability set. The system prints detected language codes to the console output (lines 53-55 of the handler file), allowing you to verify detection accuracy in real-time.
Summary
- Install
lingua-pyto enable the automatic language detection feature - Use the
parakeetSTT handler located insrc/speech_to_speech/STT/parakeet_tdt_handler.py - Set
language="auto"orlanguage=Noneto activate detection per utterance - Detection requires a minimum of 20 characters of transcribed text to trigger
- The system maintains state in
self.last_languageto handle ambiguous or short utterances
Frequently Asked Questions
Which STT handlers support automatic language detection?
Only the Parakeet TDT handler (parakeet) supports automatic language detection. Other handlers in the repository require explicit language configuration. The detection logic resides in the _detect_language_from_text method of src/speech_to_speech/STT/parakeet_tdt_handler.py.
What happens if the detected language is not supported?
If Lingua-py cannot determine the language with confidence, or if the transcribed text contains fewer than 20 characters, the method returns None. The handler then falls back to self.last_language, which stores the most recently successful detection to maintain continuity across short responses.
Can I lock the pipeline to a specific language?
Yes. Pass an explicit ISO-639-1 code (e.g., language="de" or --language de) to bypass automatic detection entirely. When a specific language is provided, the handler skips the _detect_language_from_text call and forces that language for all subsequent transcriptions.
Where is the language detection configured in the source code?
The core detection implementation is in src/speech_to_speech/STT/parakeet_tdt_handler.py (lines 72-104). CLI arguments are defined in src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py, and the pipeline wiring occurs in src/speech_to_speech/s2s_pipeline.py. The README.md (line 468) also documents this feature under "Automatic language detection."
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →