How Does --enable_lang_prompt Work with Language Auto-Detection in Speech-to-Speech?
The --enable_lang_prompt flag instructs the LLM handler to prepend a language-specific instruction to the prompt whenever the STT handler detects a language, ensuring the model responds in the same language as the input speech.
The huggingface/speech-to-speech repository implements a modular pipeline where Speech-to-Text (STT) handlers detect spoken languages and attach metadata to transcription events. When enabled, the --enable_lang_prompt argument leverages this detected language code to force the Language Model (LLM) to reply in the detected language, creating a seamless multilingual speech-to-speech experience.
Understanding Language Auto-Detection in STT Handlers
Language detection originates in the STT layer, where audio processors identify the spoken language before text generation begins.
How STT Handlers Detect Language
Individual STT handlers—such as the Parakeet, Whisper, or Qwen3 TTS handlers—implement language detection logic specific to their underlying models. For example, the Parakeet handler utilizes a _detect_language_from_text method to analyze incoming audio and determine the primary language. This detection occurs during the transcription phase, where the handler processes raw audio input and extracts not only the spoken text but also linguistic markers indicating the language used.
The Transcription Event and language_code
Once detection completes, the handler populates a language_code field on the Transcription event object. This standardized field carries the ISO language identifier (e.g., "en" for English, "fr" for French) through the pipeline. The Transcription object then propagates to the LLM handler, making the detected language available for downstream processing without requiring additional detection steps.
How --enable_lang_prompt Functions in the LLM Pipeline
The flag operates at the intersection of the STT output and LLM input, modifying the prompt construction logic based on the detected language.
Argument Definition in LanguageModelBaseArguments
The --enable_lang_prompt argument is defined in src/speech_to_speech/arguments_classes/language_model_base_arguments.py as a boolean field defaulting to False:
# src/speech_to_speech/arguments_classes/language_model_base_arguments.py#L33-L38
enable_lang_prompt: bool = field(
default=False,
metadata={
"help": "When True, append a user message instructing the model to reply in the detected/selected "
"language (e.g. 'Please reply to my message in French.'). Default is False."
},
)
When set to True, this flag activates the prompt injection mechanism in the LLM handler.
Prompt Injection Logic in LanguageModelHandler
Inside src/speech_to_speech/LLM/language_model.py, the LanguageModelHandler checks for both the presence of a detected language (lang_name) and the enabled flag before modifying the message list:
# src/speech_to_speech/LLM/language_model.py#L540-L543
if lang_name and self.enable_lang_prompt:
# prepend a language-specific instruction to the prompt
messages.insert(
0,
{
"role": "user",
"content": f"Please reply to my message in {lang_name}."
},
)
This insertion occurs immediately before the LLM generation call, ensuring the model receives an explicit instruction to respond in the detected language while maintaining the original user query context.
End-to-End Workflow
The interaction between language detection and prompt enabling follows a precise sequence:
- Audio Processing: The STT handler (e.g., Parakeet) receives audio input and executes
_detect_language_from_textto identify the language. - Metadata Attachment: The handler creates a
Transcriptionobject containing both the text and thelanguage_codefield. - Pipeline Handoff: The transcription event reaches the
LanguageModelHandler, which extracts the language name from thelanguage_code. - Conditional Prompting: If
enable_lang_promptisTrueand a language is detected, the handler inserts the instructional message: "Please reply to my message in {lang_name}." - Response Generation: The LLM processes the augmented prompt and generates a response in the specified language, which then flows to the TTS component for speech synthesis.
Practical Implementation Examples
Enable the feature via command line when starting the server:
python -m speech_to_speech.server \
--enable_lang_prompt \
--stt_handler parakeet \
--llm_handler qwen3
Or configure it programmatically when building the pipeline:
from speech_to_speech.arguments_classes.language_model_base_arguments import LanguageModelBaseArguments
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
# Configure with language prompting enabled
args = LanguageModelBaseArguments(enable_lang_prompt=True)
# Initialize pipeline with language-aware settings
pipeline = SpeechToSpeechPipeline(
language_model_handler_kwargs={"enable_lang_prompt": True}
)
pipeline.run()
Summary
- Language detection occurs in STT handlers (e.g.,
parakeet_tdt_handler.py) which populate thelanguage_codefield onTranscriptionevents. --enable_lang_promptis defined inlanguage_model_base_arguments.pyand activates automatic prompt injection when set toTrue.- The LLM handler checks for detected languages and the enabled flag, then prepends "Please reply to my message in {lang_name}." to the conversation.
- This architecture separates detection from instruction, allowing any STT handler supporting language detection to work with any LLM handler supporting prompt modification.
Frequently Asked Questions
Does --enable_lang_prompt detect the language itself?
No, the flag does not perform detection. According to the source code in language_model.py, it only consumes the language_code already detected by the STT handler. The detection logic resides in handlers like parakeet_tdt_handler.py which implement methods such as _detect_language_from_text.
Which STT handlers support language auto-detection?
The repository includes language detection capabilities in several STT handlers, including the Parakeet, Whisper, and Qwen3 TTS handlers. Each handler implements detection logic specific to its underlying model and attaches the resulting code to the Transcription event's language_code field.
What happens if language detection fails?
If the STT handler cannot determine a language, the language_code field remains unset or null. In language_model.py, the prompt injection logic explicitly checks if lang_name and self.enable_lang_prompt, meaning no language instruction is added when detection fails, and the LLM responds according to its default behavior or training.
Can I use --enable_lang_prompt with manual language selection?
Yes, while the flag is designed to work with auto-detection, the same language_code mechanism can carry manually specified languages. As long as the transcription event contains a valid language_code, the LLM handler will inject the corresponding language prompt regardless of whether the source was automatic detection or manual configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →