How Does --enable_lang_prompt Work with Language Auto-Detection in Speech-to-Speech?

The --enable_lang_prompt flag instructs the LLM handler to prepend a language-specific instruction to the prompt whenever the STT handler detects a language, ensuring the model responds in the same language as the input speech.

The huggingface/speech-to-speech repository implements a modular pipeline where Speech-to-Text (STT) handlers detect spoken languages and attach metadata to transcription events. When enabled, the --enable_lang_prompt argument leverages this detected language code to force the Language Model (LLM) to reply in the detected language, creating a seamless multilingual speech-to-speech experience.

Understanding Language Auto-Detection in STT Handlers

Language detection originates in the STT layer, where audio processors identify the spoken language before text generation begins.

How STT Handlers Detect Language

Individual STT handlers—such as the Parakeet, Whisper, or Qwen3 TTS handlers—implement language detection logic specific to their underlying models. For example, the Parakeet handler utilizes a _detect_language_from_text method to analyze incoming audio and determine the primary language. This detection occurs during the transcription phase, where the handler processes raw audio input and extracts not only the spoken text but also linguistic markers indicating the language used.

The Transcription Event and language_code

Once detection completes, the handler populates a language_code field on the Transcription event object. This standardized field carries the ISO language identifier (e.g., "en" for English, "fr" for French) through the pipeline. The Transcription object then propagates to the LLM handler, making the detected language available for downstream processing without requiring additional detection steps.

How --enable_lang_prompt Functions in the LLM Pipeline

The flag operates at the intersection of the STT output and LLM input, modifying the prompt construction logic based on the detected language.

Argument Definition in LanguageModelBaseArguments

The --enable_lang_prompt argument is defined in src/speech_to_speech/arguments_classes/language_model_base_arguments.py as a boolean field defaulting to False:


# src/speech_to_speech/arguments_classes/language_model_base_arguments.py#L33-L38

enable_lang_prompt: bool = field(
    default=False,
    metadata={
        "help": "When True, append a user message instructing the model to reply in the detected/selected "
                "language (e.g. 'Please reply to my message in French.'). Default is False."
    },
)

When set to True, this flag activates the prompt injection mechanism in the LLM handler.

Prompt Injection Logic in LanguageModelHandler

Inside src/speech_to_speech/LLM/language_model.py, the LanguageModelHandler checks for both the presence of a detected language (lang_name) and the enabled flag before modifying the message list:


# src/speech_to_speech/LLM/language_model.py#L540-L543

if lang_name and self.enable_lang_prompt:
    # prepend a language-specific instruction to the prompt

    messages.insert(
        0,
        {
            "role": "user",
            "content": f"Please reply to my message in {lang_name}."
        },
    )

This insertion occurs immediately before the LLM generation call, ensuring the model receives an explicit instruction to respond in the detected language while maintaining the original user query context.

End-to-End Workflow

The interaction between language detection and prompt enabling follows a precise sequence:

  1. Audio Processing: The STT handler (e.g., Parakeet) receives audio input and executes _detect_language_from_text to identify the language.
  2. Metadata Attachment: The handler creates a Transcription object containing both the text and the language_code field.
  3. Pipeline Handoff: The transcription event reaches the LanguageModelHandler, which extracts the language name from the language_code.
  4. Conditional Prompting: If enable_lang_prompt is True and a language is detected, the handler inserts the instructional message: "Please reply to my message in {lang_name}."
  5. Response Generation: The LLM processes the augmented prompt and generates a response in the specified language, which then flows to the TTS component for speech synthesis.

Practical Implementation Examples

Enable the feature via command line when starting the server:

python -m speech_to_speech.server \
    --enable_lang_prompt \
    --stt_handler parakeet \
    --llm_handler qwen3

Or configure it programmatically when building the pipeline:

from speech_to_speech.arguments_classes.language_model_base_arguments import LanguageModelBaseArguments
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline

# Configure with language prompting enabled

args = LanguageModelBaseArguments(enable_lang_prompt=True)

# Initialize pipeline with language-aware settings

pipeline = SpeechToSpeechPipeline(
    language_model_handler_kwargs={"enable_lang_prompt": True}
)
pipeline.run()

Summary

  • Language detection occurs in STT handlers (e.g., parakeet_tdt_handler.py) which populate the language_code field on Transcription events.
  • --enable_lang_prompt is defined in language_model_base_arguments.py and activates automatic prompt injection when set to True.
  • The LLM handler checks for detected languages and the enabled flag, then prepends "Please reply to my message in {lang_name}." to the conversation.
  • This architecture separates detection from instruction, allowing any STT handler supporting language detection to work with any LLM handler supporting prompt modification.

Frequently Asked Questions

Does --enable_lang_prompt detect the language itself?

No, the flag does not perform detection. According to the source code in language_model.py, it only consumes the language_code already detected by the STT handler. The detection logic resides in handlers like parakeet_tdt_handler.py which implement methods such as _detect_language_from_text.

Which STT handlers support language auto-detection?

The repository includes language detection capabilities in several STT handlers, including the Parakeet, Whisper, and Qwen3 TTS handlers. Each handler implements detection logic specific to its underlying model and attaches the resulting code to the Transcription event's language_code field.

What happens if language detection fails?

If the STT handler cannot determine a language, the language_code field remains unset or null. In language_model.py, the prompt injection logic explicitly checks if lang_name and self.enable_lang_prompt, meaning no language instruction is added when detection fails, and the LLM responds according to its default behavior or training.

Can I use --enable_lang_prompt with manual language selection?

Yes, while the flag is designed to work with auto-detection, the same language_code mechanism can carry manually specified languages. As long as the transcription event contains a valid language_code, the LLM handler will inject the corresponding language prompt regardless of whether the source was automatic detection or manual configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →