How the `--language auto` Flag Works in Hugging Face Speech-to-Speech

The --language auto flag enables per-utterance language detection in the Speech-to-Text (STT) handler, which extracts the language token from model outputs, stores it in a Transcription message, and forwards it through the pipeline to the Large Language Model (LLM), where resolve_auto_language normalizes the code for optional prompt injection.

The huggingface/speech-to-speech repository implements a modular pipeline for real-time voice conversations. When you enable automatic language detection via the --language auto command-line argument, the system identifies the spoken language for each audio segment and propagates that metadata through to the LLM, ensuring responses match the user's language.

The End-to-End Language Detection Flow

Setting --language auto activates a multi-stage process that bridges the STT and LLM components. The flow follows this path:

  1. CLI Argument Parsing: The shared --language argument is defined in src/speech_to_speech/arguments_classes/whisper_stt_arguments.py and passed to the selected STT handler.
  2. STT Token Extraction: The handler decodes the second token from the model's generated sequence to determine the ISO language code.
  3. Message Propagation: The code is attached to a Transcription object and yielded into the pipeline.
  4. LLM Resolution: The LLM handler calls resolve_auto_language from src/speech_to_speech/LLM/utils.py to strip suffixes and map codes to human-readable names.
  5. Prompt Injection: If --enable_lang_prompt is set, the handler prepends a system instruction forcing the LLM to reply in the detected language.

STT Handler Language Extraction

All STT handlers in the repository follow the same pattern for language detection, with the Whisper implementation serving as the reference.

Token Decoding in Whisper STT

In src/speech_to_speech/STT/whisper_stt_handler.py, the model generates a sequence of token IDs where the second token encodes the language. After generation, the handler extracts the code:

def process(self, vad_audio):
    pred_ids = self.model.generate(self.prepare_model_inputs(vad_audio.audio), **self.gen_kwargs)
    # language token is the second token

    language_code = self.processor.tokenizer.decode(pred_ids[0, 1])[2:-2]

    # attach language to the transcription message

    yield Transcription(
        text=self.processor.batch_decode(pred_ids)[0],
        language_code=language_code,
        turn_id=vad_audio.turn_id,
        turn_revision=vad_audio.turn_revision,
        speech_stopped_at_s=vad_audio.created_at_s,
    )

The slicing [2:-2] removes special characters from the decoded token to isolate the raw ISO code (e.g., en, fr).

Fallback for Unsupported Languages

If the detected language is not in the supported list, the handler falls back to the previously detected language stored in self.last_language. This prevents pipeline interruptions when the model outputs rare or invalid language codes.

Alternative STT Implementations

Other handlers implement identical extraction logic:

  • LightningWhisperSTTHandler: Extracts language tokens from the Lightning Whisper model output.
  • ParakeetTDTSTTHandler: Handles language detection for the Parakeet TDT model in src/speech_to_speech/STT/parakeet_tdt_handler.py.
  • MLXAudioWhisperSTTHandler: Optimized variant for Apple Silicon using MLX.

Pipeline Message Forwarding

The Transcription dataclass acts as the contract between pipeline stages. The language_code field populated by the STT handler travels unchanged through the pipeline queue until it reaches the LLM handler in src/speech_to_speech/LLM/language_model.py.

LLM Language Resolution and Prompting

Once the LLM handler receives a request, it processes the language metadata before generating text.

Normalizing Language Codes

The resolve_auto_language function in src/speech_to_speech/LLM/utils.py handles normalization:

language_code, lang_name = resolve_auto_language(request.language_code)

When --language auto is used, Whisper appends an -auto suffix to the code. This function strips that suffix and maps the ISO code to a human-readable language name (e.g., en becomes English).

Optional Language Enforcement

If you start the pipeline with --enable_lang_prompt, the handler injects a system message forcing the model to respond in the detected language. According to the source code at lines 39-42 of src/speech_to_speech/LLM/language_model.py:

if lang_name and self.enable_lang_prompt:
    active_chat.add_item(
        make_user_message(f"Please reply to my message in {lang_name}.")
    )

This ensures the LLM respects the user's spoken language even if the model was trained on multilingual data with English-heavy system prompts.

Configuration and Usage Examples

To enable automatic language detection with language prompt enforcement, use the following configuration via the main entry point src/speech_to_speech/s2s_pipeline.py:


# Run the pipeline with automatic language detection

python s2s_pipeline.py \
    --stt whisper \
    --language auto \
    --llm_backend transformers \
    --model_name Qwen/Qwen3-4B-Instruct-2507 \
    --enable_lang_prompt

This command activates per-utterance detection in the Whisper handler and ensures the LLM receives normalized language codes with optional prompting enabled.

Summary

  • The --language auto flag triggers token-based language detection in STT handlers like Whisper, Parakeet, and Lightning Whisper.
  • Detected language codes are extracted from the second generated token via decode(pred_ids[0, 1])[2:-2] and stored in Transcription.language_code.
  • The pipeline forwards the code to the LLM handler, where resolve_auto_language in src/speech_to_speech/LLM/utils.py normalizes ISO codes and strips the -auto suffix.
  • Enabling --enable_lang_prompt injects a system instruction at lines 39-42 of src/speech_to_speech/LLM/language_model.py, forcing the LLM to reply in the detected language.

Frequently Asked Questions

How does Whisper detect the language when using --language auto?

Whisper models generate a sequence of tokens where the second token encodes the language identifier. The handler decodes this token using the processor's tokenizer, extracts the ISO code via string slicing [2:-2], and attaches it to the Transcription message as language_code.

What happens if the detected language is not supported?

If the extracted language code is not in the supported languages list, the handler falls back to self.last_language, which stores the most recently valid detection. This prevents pipeline failures from invalid or rare language codes.

Where is the language code processed before reaching the LLM?

The LLM handler in src/speech_to_speech/LLM/language_model.py receives the code via request.language_code and passes it to resolve_auto_language in src/speech_to_speech/LLM/utils.py. This function strips the -auto suffix and maps the ISO code to a human-readable name for optional prompt injection.

Can I force the LLM to respond in the detected language automatically?

Yes. Pass the --enable_lang_prompt flag when starting the pipeline. This injects a system message instructing the LLM to reply in the specific language identified by the STT handler, ensuring coherent multilingual conversations without manual language switching.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →