How to Implement Bilingual Voice Conversations with Language Prompts

Pass a language_code parameter through the TTSInput message class to route language metadata from speech-to-text (STT) detection through the LLM to the text-to-speech (TTS) handler, enabling automatic language switching during real-time conversations.

The huggingface/speech-to-speech library enables real-time voice conversations using a modular pipeline architecture. By leveraging the language_code field in message dataclasses, developers can build bilingual voice conversations with language prompts that automatically detect, propagate, and synthesize speech in multiple languages within a single session.

How Language Flows Through the Pipeline

The pipeline transports language metadata across three stages using specific message types. Understanding this flow is essential for implementing robust multilingual support.

When audio enters the system, the STT handler (such as Whisper or Parakeet) detects the spoken language and creates a Transcription object containing the text and language_code. This code propagates through the pipeline until it reaches the TTS handler (such as Qwen‑3‑TTS or Pocket‑TTS), which uses it to select the appropriate synthesis model.

According to the source code in src/speech_to_speech/pipeline/messages.py, the TTSInput dataclass carries the language hint:


# src/speech_to_speech/pipeline/messages.py

class TTSInput(PipelineMessage):
    text: str
    language_code: Optional[str] = None   # Language hint for the TTS backend

If the language_code field is None, the TTS handler falls back to the last-used language or defaults to "en".

The Three Core Message Types

Transcription Messages

Transcription objects are produced by STT handlers and carry the detected language_code from the audio stream. In src/speech_to_speech/STT/whisper_stt_handler.py, the Whisper handler normalizes language codes and appends an -auto suffix when detection is automatic (e.g., "en-auto").

TTSInput Messages

TTSInput messages hold the text to be spoken and the optional language_code that instructs the TTS model which language to synthesize. The pipeline's transcription notifier forwards the language from the STT output to the next turn.

LLMResponse Messages

LLMResponse objects may embed a voice-system prompt that guides the model to prefer spoken output. The build_voice_system_prompt() function in src/speech_to_speech/LLM/voice_prompt.py includes rules such as "Speak naturally" and "If unsure whether a tool is needed, just speak", ensuring the model produces spoken replies even when language switching occurs.

Step-by-Step Implementation

1. Configure Explicit Language Codes

For controlled bilingual conversations, initialize TTSInput objects with explicit language codes:

from speech_to_speech.pipeline.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.pipeline.messages import TTSInput

# Initialize the pipeline

pipeline = SpeechToSpeechPipeline()

# English turn

english_input = TTSInput(
    text="Hello, how are you?", 
    language_code="en"
)
pipeline.send_user_input(english_input)

# Spanish turn  

spanish_input = TTSInput(
    text="¿Cómo estás?", 
    language_code="es"
)
pipeline.send_user_input(spanish_input)

The send_user_input() method in src/speech_to_speech/pipeline/s2s_pipeline.py wraps the TTSInput into a UserAudioChunk, processes it through the STT (which may verify or override the language), and forwards the language_code to the LLM. The TTS handler then reads this code and switches the synthesis model accordingly.

2. Enable Automatic Language Detection

To let the system detect language on the fly, set the session configuration to "auto":

pipeline.set_session_config({"audio": {"input": {"language": "auto"}}})

When using automatic detection, the STT handler emits codes like "es-auto". The helper function resolve_auto_language() in src/speech_to_speech/LLM/utils.py strips the -auto suffix and maps Whisper language codes to the LLM's internal identifiers:

from speech_to_speech.LLM.utils import resolve_auto_language

lang, llm_lang = resolve_auto_language("es-auto")

# Returns: lang == "es", llm_lang == "es"

This approach allows the pipeline to detect the spoken language via STT, propagate the detected code to the TTS handler, and fall back to the last-known language if detection fails.

Complete Working Example

Below is a full implementation demonstrating automatic bilingual conversation handling:


# demo/full_bilingual.py

from speech_to_speech.pipeline.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.pipeline.messages import TTSInput

# Initialize with auto-detection enabled

pipeline = SpeechToSpeechPipeline()
pipeline.set_session_config({"audio": {"input": {"language": "auto"}}})

# Turn 1 – English (detected automatically)

pipeline.send_user_input(
    TTSInput(text="Hey, what's the weather?", language_code=None)
)

# Turn 2 – Spanish (detected automatically)  

pipeline.send_user_input(
    TTSInput(text="¿Cuál es la temperatura en Madrid?", language_code=None)
)

Under the hood, this script:

  1. Detects "en" for the first turn and "es" for the second via the Whisper STT handler
  2. Maintains the voice-system prompt from voice_prompt.py to keep responses spoken-only
  3. Switches the TTS language on-the-fly (e.g., Qwen‑3‑TTS with language="es")

Key Source Files and Functions

File Purpose
src/speech_to_speech/LLM/voice_prompt.py Defines build_voice_system_prompt() for spoken-only output guidance
src/speech_to_speech/pipeline/messages.py Contains TTSInput dataclass with language_code field
src/speech_to_speech/LLM/utils.py Implements resolve_auto_language() for normalizing auto-detected codes
src/speech_to_speech/STT/whisper_stt_handler.py Performs language detection and manages the -auto suffix logic
src/speech_to_speech/TTS/qwen3_tts_handler.py Reads language_code to select multilingual synthesis models
src/speech_to_speech/pipeline/s2s_pipeline.py Orchestrates the end-to-end flow via send_user_input()

Summary

  • Use TTSInput.language_code to explicitly set or propagate language metadata through the pipeline
  • Enable "auto" detection via session configuration to let STT handlers automatically identify spoken languages
  • Handle -auto suffixes with resolve_auto_language() in src/speech_to_speech/LLM/utils.py when processing automatic detections
  • Leverage voice-system prompts from voice_prompt.py to ensure the LLM produces spoken output suitable for voice synthesis
  • Configure fallback behavior by understanding that TTS handlers default to "en" or the last-used language when language_code is None

Frequently Asked Questions

How does the pipeline handle language switching mid-conversation?

The pipeline treats each turn independently while maintaining a language state. When send_user_input() receives a new TTSInput, it processes the contained language_code (or detects it via STT if set to "auto") and passes this metadata to the TTS handler. The TTS backend in qwen3_tts_handler.py (or equivalent) then switches synthesis models or voice parameters accordingly, allowing seamless transitions between languages within the same session.

What happens if language detection fails?

According to the logic in src/speech_to_speech/STT/whisper_stt_handler.py, the system falls back to the last-known language code when detection is uncertain. If no previous language exists, the TTS handler defaults to "en". This ensures the conversation continues without interruption, though developers can override this behavior by explicitly setting language_code in TTSInput messages.

Can I force a specific language instead of using auto-detection?

Yes. Instead of setting the session configuration to "auto", omit the configuration or set explicit language_code values in your TTSInput objects (e.g., "en", "es", "fr"). When provided, these codes bypass automatic detection and instruct the TTS handler to use the specified language model directly, as implemented in the message handling logic of s2s_pipeline.py.

Which TTS handlers support multilingual output?

The reference implementation includes support in src/speech_to_speech/TTS/qwen3_tts_handler.py, which reads the language_code field and passes it to the underlying Qwen‑3‑TTS model. Other handlers like pocket_tts_handler.py implement similar logic, reading the language_code from TTSInput to select appropriate synthesis parameters. Check your specific TTS handler implementation to confirm multilingual support.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →