How to Convert Speech from One Language to Another with Hugging Face Speech-to-Speech

The speech-to-speech library converts spoken input to spoken output in a different language through a three-stage pipeline: Speech-to-Text (STT) detects the source language, a Language Model translates the text, and Text-to-Speech (TTS) synthesizes the target language audio—all coordinated by S2SPipeline in s2s_pipeline.py.

This guide walks through the complete process of building a multilingual speech translator using the huggingface/speech-to-speech repository. The library provides a modular, full-duplex pipeline that chains together STT, translation, and TTS components with automatic language code propagation.

Understanding the Multilingual Pipeline Architecture

The speech-to-speech translation flow consists of three core stages that communicate via typed messages carrying language_code fields:

Stage Handler Location Function
Speech-to-Text src/speech_to_speech/STT/whisper_stt_handler.py Captures audio, runs Whisper, emits Transcription with detected source language_code
Language Model / Translation src/speech_to_speech/LLM/language_model.py via LanguageModelHandler Receives transcription, generates translation, returns LLMResponseChunk with target language_code
Text-to-Speech src/speech_to_speech/TTS/qwen3_tts_handler.py Synthesizes audio using the target language_code from the LM stage

The S2SPipeline class in src/speech_to_speech/s2s_pipeline.py orchestrates these stages by creating async queues and wiring handlers together. It automatically propagates language metadata through each message type defined in src/speech_to_speech/pipeline/messages.py.

Language Handling Mechanics

Source Language Detection

When WhisperSTTHandler receives --language auto, Whisper detects the spoken language and stores it in handler.last_language. This value attaches to the Transcription message as language_code (e.g., "en", "fr", "de").

Target Language Selection

The target language flows through the pipeline via runtime_config.language_code. When LanguageModelHandler builds a GenerateResponseRequest, the language_code field signals which language the LLM should generate. Handlers normalize aliases (e.g., "English" → "en"; "nb" → "no") and fall back to the last valid language if unsupported.

Running Speech-to-Speech Translation

Command-Line Approach with listen_and_play.py

The repository includes scripts/listen_and_play.py as a complete working example:


# scripts/listen_and_play.py

import argparse
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelArguments

parser = argparse.ArgumentParser(description="Speech-to-speech translation")
parser.add_argument("--source-language", default="auto", help="Source language (auto for detection)")
parser.add_argument("--target-language", required=True, help="Target language code, e.g. 'fr' or 'de'")
args = parser.parse_args()

runtime_config = {
    "language_code": args.target_language,   # Target language propagated to LM and TTS

    "stt": WhisperSTTArguments(language=args.source_language),
    "tts": Qwen3TTSArguments(),
    "lm": LanguageModelArguments(),
}

pipeline = S2SPipeline(runtime_config)
pipeline.run()  # Starts interactive microphone-to-speaker loop

The pipeline executes as follows:

  1. WhisperSTTHandler detects source language and emits Transcription(language_code="en")
  2. LanguageModelHandler prompts the LLM to translate, returns LLMResponseChunk(language_code="fr")
  3. Qwen3TTSHandler synthesizes French speech using the target language_code

Programmatic API Without CLI

For integration into larger applications, instantiate S2SPipeline directly:

from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelArguments

# Configure for Spanish output

config = {
    "language_code": "es",
    "stt": WhisperSTTArguments(language="auto"),
    "tts": Qwen3TTSArguments(),
    "lm": LanguageModelArguments(),
}
pipeline = S2SPipeline(config)

# Feed audio from file, microphone, or stream

with open("input_english.wav", "rb") as f:
    audio_bytes = f.read()
pipeline.feed_audio(audio_bytes)

# Retrieve translated audio

translated_audio = pipeline.get_output_audio()

Fixed Source Language (Skip Auto-Detection)

For known input languages, disable detection to reduce latency:

config["stt"] = WhisperSTTArguments(language="de")  # German input

config["language_code"] = "en"                      # English output

pipeline = S2SPipeline(config)

Key Source Files for Language Conversion

File Path Purpose
Pipeline orchestrator src/speech_to_speech/s2s_pipeline.py S2SPipeline class creates queues and runs the async event loop
Message types src/speech_to_speech/pipeline/messages.py Transcription, TTSInput, AudioOutput with language_code fields
Event types src/speech_to_speech/pipeline/events.py TranscriptionCompletedEvent and other pipeline events
Whisper STT src/speech_to_speech/STT/whisper_stt_handler.py Speech recognition with language detection
Qwen-3 TTS src/speech_to_speech/TTS/qwen3_tts_handler.py Speech synthesis respecting target language_code
LLM translation src/speech_to_speech/LLM/language_model.py Translation via LanguageModelHandler
Demo script scripts/listen_and_play.py Complete end-to-end example

Swappable Backend Components

The modular design allows mixing STT, LM, and TTS backends. Replace argument classes to use alternatives:

  • STT: WhisperSTTArguments, or custom handlers in src/speech_to_speech/STT/
  • LM: Any OpenAI-compatible API via LanguageModelArguments, or local models
  • TTS: Qwen3TTSArguments, or other handlers in src/speech_to_speech/TTS/

All handlers respect the language_code field in their input messages, ensuring consistent language propagation regardless of backend choice.

Summary

  • The speech-to-speech library implements translation through a three-stage pipeline coordinated by S2SPipeline in src/speech_to_speech/s2s_pipeline.py
  • Source language is detected by Whisper or specified explicitly; stored in Transcription.language_code
  • Target language is set via runtime_config.language_code and propagated through LM and TTS stages
  • Message types in src/speech_to_speech/pipeline/messages.py carry language_code fields ensuring metadata flows through the pipeline
  • scripts/listen_and_play.py provides a complete CLI example for interactive translation
  • Backend components are swappable—mix Whisper STT, any OpenAI-compatible LLM, and Qwen-3 TTS or alternatives

Frequently Asked Questions

What language codes does the speech-to-speech library support?

The library uses standard ISO 639-1 codes ("en", "fr", "de", "es", "zh", etc.). Handlers in whisper_stt_handler.py and qwen3_tts_handler.py normalize aliases automatically, mapping "English" to "en" and "nb" (Norwegian Bokmål) to "no". Unsupported codes trigger fallback to the last valid language.

Can I use a different translation model instead of the default LLM?

Yes. The LanguageModelArguments class accepts any OpenAI-compatible API endpoint. Configure the base_url and model parameters to switch to local models (via vLLM, llama.cpp) or alternative APIs. The LanguageModelHandler builds translation prompts and returns LLMResponseChunk with the target language_code regardless of backend.

How does the pipeline handle language detection failures?

If Whisper cannot determine the source language, whisper_stt_handler.py falls back to handler.last_language (previously detected) or a configured default. The language_code field is never empty—handlers ensure a valid code propagates to prevent TTS synthesis errors.

Is real-time streaming translation supported?

Yes. The S2SPipeline uses asynchronous queues between stages, enabling chunked processing. scripts/listen_and_play.py demonstrates real-time microphone input with immediate speaker output. For lower latency, disable auto-detection with a fixed source language and use smaller Whisper models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →