# How to Convert Speech from One Language to Another with Hugging Face Speech-to-Speech

> Easily convert speech to another language using Hugging Face speech-to-speech. Learn the three-stage pipeline STT LM TTS for seamless voice translation.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-02

---

**The speech-to-speech library converts spoken input to spoken output in a different language through a three-stage pipeline: Speech-to-Text (STT) detects the source language, a Language Model translates the text, and Text-to-Speech (TTS) synthesizes the target language audio—all coordinated by `S2SPipeline` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).**

This guide walks through the complete process of building a multilingual speech translator using the huggingface/speech-to-speech repository. The library provides a modular, full-duplex pipeline that chains together STT, translation, and TTS components with automatic language code propagation.

## Understanding the Multilingual Pipeline Architecture

The speech-to-speech translation flow consists of three core stages that communicate via typed messages carrying `language_code` fields:

| Stage | Handler Location | Function |
|-------|-----------------|----------|
| **Speech-to-Text** | [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) | Captures audio, runs Whisper, emits `Transcription` with detected source `language_code` |
| **Language Model / Translation** | [`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py) via `LanguageModelHandler` | Receives transcription, generates translation, returns `LLMResponseChunk` with target `language_code` |
| **Text-to-Speech** | [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) | Synthesizes audio using the target `language_code` from the LM stage |

The `S2SPipeline` class in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) orchestrates these stages by creating async queues and wiring handlers together. It automatically propagates language metadata through each message type defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py).

## Language Handling Mechanics

### Source Language Detection

When `WhisperSTTHandler` receives `--language auto`, Whisper detects the spoken language and stores it in `handler.last_language`. This value attaches to the `Transcription` message as `language_code` (e.g., `"en"`, `"fr"`, `"de"`).

### Target Language Selection

The target language flows through the pipeline via `runtime_config.language_code`. When `LanguageModelHandler` builds a `GenerateResponseRequest`, the `language_code` field signals which language the LLM should generate. Handlers normalize aliases (e.g., `"English"` → `"en"`; `"nb"` → `"no"`) and fall back to the last valid language if unsupported.

## Running Speech-to-Speech Translation

### Command-Line Approach with listen_and_play.py

The repository includes [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py) as a complete working example:

```python

# scripts/listen_and_play.py

import argparse
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelArguments

parser = argparse.ArgumentParser(description="Speech-to-speech translation")
parser.add_argument("--source-language", default="auto", help="Source language (auto for detection)")
parser.add_argument("--target-language", required=True, help="Target language code, e.g. 'fr' or 'de'")
args = parser.parse_args()

runtime_config = {
    "language_code": args.target_language,   # Target language propagated to LM and TTS

    "stt": WhisperSTTArguments(language=args.source_language),
    "tts": Qwen3TTSArguments(),
    "lm": LanguageModelArguments(),
}

pipeline = S2SPipeline(runtime_config)
pipeline.run()  # Starts interactive microphone-to-speaker loop

```

The pipeline executes as follows:

1. **WhisperSTTHandler** detects source language and emits `Transcription(language_code="en")`
2. **LanguageModelHandler** prompts the LLM to translate, returns `LLMResponseChunk(language_code="fr")`
3. **Qwen3TTSHandler** synthesizes French speech using the target `language_code`

### Programmatic API Without CLI

For integration into larger applications, instantiate `S2SPipeline` directly:

```python
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelArguments

# Configure for Spanish output

config = {
    "language_code": "es",
    "stt": WhisperSTTArguments(language="auto"),
    "tts": Qwen3TTSArguments(),
    "lm": LanguageModelArguments(),
}
pipeline = S2SPipeline(config)

# Feed audio from file, microphone, or stream

with open("input_english.wav", "rb") as f:
    audio_bytes = f.read()
pipeline.feed_audio(audio_bytes)

# Retrieve translated audio

translated_audio = pipeline.get_output_audio()

```

### Fixed Source Language (Skip Auto-Detection)

For known input languages, disable detection to reduce latency:

```python
config["stt"] = WhisperSTTArguments(language="de")  # German input

config["language_code"] = "en"                      # English output

pipeline = S2SPipeline(config)

```

## Key Source Files for Language Conversion

| File | Path | Purpose |
|------|------|---------|
| Pipeline orchestrator | [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) | `S2SPipeline` class creates queues and runs the async event loop |
| Message types | [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py) | `Transcription`, `TTSInput`, `AudioOutput` with `language_code` fields |
| Event types | [`src/speech_to_speech/pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/events.py) | `TranscriptionCompletedEvent` and other pipeline events |
| Whisper STT | [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) | Speech recognition with language detection |
| Qwen-3 TTS | [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) | Speech synthesis respecting target `language_code` |
| LLM translation | [`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py) | Translation via `LanguageModelHandler` |
| Demo script | [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py) | Complete end-to-end example |

## Swappable Backend Components

The modular design allows mixing STT, LM, and TTS backends. Replace argument classes to use alternatives:

- **STT**: `WhisperSTTArguments`, or custom handlers in `src/speech_to_speech/STT/`
- **LM**: Any OpenAI-compatible API via `LanguageModelArguments`, or local models
- **TTS**: `Qwen3TTSArguments`, or other handlers in `src/speech_to_speech/TTS/`

All handlers respect the `language_code` field in their input messages, ensuring consistent language propagation regardless of backend choice.

## Summary

- The **speech-to-speech library** implements translation through a **three-stage pipeline** coordinated by `S2SPipeline` in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)
- **Source language** is detected by Whisper or specified explicitly; stored in `Transcription.language_code`
- **Target language** is set via `runtime_config.language_code` and propagated through LM and TTS stages
- **Message types** in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py) carry `language_code` fields ensuring metadata flows through the pipeline
- **[`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py)** provides a complete CLI example for interactive translation
- **Backend components are swappable**—mix Whisper STT, any OpenAI-compatible LLM, and Qwen-3 TTS or alternatives

## Frequently Asked Questions

### What language codes does the speech-to-speech library support?

The library uses standard ISO 639-1 codes (`"en"`, `"fr"`, `"de"`, `"es"`, `"zh"`, etc.). Handlers in [`whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_handler.py) and [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) normalize aliases automatically, mapping `"English"` to `"en"` and `"nb"` (Norwegian Bokmål) to `"no"`. Unsupported codes trigger fallback to the last valid language.

### Can I use a different translation model instead of the default LLM?

Yes. The `LanguageModelArguments` class accepts any OpenAI-compatible API endpoint. Configure the `base_url` and `model` parameters to switch to local models (via vLLM, llama.cpp) or alternative APIs. The `LanguageModelHandler` builds translation prompts and returns `LLMResponseChunk` with the target `language_code` regardless of backend.

### How does the pipeline handle language detection failures?

If Whisper cannot determine the source language, [`whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_handler.py) falls back to `handler.last_language` (previously detected) or a configured default. The `language_code` field is never empty—handlers ensure a valid code propagates to prevent TTS synthesis errors.

### Is real-time streaming translation supported?

Yes. The `S2SPipeline` uses asynchronous queues between stages, enabling chunked processing. [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py) demonstrates real-time microphone input with immediate speaker output. For lower latency, disable auto-detection with a fixed source language and use smaller Whisper models.