# Speech-to-Speech Tasks Supported by the Hugging Face speech-to-speech Repository

> Explore speech-to-speech tasks like VAD STT LLM generation and TTS supported by Hugging Face. Mix and match pluggable backends easily via CLI for custom solutions.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: getting-started
- Published: 2026-08-01

---

**The huggingface/speech-to-speech repository supports four modular speech-to-speech tasks—Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM) generation, and Text-to-Speech (TTS)—each with pluggable backends ranging from Silero VAD to Qwen3-TTS that can be mixed and matched via CLI arguments.**

The `huggingface/speech-to-speech` package implements a production-ready pipeline for end-to-end voice conversations, decomposing speech-to-speech tasks into four independent stages that communicate via async queues. According to the source code in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), each task runs in its own thread and can be swapped without modifying core pipeline logic, enabling deployment across CUDA, CPU, and Apple Silicon hardware.

## The Four Core Speech-to-Speech Tasks

The repository organizes functionality into four distinct handlers, each defined in dedicated subdirectories under `src/speech_to_speech/`.

### Voice Activity Detection (VAD)

**VAD** detects when a user starts and stops speaking, providing turn-taking information to downstream components. The default implementation uses **Silero VAD v5** via the handler defined in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py). This handler processes raw microphone streams and emits `SpeechStartedEvent` and `SpeechStoppedEvent` objects, forwarding audio chunks as `VADAudio` items to the STT stage.

### Speech-to-Text (STT)

**STT** converts spoken audio into textual transcripts, optionally streaming partial results for real-time displays. The default backend is **Parakeet TDT 0.6B v3**, implemented in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py). Alternative backends include:
- **Whisper** (via Transformers)
- **Faster Whisper**
- **Lightning Whisper MLX** (Apple Silicon)
- **MLX Audio Whisper** (Apple Silicon)
- **Paraformer** (FunASR)

The STT handler pushes transcribed output as a `TextEventItem` onto the LLM queue.

### Language Model (LLM)

**LLM** generates the assistant’s response text, optionally streaming tool-call events. The default backend uses the **OpenAI-compatible Responses API**, with handler logic referenced in [`src/speech_to_speech/pipeline/handler_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/handler_types.py). Alternative backends include:
- **Transformers** (CUDA/CPU)
- **mlx-lm** (Apple Silicon)
- Self-hosted servers via **Responses API** or **Chat Completions** (e.g., vLLM, llama.cpp)

This stage consumes `TextEventItem` objects from the STT stage and produces new `TextEventItem` instances for the TTS stage.

### Text-to-Speech (TTS)

**TTS** synthesizes the LLM’s textual output into audio waveforms streamed back to the client. The default is **Qwen3-TTS**, using GGML on Linux and `mlx-audio` on macOS, implemented in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py). Alternative backends include:
- **Kokoro-82M** (CUDA/CPU, Apple Silicon)
- **Pocket TTS** (CPU/CUDA)
- **ChatTTS** (CUDA/CPU)
- **MMS TTS** (Transformers)

## Task Pipeline Architecture

The four speech-to-speech tasks execute as a cascade of independent handlers linked by async queues. As implemented in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), the flow is:

1. **VAD** processes raw audio, detects speech boundaries, and forwards `VADAudio` chunks.
2. **STT** receives chunks, transcribes them, and pushes a `TextEventItem` to the LLM queue.
3. **LLM** consumes the transcript, generates a response, and pushes a `TextEventItem` to the TTS queue.
4. **TTS** synthesizes the final audio output for delivery via WebSocket or local playback.

Each handler inherits from the base `Handler` class defined in [`baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/baseHandler.py), allowing seamless backend substitution through CLI flags like `--stt`, `--llm_backend`, and `--tts`.

## Execution Modes for Speech-to-Speech Tasks

The repository supports multiple runtime configurations defined in [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py):

- **Realtime mode** (`--mode realtime`): Uses the OpenAI Realtime protocol over WebSocket/WebRTC, enabling live transcription (`--enable_live_transcription`) and speculative turn handling.
- **Local mode** (`--mode local`): Runs the entire pipeline on a single machine without the Realtime protocol wrapper.
- **Raw-WebSocket** (`--mode raw-websocket`): Streams raw PCM audio without protocol overhead.
- **Socket** (`--mode socket`): Minimal TCP-socket streaming for embedded applications.

## Configuring Speech-to-Speech Tasks: Code Examples

### Running the Default Realtime Server

Install the core package and start the server:

```bash
pip install speech-to-speech

# Starts server on ws://localhost:8765/v1/realtime

speech-to-speech

```

### Switching to Faster Whisper for STT

Select an alternative STT backend using the `--stt` flag, which loads the handler from [`src/speech_to_speech/STT/faster_whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/faster_whisper_handler.py):

```bash
speech-to-speech \
    --stt faster-whisper \
    --stt_model_name base \
    --mode realtime

```

### Using Local LLM Inference with mlx-lm

Deploy Apple Silicon-optimized LLM inference via the `--llm_backend` flag:

```bash
speech-to-speech \
    --llm_backend mlx-lm \
    --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16 \
    --mode realtime

```

The `mlx-lm` backend is defined in [`src/speech_to_speech/pipeline/handler_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/handler_types.py).

### Configuring Pocket TTS for Speech Synthesis

Switch to the Pocket TTS backend implemented in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py):

```bash
speech-to-speech \
    --tts pocket \
    --pocket_tts_voice jean \
    --pocket_tts_device cpu \
    --mode realtime

```

## Summary

- The huggingface/speech-to-speech repository implements **four modular speech-to-speech tasks**: VAD, STT, LLM, and TTS.
- Each task supports **multiple interchangeable backends** selectable via CLI flags, including Silero VAD, Parakeet/Whisper STT, OpenAI/Transformers/MLX LLMs, and Qwen3/Kokoro TTS.
- Tasks communicate via **async queues** in a multi-threaded pipeline orchestrated by [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).
- The system supports **four execution modes**: realtime (OpenAI protocol), local, raw-websocket, and socket.

## Frequently Asked Questions

### What is the default STT backend in the speech-to-speech repository?

The default Speech-to-Text backend is **Parakeet TDT 0.6B v3**, implemented in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py). The system also supports Whisper, Faster Whisper, Lightning Whisper MLX, and Paraformer via the `--stt` CLI flag.

### Can I run speech-to-speech tasks entirely on Apple Silicon without CUDA?

Yes. The repository supports Apple Silicon-optimized backends for all four tasks: Silero VAD for voice detection, Lightning Whisper MLX or MLX Audio Whisper for transcription, `mlx-lm` for language modeling, and `mlx-audio` for Qwen3-TTS synthesis.

### How do I enable live transcription streaming during conversations?

Pass the `--enable_live_transcription` flag when starting in realtime mode (`--mode realtime`). This configures the STT handler to emit partial transcripts while speech is still ongoing, rather than waiting for the VAD to detect speech completion.

### Where are the CLI arguments for each speech-to-speech task defined?

Argument definitions for VAD, STT, LLM, and TTS handlers are located in `src/speech_to_speech/arguments_classes/`, with files such as [`whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_arguments.py) defining the parameters for specific backends.