# Speech-to-Speech Model Architecture: Inside the Hugging Face Real-Time Voice Pipeline

> Explore the four-stage speech-to-speech model architecture from Hugging Face: VAD, STT, LLM, and TTS. Enable real-time voice conversations with this low-latency pipeline.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: architecture
- Published: 2026-07-07

---

**The huggingface/speech-to-speech repository implements a modular, four-stage pipeline architecture consisting of Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS) handlers that communicate via thread-safe queues to enable low-latency, real-time voice conversation.**

The open-source speech-to-speech model architecture provides a flexible framework for building real-time voice assistants on consumer hardware. This system chains together specialized handlers in a producer-consumer pattern, allowing developers to swap individual components—such as replacing Whisper with Parakeet TDT or switching between Qwen3-TTS and Kokoro—without modifying the core orchestration logic in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

## Four-Stage Modular Architecture

The architecture processes audio through four sequential stages, each running as an independent thread with typed input and output queues defined in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py).

### Voice Activity Detection (VAD)

The **VAD** stage detects when users start and stop speaking, splitting continuous audio into discrete turns. By default, the pipeline uses **Silero VAD v5** via the `VADHandler` class in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py). This handler manages turn boundaries and emits metadata events that downstream stages consume.

When `enable_realtime_transcription` is active, the VAD emits progressive audio chunks for live transcription rather than waiting for complete turns. The handler also supports **speculative turn tracking**—if the user resumes speaking within a configurable timeout (`speculative_reopen_ms`), the VAD reopens the current turn using methods like `_should_reopen_current_turn` and `_reopen_current_turn`, avoiding the latency of starting a new LLM inference.

### Speech-to-Text (STT)

The **STT** stage transcribes audio segments into text. The default backend is **Parakeet TDT**, though the architecture supports multiple implementations selected via the `get_stt_handler` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (around line 500).

Available backends include:
- **Parakeet TDT** – Optimized for real-time streaming
- **Whisper** – OpenAI's original implementation
- **Faster-Whisper** – Optimized Whisper variant
- **Lightning-Whisper-MLX** – Apple Silicon optimized
- **Paraformer** – Alternative architecture

Each implementation resides in `src/speech_to_speech/STT/` and inherits from `BaseHandler`, ensuring consistent queue-based interfaces.

### Language Model (LLM)

The **LLM** stage generates text responses from the transcribed user input. By default, the pipeline uses the **OpenAI-compatible Responses API**, though it supports local inference via Transformers, `mlx-lm` (for Apple Silicon), and Chat-Completions compatible endpoints.

The `get_llm_handler` function (around line 549 in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)) instantiates the appropriate handler from `src/speech_to_speech/LLM/`, which includes `LanguageModelHandler`, `ResponsesApiModelHandler`, and `VisionLanguageModelHandler`. This stage can stream partial responses to the TTS stage before generation completes, reducing perceived latency.

### Text-to-Speech (TTS)

The final **TTS** stage synthesizes the LLM output into audio streams sent to the client. The default backend is **Qwen3-TTS**, with alternatives including **Kokoro**, **Pocket**, **ChatTTS**, and **Facebook-MMS**.

The `get_tts_handler` function (around line 689 in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)) selects the appropriate implementation from `src/speech_to_speech/TTS/`. Each handler converts text chunks into audio frames, enabling true streaming synthesis where the user hears audio before the LLM finishes generating the complete response.

## Threading and Data Flow

The pipeline orchestration relies on a **thread-per-stage** architecture coordinated by the `build_pipeline` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (around line 681). This function:

1. Creates typed queues for inter-stage communication
2. Instantiates each handler (VAD, STT, LLM, TTS)
3. Returns a `ThreadManager` that starts and supervises all threads

```

Audio Source → [VAD Thread] → [STT Thread] → [LLM Thread] → [TTS Thread] → Output

```

Each handler runs in a loop, blocking on its input queue, processing data, and placing results on the next stage's queue. This design ensures that slow operations (like LLM generation) do not block audio capture or VAD processing.

## Realtime Mode and Pipeline Pools

For production deployments, the **RealtimeServer** creates isolated pipeline instances for each concurrent client. The `_build_realtime_pipeline_unit` function (around line 480 in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)) constructs these pipeline units, which expose the OpenAI Realtime WebSocket API at `/v1/realtime`.

Each client receives a dedicated set of threads and queues, ensuring that heavy processing by one user does not affect others' latency.

## Speculative Turn-Taking for Interruption Handling

Low-latency interruption handling relies on the **SpeculativeTurnTracker** in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py). When `speculative_turns` is enabled:

- The VAD tracks speculative turn IDs alongside confirmed turns
- If speech resumes within `speculative_reopen_ms`, the system reopens the existing turn rather than creating a new one
- This prevents the LLM from starting fresh inferences when users briefly pause mid-thought

The logic resides in `VADHandler` methods like `_begin_pending_reopen_if_needed`, which coordinate with the speculative tracker to minimize response latency during natural conversation pauses.

## Implementation Examples

### Running the Default Realtime Server

Configure the complete four-stage pipeline with default backends using the CLI:

```bash
pip install speech-to-speech

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --enable_live_transcription

```

This command maps to the `parse_arguments` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), which validates the configuration and prepares the `ParsedArguments` dataclass.

### Building a Local Pipeline via Python API

For custom integrations, instantiate the pipeline programmatically:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse configuration

args = parse_arguments()

# Normalize device arguments and backend-specific parameters

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    # ... other handler kwargs

)

# Initialize communication queues

queues = initialize_queues_and_events()

# Construct thread manager

pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues,
)

# Start all handler threads

pipeline_manager.start()
pipeline_manager.wait()

```

This mirrors the `main()` entry point in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), giving you direct control over queue initialization and thread lifecycle.

### Swapping Backends (Faster-Whisper Example)

Replace the default STT backend via command-line flags:

```bash
speech-to-speech \
    --mode realtime \
    --stt faster-whisper \
    --faster_whisper_stt_model_name large-v2 \
    --llm_backend transformers \
    --model_name meta-llama/Meta-Llama-3.1-8B-Instruct \
    --tts qwen3

```

The `--stt faster-whisper` flag triggers the `elif module_kwargs.stt == "faster-whisper"` branch in `get_stt_handler`, loading the appropriate handler from [`src/speech_to_speech/STT/faster_whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/faster_whisper_handler.py).

## Summary

- **The architecture comprises four sequential stages** (VAD → STT → LLM → TTS) communicating through thread-safe queues defined in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py).
- **Each stage runs in its own thread**, managed by `build_pipeline` and supervised by `ThreadManager` to prevent blocking.
- **Backends are fully interchangeable** via `get_stt_handler`, `get_llm_handler`, and `get_tts_handler` functions, supporting both cloud APIs and local MLX-optimized models.
- **Speculative turn tracking** in `VADHandler` and `SpeculativeTurnTracker` enables low-latency interruption handling without restarting LLM inference.
- **Realtime mode** creates isolated pipeline pools per client, exposing an OpenAI-compatible WebSocket API via `RealtimeServer`.

## Frequently Asked Questions

### What is the default speech-to-speech model architecture in the Hugging Face repository?

The default architecture chains **Silero VAD v5** for voice detection, **Parakeet TDT** for speech recognition, the **OpenAI Responses API** for language modeling, and **Qwen3-TTS** for speech synthesis. These components communicate through Python's `queue.Queue` structures managed by the `build_pipeline` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

### How does the pipeline handle interruptions during real-time conversation?

The system uses **speculative turn tracking** via the `SpeculativeTurnTracker` class in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py). When enabled, the VAD handler tracks when turns might reopen using `_should_reopen_current_turn`. If the user resumes speaking within the `speculative_reopen_ms` window, the existing turn continues rather than starting a new LLM inference, maintaining conversational continuity with minimal latency.

### Can I use local models instead of cloud APIs in this architecture?

Yes. The architecture supports local inference through multiple backends. For LLMs, use `--llm_backend transformers` or `--llm_backend mlx-lm` (for Apple Silicon). For STT, select `--stt faster-whisper` or `--stt mlx-audio-whisper`. For TTS, options include `--tts kokoro` and `--tts qwen3` with `--device mps`. The `prepare_all_args` function automatically applies macOS-specific optimizations when `--local_mac_optimal_settings` is enabled.

### What files control the queue communication between pipeline stages?

Queue definitions reside in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py), which provides typed `queue.Queue` instances for audio, text, and control messages. The `initialize_queues_and_events` function creates these queues, which are then passed to `build_pipeline` to wire together the VAD, STT, LLM, and TTS handlers. Each handler runs in its own thread, blocking on its input queue and placing results on the output queue for the next stage.