# Speech-to-Speech Model Architecture in the Hugging Face Speech-to-Speech Repository

> Explore the modular four-stage speech-to-speech model architecture in the Hugging Face repository. Discover how VAD STT LLM TTS stages enable low-latency voice assistants.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: architecture
- Published: 2026-08-02

---

**The huggingface/speech-to-speech repository implements a modular, four-stage pipeline (VAD → STT → LLM → TTS) where each stage runs in its own thread and communicates via typed queues, enabling low-latency, pluggable voice assistants.**

The **speech-to-speech** project provides a production-ready foundation for building real-time voice agents. Its architecture decouples voice activity detection, speech recognition, language generation, and speech synthesis into independent, swappable components. This design allows developers to mix open-source and proprietary models while maintaining thread-safe, low-latency audio streaming.

## Four-Stage Pipeline Architecture

The core architecture follows a linear data flow with four sequential handlers. Each stage is implemented as a subclass of `BaseHandler` and operates in its own thread.

| Stage | Responsibility | Default Backend | Key File |
|-------|---------------|-----------------|----------|
| **Voice Activity Detection (VAD)** | Detects speech boundaries, splits audio into turns, emits turn metadata | Silero VAD v5 | [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) |
| **Speech-to-Text (STT)** | Transcribes audio turns into text (with optional streaming) | Parakeet TDT | [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) |
| **Language Model (LLM)** | Generates assistant responses, optionally with tool calls | OpenAI-compatible Responses API | [`src/speech_to_speech/LLM/responses_api_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_model.py) |
| **Text-to-Speech (TTS)** | Synthesizes text back to audio and streams to client | Qwen3-TTS | [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) |

### Data Flow Diagram

```

┌─────────────┐   audio   ┌───────┐   text   ┌───────┐   text   ┌───────┐
│  Audio      │ ───────► │ VAD   │ ─────► │ STT   │ ─────► │ LLM   │
│  Source     │          │       │          │       │          │
│ (mic/socket│          │       │          │       │          │
│  /file)     │          │       │          │       │          │
└─────────────┘          └───────┘          └───────┘          └───────┘
                                                                   │
                                                                   ▼
                                                            audio (synth)
                                                               │
                                                               ▼
                                                          ┌───────┐
                                                          │ TTS   │
                                                          └───────┘
                                                               │
                                                            output

```

## Threading and Queue Communication

Inter-stage communication uses **typed queues** defined in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py). Each handler reads from its input queue and writes to its output queue, enabling asynchronous, non-blocking data flow.

The `build_pipeline` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (line 681) orchestrates this wiring:

- Creates queue instances for each stage transition
- Instantiates handler threads with appropriate kwargs
- Returns a `ThreadManager` that starts and stops all threads

This design ensures that slow stages (like LLM generation) don't block audio capture or playback.

## Realtime Mode and Pipeline Pooling

For production deployments, the repository provides **realtime mode** via `RealtimeServer`. This exposes an OpenAI Realtime-compatible WebSocket API at `/v1/realtime`.

Key implementation details from [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py):

- `_build_realtime_pipeline_unit` (line 480) creates isolated pipeline instances
- Each WebSocket connection gets its own VAD→STT→LLM→TTS chain
- The server pools these units to handle concurrent clients

```bash

# Start the realtime server with default backend stack

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --enable_live_transcription

```

## Turn-Taking and Speculative Interruptions

The VAD handler implements sophisticated **turn management** to minimize latency and handle interruptions gracefully.

### Progressive Transcription

When `enable_realtime_transcription` is set, the VAD emits partial audio chunks before the turn completes. This allows the STT to begin transcription early, reducing perceived latency.

### Speculative Turn Reopening

The `speculative_turns` feature (implemented in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py)) handles cases where users pause briefly then resume speaking:

- `SpeculativeTurnTracker` monitors turn state
- `_should_reopen_current_turn` evaluates whether to continue the current turn
- `_reopen_current_turn` and `_begin_pending_reopen_if_needed` manage the transition

This avoids triggering a full LLM inference cycle for brief pauses, with configurable timeout via `speculative_reopen_ms`.

## Backend Flexibility and Device Adaptation

Every stage supports multiple backend implementations. Selection happens at runtime through CLI flags parsed by `ParsedArguments` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).

### STT Options

The `get_stt_handler` function (line 500) branches on the `--stt` flag:

- `parakeet-tdt` — NVIDIA Parakeet TDT (default)
- `whisper` — OpenAI Whisper
- `faster-whisper` — Optimized Whisper implementation
- `lightning-whisper-mlx` — Apple Silicon optimized
- `mlx-audio-whisper` — MLX-native Whisper
- `paraformer` — Alibaba Paraformer

### LLM Options

`get_llm_handler` (line 549) supports:

- `responses-api` — OpenAI-compatible Responses API (default)
- `transformers` — Hugging Face Transformers
- `mlx-lm` — Apple Silicon optimized inference
- `chat-completions` — OpenAI Chat Completions compatibility

### TTS Options

`get_tts_handler` (line 689) includes:

- `qwen3` — Qwen3-TTS (default)
- `kokoro` — Kokoro TTS
- `pocket` — Pocket TTS (lightweight)
- `chat-tts` — ChatTTS
- `facebook-mms` — Meta MMS

### macOS and MLX Integration

The `prepare_all_args` helper automatically selects MLX backends when `--device mps` is detected. The `--local_mac_optimal_settings` flag applies tuned defaults for Apple Silicon.

## Building a Custom Pipeline

The Python API provides fine-grained control over **speech-to-speech model architecture** composition:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse arguments programmatically

args = parse_arguments()

# Normalize device settings and backend kwargs

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    # ... additional handler kwargs

)

# Initialize communication queues

queues = initialize_queues_and_events()

# Construct the pipeline

pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.vad_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues,
)

pipeline_manager.start()
pipeline_manager.wait()

```

This mirrors the `main()` entry point but allows dynamic configuration without CLI invocation.

## Swapping Backends: Practical Example

Replace the default STT with Faster-Whisper and use a local Transformers LLM:

```bash
speech-to-speech \
    --mode realtime \
    --stt faster-whisper \
    --faster_whisper_stt_model_name large-v2 \
    --llm_backend transformers \
    --model_name meta-llama/Meta-Llama-3.1-8B-Instruct \
    --tts qwen3 \
    --enable_live_transcription

```

The `--stt faster-whisper` branch in `get_stt_handler` instantiates `FasterWhisperSTTHandler` from [`src/speech_to_speech/STT/faster_whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/faster_whisper_handler.py).

## Extending the Architecture

Adding new backends requires minimal changes:

1. Implement a subclass of `BaseHandler` in the appropriate `STT/`, `LLM/`, or `TTS/` directory
2. Register it in the corresponding `get_*_handler` function
3. Add argument dataclass in `src/speech_to_speech/arguments_classes/`
4. Expose CLI flags through `ParsedArguments`

The `arguments_classes/` pattern keeps configuration modular—each handler owns its parameter definitions without modifying core pipeline code.

## Summary

- The **speech-to-speech model architecture** uses four pluggable stages: **VAD → STT → LLM → TTS**
- Each stage runs in its own thread with **typed queue** communication via [`queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/queue_types.py)
- **Realtime mode** provides OpenAI-compatible WebSocket API with per-client pipeline isolation
- **Speculative turn reopening** enables low-latency interruption handling without full pipeline resets
- **Multiple backends** per stage (6 STT, 4 LLM, 5 TTS) support diverse hardware and latency requirements
- **Device adaptation** automatically selects MLX backends on Apple Silicon

## Frequently Asked Questions

### What is the default speech-to-speech pipeline configuration?

The default stack uses **Silero VAD v5** for voice detection, **Parakeet TDT** for speech-to-text, **OpenAI Responses API** for language modeling, and **Qwen3-TTS** for text-to-speech. This configuration prioritizes quality and low latency on GPU-equipped servers. Run with `speech-to-speech --mode realtime` to activate.

### How does the pipeline handle user interruptions?

The **speculative turn** mechanism in `VADHandler` tracks provisional turn boundaries. If speech resumes within `speculative_reopen_ms` milliseconds, the turn continues without triggering a new LLM inference. This logic lives in `_should_reopen_current_turn` and uses `SpeculativeTurnTracker` from [`speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/speculative_turns.py) to manage state.

### Can I run the pipeline on Apple Silicon Macs?

Yes. Set `--device mps` and optionally `--local_mac_optimal_settings`. The `prepare_all_args` helper automatically routes to MLX-compatible backends: `mlx-lm` for LLM inference, `lightning-whisper-mlx` or `mlx-audio-whisper` for STT, and Qwen3-TTS supports MPS directly.

### How do I add a custom TTS model to the pipeline?

Implement a `BaseHandler` subclass in `src/speech_to_speech/TTS/`, add a factory branch in `get_tts_handler` (line 689), create an arguments dataclass in `arguments_classes/`, and register the CLI flag. The handler receives text via its input queue and should emit audio chunks to its output queue in the expected format.