# Architecture of the Default Speech-to-Speech Models in Hugging Face's Pipeline

> Explore the default speech-to-speech architecture in Hugging Face's pipeline. Understand how VAD, TDT, LLM, and TTS handlers create a real-time speech conversion system.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: architecture
- Published: 2026-08-01

---

**The default architecture chains four specialized handlers—VADHandler, ParakeetTDTSTTHandler, ResponsesApiModelHandler, and Qwen3TTSHandler—into a real-time pipeline that converts speech to text, processes it through a remote LLM, and synthesizes the response back into speech.**

The `huggingface/speech-to-speech` repository implements a modular, low-latency inference system designed for conversational AI applications. When invoked without custom overrides, the pipeline automatically instantiates a specific quartet of models optimized for speed and quality. Understanding this default architecture helps developers debug latency issues and swap components effectively.

## The Four-Stage Default Pipeline

The architecture follows a strict linear data flow defined in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). Each stage communicates via strongly-typed queues, ensuring type safety across asynchronous boundaries.

### Voice Activity Detection (VAD)

The pipeline begins with `VADHandler`, located in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py). This component captures raw audio chunks and determines when speech starts and stops.

In `_build_pipeline_handlers` (lines 16–22 of [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)), the VAD handler is instantiated first. It performs two critical functions: it feeds active audio segments downstream to the STT handler and simultaneously forwards spoken prompts to the `TranscriptionNotifier` for UI display.

### Speech-to-Text (STT) - Parakeet-TDT

The default speech recognizer is `ParakeetTDTSTTHandler`, implemented in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py). This handler uses the *parakeet-tdt* model, selected specifically for its fast, low-latency transcription capabilities.

The handler is created by `get_stt_handler` (lines 81–88) according to the default arguments defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py). Recognized text is packaged into `STTOutItem` objects and queued for the language model stage.

### Language Model (LLM) - Responses API

The default LLM backend is `ResponsesApiModelHandler`, found in [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py). Rather than running a local model, this handler connects to a remote OpenAI-compatible API endpoint (configured via the *responses-api* backend).

Built by `get_llm_handler` (lines 88–109), this component consumes text prompts from the STT stage, generates responses, and pushes them downstream as `LMOutItem` objects. This remote architecture keeps local resource usage low while providing access to large-scale models.

### Text-to-Speech (TTS) - Qwen3

The final stage uses `Qwen3TTSHandler` from [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), leveraging the *qwen3* model accelerated via MLX. Before reaching TTS, raw LLM output passes through `LMOutputProcessor` (instantiated in `_build_pipeline_handlers`, lines 53–57), which handles speculative-turn detection and optional text events.

The TTS handler is constructed by `get_tts_handler` (lines 98–115) and outputs synthesized audio that is then transmitted to the client via the selected communication mode.

## Pipeline Orchestration and Communication

### Queue Types and Typed Communication

All inter-stage communication uses strongly-typed queue items defined in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py). The four primary data structures are:

- `AudioInItem` – Raw audio from VAD to STT
- `STTOutItem` – Transcribed text from STT to LLM
- `LMOutItem` – Generated responses from LLM to TTS
- `TTSInItem` – Processed text ready for speech synthesis

This typing ensures that handlers receive expected data formats without runtime ambiguity.

### Event Synchronization

Pipeline control relies on synchronization primitives defined in [`src/speech_to_speech/pipeline/control.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/control.py). Key events include `stop_event` for graceful shutdown and `should_listen` for managing the listening state. These events coordinate the four handlers across different execution threads or processes.

In *realtime* mode (the default), the system creates isolated pipeline pools via `_build_realtime_pipeline_unit`, where each session maintains its own VAD, STT, LLM, and TTS instances to prevent cross-contamination between simultaneous users.

## Running the Default Pipeline

To launch the default architecture with all standard components, use the following command:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3

```

For programmatic access, instantiate the pipeline through the Python API:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments, 
    build_pipeline, 
    initialize_queues_and_events,
    prepare_all_args
)

# Load default configurations

args = parse_arguments()

# Initialize device-specific and model-specific arguments

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Create queues and events

queues = initialize_queues_and_events()

# Build and start the pipeline

pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues,
)

pipeline_manager.start()
pipeline_manager.wait()

```

## Summary

- The default architecture implements a **VAD → Parakeet-TDT STT → Responses-API LLM → Qwen3 TTS** pipeline.
- Component selection is governed by `ModuleArguments` in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py).
- Inter-process communication uses typed queues (`AudioInItem`, `STTOutItem`, `LMOutItem`, `TTSInItem`) from [`queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/queue_types.py).
- Event synchronization relies on [`control.py`](https://github.com/huggingface/speech-to-speech/blob/main/control.py) for managing pipeline state across concurrent sessions.
- The system supports four execution modes (`local`, `socket`, `raw-websocket`, `realtime`), with realtime mode creating isolated pipeline pools for multi-user scenarios.

## Frequently Asked Questions

### Can I replace the default Parakeet-TDT recognizer with Whisper?

Yes. The pipeline supports multiple STT handlers including Whisper variants. Override the `--stt` argument with `whisper`, `faster-whisper`, or `mlx-audio-whisper` to switch recognizers. The `get_stt_handler` factory function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) automatically instantiates the appropriate handler class based on this argument.

### Why does the default architecture use a remote LLM instead of a local model?

The default `ResponsesApiModelHandler` minimizes local GPU/CPU memory requirements by offloading inference to remote API endpoints. This allows the pipeline to run on edge devices while still leveraging large language models. You can switch to local LLM inference by changing `--llm_backend` to a local-compatible handler, though this requires additional configuration in the language model handler arguments.

### How does the pipeline handle multiple simultaneous conversations?

In the default `realtime` mode, the `_build_realtime_pipeline_unit` function creates isolated pipeline instances for each connection. Each instance maintains its own VADHandler, ParakeetTDTSTTHandler, ResponsesApiModelHandler, and Qwen3TTSHandler with separate queue sets, ensuring that audio and text from different users never intermingle.

### What is the purpose of the LMOutputProcessor between the LLM and TTS?

The `LMOutputProcessor` (instantiated in `_build_pipeline_handlers`) performs intermediate text processing before synthesis. It handles speculative-turn detection—identifying when the model has generated a complete thought versus a partial sentence—and manages optional text events. This preprocessing ensures that the TTS handler receives properly segmented text for natural-sounding speech synthesis.