# How the VAD → STT → LLM → TTS Pipeline Works in Hugging Face Speech-to-Speech

> Understand the VAD STT LLM TTS pipeline in Hugging Face Speech-to-Speech. Explore this asynchronous, real-time architecture for low-latency voice processing.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: architecture
- Published: 2026-07-08

---

**The VAD → STT → LLM → TTS pipeline in huggingface/speech-to-speech is an asynchronous, queue-driven architecture that processes live microphone audio through four sequential stages—Voice Activity Detection, Speech-to-Text, Large Language Model, and Text-to-Speech—using threaded handlers that communicate via typed Python queues to enable real-time, low-latency conversation.**

The huggingface/speech-to-speech repository implements a modular voice conversation system that converts spoken input into synthesized responses. At its core lies a **VAD → STT → LLM → TTS pipeline** that orchestrates audio processing through distinct, interchangeable components. This architecture enables real-time speech-to-speech interaction by buffering audio fragments, transcribing speech, generating contextual responses, and synthesizing audio output in a continuous stream.

## The Four Pipeline Stages

The pipeline consists of four specialized handlers, each implemented as a standalone module that adheres to the `BaseHandler` interface.

### Voice Activity Detection (VAD)

The **VAD** stage detects when a user starts and stops speaking, buffering audio fragments for downstream processing. Implemented in [`speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/VAD/vad_handler.py), this handler uses the Silero VAD model (wrapped by [`speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/VAD/vad_iterator.py)) to emit `SpeechStartedEvent` and `SpeechStoppedEvent` markers. When the `enable_realtime_transcription` flag is active, the VAD handler can publish interim transcription events via `text_output_queue` while still listening, enabling the OpenAI realtime server mode found in [`api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/api/openai_realtime/server.py).

### Speech-to-Text (STT)

The **STT** stage converts voiced audio chunks into raw text. Multiple backends are supported, including Whisper, Faster-Whisper, Paraformer, and MLX Audio variants. The [`speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/STT/whisper_stt_handler.py) provides a concrete implementation using OpenAI Whisper, while [`speech_to_speech/STT/base_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/STT/base_stt_handler.py) defines the common interface. Each STT handler pulls `VADOutItem` objects from the input queue and pushes `STTOutItem` transcriptions to the next stage.

### Large Language Model (LLM)

The **LLM** stage receives transcriptions and generates textual responses. Two primary implementations exist: [`speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/language_model.py) for local transformer-based models (such as Qwen-3), and [`speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/responses_api_language_model.py) for OpenAI API-compatible endpoints. The handler enriches input with chat history and returns streaming `LMOutItem` objects. Before passing to TTS, [`speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/lm_output_processor.py) post-processes the output to handle speculative turns and text compaction.

### Text-to-Speech (TTS)

The **TTS** stage synthesizes the LLM response into audible speech. Supported models include Qwen-3, Pocket, Kokoro, ChatTTS, and Facebook MMS. The [`speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/TTS/qwen3_tts_handler.py) demonstrates the implementation pattern for local models, consuming `TTSInItem` objects and producing `AudioOutItem` speech data. Like other stages, TTS handlers are swappable via command-line arguments parsed in the main pipeline.

## Pipeline Orchestration and Data Flow

The entire pipeline is assembled and executed by [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py). This orchestration layer manages the complex data flow between stages using a queue-based architecture.

### Queue-Driven Thread Management

The `initialize_queues_and_events()` function creates typed `queue.Queue` objects that link each stage. Every handler runs in its own thread via `ThreadManager` ([`speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/utils/thread_manager.py)), pulling from an input queue and pushing to an output queue. This design creates natural back-pressure: if a downstream consumer slows down, the upstream stage blocks on `put()` until space becomes available.

The data flow follows this path:

```

Audio chunks → VADHandler → VADOutItem → STTHandler → STTOutItem 
               → TranscriptionNotifier → LLMHandler → LMOutItem 
               → LMOutputProcessor → TTSInItem → TTSHandler → AudioOutItem

```

### Handler Selection and Registration

The `parse_arguments()` function maps user-provided `--stt`, `--llm_backend`, and `--tts` options to concrete classes via `get_stt_handler()`, `get_llm_handler()`, and `get_tts_handler()`. Adding a new model requires only implementing the `BaseHandler` interface and registering it in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).

## Key Architectural Features

### Asynchronous, Back-Pressure-Aware Design

Each stage operates independently with thread-safe queues mediating communication. The asynchronous design ensures that VAD can continue buffering new audio while the TTS stage synthesizes previous responses, maintaining pipeline throughput.

### Speculative Turn Handling

The `SpeculativeTurnTracker` (located in [`pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline/speculative_turns.py)) enables the system to treat partial LLM responses as tentative "turns." This allows early audio playback while the LLM continues generating text, reducing perceived latency. The VAD handler can merge speculative audio prefixes with final output to eliminate gaps between turns.

### Graceful Shutdown

A `CancelScope` and signal handlers in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) ensure all threads stop cleanly when the process receives SIGINT or SIGTERM signals, preventing audio device locks or zombie threads.

## Implementation Examples

### Running a Local Pipeline

To launch the complete pipeline with Whisper STT and Qwen-3 TTS:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse CLI flags or build ParsedArguments manually

args = parse_arguments()

# Prepare device-specific arguments

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    # ... other handler kwargs

)

# Initialize queues and build pipeline

queues = initialize_queues_and_events()
pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    # ... other kwargs

    queues,
)

# Start and block until shutdown

pipeline_manager.start()
pipeline_manager.wait()

```

### WebSocket Realtime Server

To expose the pipeline via OpenAI-compatible WebSocket:

```python
from speech_to_speech.s2s_pipeline import build_pipeline, parse_arguments, prepare_all_args, initialize_queues_and_events

args = parse_arguments()
prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    # ... other handlers

)

queues = initialize_queues_and_events()
pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    # ... other handlers

    queues,
)

# Listens on ws://127.0.0.1:8765 by default

pipeline_manager.start()
pipeline_manager.wait()

```

## Summary

- The **VAD → STT → LLM → TTS pipeline** processes audio through four distinct stages, each implemented as a threaded handler in the huggingface/speech-to-speech repository.
- **Queue-based communication** via `initialize_queues_and_events()` enables asynchronous, back-pressure-aware data flow between stages.
- **Speculative turn tracking** reduces latency by allowing early audio playback while the LLM completes generation.
- **Extensible architecture** supports multiple backends (Whisper, Faster-Whisper, Qwen-3, Kokoro) via command-line selection in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).
- **Graceful shutdown** is handled through `CancelScope` and signal management to ensure clean thread termination.

## Frequently Asked Questions

### How does the VAD handler know when to start and stop recording?

The VAD handler uses the Silero VAD model through [`speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/VAD/vad_iterator.py) to detect voice activity thresholds. When audio levels exceed the threshold, it emits a `SpeechStartedEvent` and begins buffering chunks into a `VADOutItem` queue. When silence persists beyond a configured timeout, it emits a `SpeechStoppedEvent` and sends the complete buffered audio to the STT handler.

### Can I swap individual components without modifying the pipeline code?

Yes. The `parse_arguments()` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) dynamically maps command-line flags such as `--stt`, `--llm_backend`, and `--tts` to handler classes via `get_stt_handler()`, `get_llm_handler()`, and `get_tts_handler()`. To add a new model, implement the `BaseHandler` interface and register the mapping in the argument parser.

### What prevents the pipeline from overwhelming system resources?

The queue-based architecture implements natural back-pressure. Each stage blocks on `queue.put()` when the downstream queue is full, automatically throttling upstream processing. Additionally, the `ThreadManager` monitors thread health, and the `CancelScope` ensures immediate resource cleanup during shutdown.

### How does realtime transcription work alongside the main pipeline?

When `enable_realtime_transcription` is set to `True`, the VAD handler in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py) emits interim text events through a separate `text_output_queue` while continuing to buffer audio for the full STT pass. This integrates with the OpenAI realtime WebSocket server in [`api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/api/openai_realtime/server.py), allowing clients to receive partial transcriptions before the LLM response completes.