How the VAD → STT → LLM → TTS Pipeline Works in Hugging Face Speech-to-Speech
The VAD → STT → LLM → TTS pipeline in huggingface/speech-to-speech is an asynchronous, queue-driven architecture that processes live microphone audio through four sequential stages—Voice Activity Detection, Speech-to-Text, Large Language Model, and Text-to-Speech—using threaded handlers that communicate via typed Python queues to enable real-time, low-latency conversation.
The huggingface/speech-to-speech repository implements a modular voice conversation system that converts spoken input into synthesized responses. At its core lies a VAD → STT → LLM → TTS pipeline that orchestrates audio processing through distinct, interchangeable components. This architecture enables real-time speech-to-speech interaction by buffering audio fragments, transcribing speech, generating contextual responses, and synthesizing audio output in a continuous stream.
The Four Pipeline Stages
The pipeline consists of four specialized handlers, each implemented as a standalone module that adheres to the BaseHandler interface.
Voice Activity Detection (VAD)
The VAD stage detects when a user starts and stops speaking, buffering audio fragments for downstream processing. Implemented in speech_to_speech/VAD/vad_handler.py, this handler uses the Silero VAD model (wrapped by speech_to_speech/VAD/vad_iterator.py) to emit SpeechStartedEvent and SpeechStoppedEvent markers. When the enable_realtime_transcription flag is active, the VAD handler can publish interim transcription events via text_output_queue while still listening, enabling the OpenAI realtime server mode found in api/openai_realtime/server.py.
Speech-to-Text (STT)
The STT stage converts voiced audio chunks into raw text. Multiple backends are supported, including Whisper, Faster-Whisper, Paraformer, and MLX Audio variants. The speech_to_speech/STT/whisper_stt_handler.py provides a concrete implementation using OpenAI Whisper, while speech_to_speech/STT/base_stt_handler.py defines the common interface. Each STT handler pulls VADOutItem objects from the input queue and pushes STTOutItem transcriptions to the next stage.
Large Language Model (LLM)
The LLM stage receives transcriptions and generates textual responses. Two primary implementations exist: speech_to_speech/LLM/language_model.py for local transformer-based models (such as Qwen-3), and speech_to_speech/LLM/responses_api_language_model.py for OpenAI API-compatible endpoints. The handler enriches input with chat history and returns streaming LMOutItem objects. Before passing to TTS, speech_to_speech/LLM/lm_output_processor.py post-processes the output to handle speculative turns and text compaction.
Text-to-Speech (TTS)
The TTS stage synthesizes the LLM response into audible speech. Supported models include Qwen-3, Pocket, Kokoro, ChatTTS, and Facebook MMS. The speech_to_speech/TTS/qwen3_tts_handler.py demonstrates the implementation pattern for local models, consuming TTSInItem objects and producing AudioOutItem speech data. Like other stages, TTS handlers are swappable via command-line arguments parsed in the main pipeline.
Pipeline Orchestration and Data Flow
The entire pipeline is assembled and executed by speech_to_speech/s2s_pipeline.py. This orchestration layer manages the complex data flow between stages using a queue-based architecture.
Queue-Driven Thread Management
The initialize_queues_and_events() function creates typed queue.Queue objects that link each stage. Every handler runs in its own thread via ThreadManager (speech_to_speech/utils/thread_manager.py), pulling from an input queue and pushing to an output queue. This design creates natural back-pressure: if a downstream consumer slows down, the upstream stage blocks on put() until space becomes available.
The data flow follows this path:
Audio chunks → VADHandler → VADOutItem → STTHandler → STTOutItem
→ TranscriptionNotifier → LLMHandler → LMOutItem
→ LMOutputProcessor → TTSInItem → TTSHandler → AudioOutItem
Handler Selection and Registration
The parse_arguments() function maps user-provided --stt, --llm_backend, and --tts options to concrete classes via get_stt_handler(), get_llm_handler(), and get_tts_handler(). Adding a new model requires only implementing the BaseHandler interface and registering it in s2s_pipeline.py.
Key Architectural Features
Asynchronous, Back-Pressure-Aware Design
Each stage operates independently with thread-safe queues mediating communication. The asynchronous design ensures that VAD can continue buffering new audio while the TTS stage synthesizes previous responses, maintaining pipeline throughput.
Speculative Turn Handling
The SpeculativeTurnTracker (located in pipeline/speculative_turns.py) enables the system to treat partial LLM responses as tentative "turns." This allows early audio playback while the LLM continues generating text, reducing perceived latency. The VAD handler can merge speculative audio prefixes with final output to eliminate gaps between turns.
Graceful Shutdown
A CancelScope and signal handlers in s2s_pipeline.py ensure all threads stop cleanly when the process receives SIGINT or SIGTERM signals, preventing audio device locks or zombie threads.
Implementation Examples
Running a Local Pipeline
To launch the complete pipeline with Whisper STT and Qwen-3 TTS:
from speech_to_speech.s2s_pipeline import (
parse_arguments,
prepare_all_args,
initialize_queues_and_events,
build_pipeline,
)
# Parse CLI flags or build ParsedArguments manually
args = parse_arguments()
# Prepare device-specific arguments
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.qwen3_tts_handler_kwargs,
# ... other handler kwargs
)
# Initialize queues and build pipeline
queues = initialize_queues_and_events()
pipeline_manager = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.qwen3_tts_handler_kwargs,
# ... other kwargs
queues,
)
# Start and block until shutdown
pipeline_manager.start()
pipeline_manager.wait()
WebSocket Realtime Server
To expose the pipeline via OpenAI-compatible WebSocket:
from speech_to_speech.s2s_pipeline import build_pipeline, parse_arguments, prepare_all_args, initialize_queues_and_events
args = parse_arguments()
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.qwen3_tts_handler_kwargs,
# ... other handlers
)
queues = initialize_queues_and_events()
pipeline_manager = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.qwen3_tts_handler_kwargs,
# ... other handlers
queues,
)
# Listens on ws://127.0.0.1:8765 by default
pipeline_manager.start()
pipeline_manager.wait()
Summary
- The VAD → STT → LLM → TTS pipeline processes audio through four distinct stages, each implemented as a threaded handler in the huggingface/speech-to-speech repository.
- Queue-based communication via
initialize_queues_and_events()enables asynchronous, back-pressure-aware data flow between stages. - Speculative turn tracking reduces latency by allowing early audio playback while the LLM completes generation.
- Extensible architecture supports multiple backends (Whisper, Faster-Whisper, Qwen-3, Kokoro) via command-line selection in
s2s_pipeline.py. - Graceful shutdown is handled through
CancelScopeand signal management to ensure clean thread termination.
Frequently Asked Questions
How does the VAD handler know when to start and stop recording?
The VAD handler uses the Silero VAD model through speech_to_speech/VAD/vad_iterator.py to detect voice activity thresholds. When audio levels exceed the threshold, it emits a SpeechStartedEvent and begins buffering chunks into a VADOutItem queue. When silence persists beyond a configured timeout, it emits a SpeechStoppedEvent and sends the complete buffered audio to the STT handler.
Can I swap individual components without modifying the pipeline code?
Yes. The parse_arguments() function in s2s_pipeline.py dynamically maps command-line flags such as --stt, --llm_backend, and --tts to handler classes via get_stt_handler(), get_llm_handler(), and get_tts_handler(). To add a new model, implement the BaseHandler interface and register the mapping in the argument parser.
What prevents the pipeline from overwhelming system resources?
The queue-based architecture implements natural back-pressure. Each stage blocks on queue.put() when the downstream queue is full, automatically throttling upstream processing. Additionally, the ThreadManager monitors thread health, and the CancelScope ensures immediate resource cleanup during shutdown.
How does realtime transcription work alongside the main pipeline?
When enable_realtime_transcription is set to True, the VAD handler in vad_handler.py emits interim text events through a separate text_output_queue while continuing to buffer audio for the full STT pass. This integrates with the OpenAI realtime WebSocket server in api/openai_realtime/server.py, allowing clients to receive partial transcriptions before the LLM response completes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →