Speech-to-Speech Model Architecture: Inside the Hugging Face Real-Time Voice Pipeline
The huggingface/speech-to-speech repository implements a modular, four-stage pipeline architecture consisting of Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS) handlers that communicate via thread-safe queues to enable low-latency, real-time voice conversation.
The open-source speech-to-speech model architecture provides a flexible framework for building real-time voice assistants on consumer hardware. This system chains together specialized handlers in a producer-consumer pattern, allowing developers to swap individual components—such as replacing Whisper with Parakeet TDT or switching between Qwen3-TTS and Kokoro—without modifying the core orchestration logic in src/speech_to_speech/s2s_pipeline.py.
Four-Stage Modular Architecture
The architecture processes audio through four sequential stages, each running as an independent thread with typed input and output queues defined in src/speech_to_speech/pipeline/queue_types.py.
Voice Activity Detection (VAD)
The VAD stage detects when users start and stop speaking, splitting continuous audio into discrete turns. By default, the pipeline uses Silero VAD v5 via the VADHandler class in src/speech_to_speech/VAD/vad_handler.py. This handler manages turn boundaries and emits metadata events that downstream stages consume.
When enable_realtime_transcription is active, the VAD emits progressive audio chunks for live transcription rather than waiting for complete turns. The handler also supports speculative turn tracking—if the user resumes speaking within a configurable timeout (speculative_reopen_ms), the VAD reopens the current turn using methods like _should_reopen_current_turn and _reopen_current_turn, avoiding the latency of starting a new LLM inference.
Speech-to-Text (STT)
The STT stage transcribes audio segments into text. The default backend is Parakeet TDT, though the architecture supports multiple implementations selected via the get_stt_handler function in s2s_pipeline.py (around line 500).
Available backends include:
- Parakeet TDT – Optimized for real-time streaming
- Whisper – OpenAI's original implementation
- Faster-Whisper – Optimized Whisper variant
- Lightning-Whisper-MLX – Apple Silicon optimized
- Paraformer – Alternative architecture
Each implementation resides in src/speech_to_speech/STT/ and inherits from BaseHandler, ensuring consistent queue-based interfaces.
Language Model (LLM)
The LLM stage generates text responses from the transcribed user input. By default, the pipeline uses the OpenAI-compatible Responses API, though it supports local inference via Transformers, mlx-lm (for Apple Silicon), and Chat-Completions compatible endpoints.
The get_llm_handler function (around line 549 in s2s_pipeline.py) instantiates the appropriate handler from src/speech_to_speech/LLM/, which includes LanguageModelHandler, ResponsesApiModelHandler, and VisionLanguageModelHandler. This stage can stream partial responses to the TTS stage before generation completes, reducing perceived latency.
Text-to-Speech (TTS)
The final TTS stage synthesizes the LLM output into audio streams sent to the client. The default backend is Qwen3-TTS, with alternatives including Kokoro, Pocket, ChatTTS, and Facebook-MMS.
The get_tts_handler function (around line 689 in s2s_pipeline.py) selects the appropriate implementation from src/speech_to_speech/TTS/. Each handler converts text chunks into audio frames, enabling true streaming synthesis where the user hears audio before the LLM finishes generating the complete response.
Threading and Data Flow
The pipeline orchestration relies on a thread-per-stage architecture coordinated by the build_pipeline function in s2s_pipeline.py (around line 681). This function:
- Creates typed queues for inter-stage communication
- Instantiates each handler (VAD, STT, LLM, TTS)
- Returns a
ThreadManagerthat starts and supervises all threads
Audio Source → [VAD Thread] → [STT Thread] → [LLM Thread] → [TTS Thread] → Output
Each handler runs in a loop, blocking on its input queue, processing data, and placing results on the next stage's queue. This design ensures that slow operations (like LLM generation) do not block audio capture or VAD processing.
Realtime Mode and Pipeline Pools
For production deployments, the RealtimeServer creates isolated pipeline instances for each concurrent client. The _build_realtime_pipeline_unit function (around line 480 in s2s_pipeline.py) constructs these pipeline units, which expose the OpenAI Realtime WebSocket API at /v1/realtime.
Each client receives a dedicated set of threads and queues, ensuring that heavy processing by one user does not affect others' latency.
Speculative Turn-Taking for Interruption Handling
Low-latency interruption handling relies on the SpeculativeTurnTracker in src/speech_to_speech/pipeline/speculative_turns.py. When speculative_turns is enabled:
- The VAD tracks speculative turn IDs alongside confirmed turns
- If speech resumes within
speculative_reopen_ms, the system reopens the existing turn rather than creating a new one - This prevents the LLM from starting fresh inferences when users briefly pause mid-thought
The logic resides in VADHandler methods like _begin_pending_reopen_if_needed, which coordinate with the speculative tracker to minimize response latency during natural conversation pauses.
Implementation Examples
Running the Default Realtime Server
Configure the complete four-stage pipeline with default backends using the CLI:
pip install speech-to-speech
speech-to-speech \
--mode realtime \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3 \
--enable_live_transcription
This command maps to the parse_arguments function in s2s_pipeline.py, which validates the configuration and prepares the ParsedArguments dataclass.
Building a Local Pipeline via Python API
For custom integrations, instantiate the pipeline programmatically:
from speech_to_speech.s2s_pipeline import (
parse_arguments,
prepare_all_args,
initialize_queues_and_events,
build_pipeline,
)
# Parse configuration
args = parse_arguments()
# Normalize device arguments and backend-specific parameters
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.qwen3_tts_handler_kwargs,
# ... other handler kwargs
)
# Initialize communication queues
queues = initialize_queues_and_events()
# Construct thread manager
pipeline_manager = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.qwen3_tts_handler_kwargs,
queues,
)
# Start all handler threads
pipeline_manager.start()
pipeline_manager.wait()
This mirrors the main() entry point in s2s_pipeline.py, giving you direct control over queue initialization and thread lifecycle.
Swapping Backends (Faster-Whisper Example)
Replace the default STT backend via command-line flags:
speech-to-speech \
--mode realtime \
--stt faster-whisper \
--faster_whisper_stt_model_name large-v2 \
--llm_backend transformers \
--model_name meta-llama/Meta-Llama-3.1-8B-Instruct \
--tts qwen3
The --stt faster-whisper flag triggers the elif module_kwargs.stt == "faster-whisper" branch in get_stt_handler, loading the appropriate handler from src/speech_to_speech/STT/faster_whisper_handler.py.
Summary
- The architecture comprises four sequential stages (VAD → STT → LLM → TTS) communicating through thread-safe queues defined in
src/speech_to_speech/pipeline/queue_types.py. - Each stage runs in its own thread, managed by
build_pipelineand supervised byThreadManagerto prevent blocking. - Backends are fully interchangeable via
get_stt_handler,get_llm_handler, andget_tts_handlerfunctions, supporting both cloud APIs and local MLX-optimized models. - Speculative turn tracking in
VADHandlerandSpeculativeTurnTrackerenables low-latency interruption handling without restarting LLM inference. - Realtime mode creates isolated pipeline pools per client, exposing an OpenAI-compatible WebSocket API via
RealtimeServer.
Frequently Asked Questions
What is the default speech-to-speech model architecture in the Hugging Face repository?
The default architecture chains Silero VAD v5 for voice detection, Parakeet TDT for speech recognition, the OpenAI Responses API for language modeling, and Qwen3-TTS for speech synthesis. These components communicate through Python's queue.Queue structures managed by the build_pipeline function in src/speech_to_speech/s2s_pipeline.py.
How does the pipeline handle interruptions during real-time conversation?
The system uses speculative turn tracking via the SpeculativeTurnTracker class in src/speech_to_speech/pipeline/speculative_turns.py. When enabled, the VAD handler tracks when turns might reopen using _should_reopen_current_turn. If the user resumes speaking within the speculative_reopen_ms window, the existing turn continues rather than starting a new LLM inference, maintaining conversational continuity with minimal latency.
Can I use local models instead of cloud APIs in this architecture?
Yes. The architecture supports local inference through multiple backends. For LLMs, use --llm_backend transformers or --llm_backend mlx-lm (for Apple Silicon). For STT, select --stt faster-whisper or --stt mlx-audio-whisper. For TTS, options include --tts kokoro and --tts qwen3 with --device mps. The prepare_all_args function automatically applies macOS-specific optimizations when --local_mac_optimal_settings is enabled.
What files control the queue communication between pipeline stages?
Queue definitions reside in src/speech_to_speech/pipeline/queue_types.py, which provides typed queue.Queue instances for audio, text, and control messages. The initialize_queues_and_events function creates these queues, which are then passed to build_pipeline to wire together the VAD, STT, LLM, and TTS handlers. Each handler runs in its own thread, blocking on its input queue and placing results on the output queue for the next stage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →