Architecture of the Default Speech-to-Speech Models in Hugging Face's Pipeline
The default architecture chains four specialized handlers—VADHandler, ParakeetTDTSTTHandler, ResponsesApiModelHandler, and Qwen3TTSHandler—into a real-time pipeline that converts speech to text, processes it through a remote LLM, and synthesizes the response back into speech.
The huggingface/speech-to-speech repository implements a modular, low-latency inference system designed for conversational AI applications. When invoked without custom overrides, the pipeline automatically instantiates a specific quartet of models optimized for speed and quality. Understanding this default architecture helps developers debug latency issues and swap components effectively.
The Four-Stage Default Pipeline
The architecture follows a strict linear data flow defined in src/speech_to_speech/s2s_pipeline.py. Each stage communicates via strongly-typed queues, ensuring type safety across asynchronous boundaries.
Voice Activity Detection (VAD)
The pipeline begins with VADHandler, located in src/speech_to_speech/VAD/vad_handler.py. This component captures raw audio chunks and determines when speech starts and stops.
In _build_pipeline_handlers (lines 16–22 of s2s_pipeline.py), the VAD handler is instantiated first. It performs two critical functions: it feeds active audio segments downstream to the STT handler and simultaneously forwards spoken prompts to the TranscriptionNotifier for UI display.
Speech-to-Text (STT) - Parakeet-TDT
The default speech recognizer is ParakeetTDTSTTHandler, implemented in src/speech_to_speech/STT/parakeet_tdt_handler.py. This handler uses the parakeet-tdt model, selected specifically for its fast, low-latency transcription capabilities.
The handler is created by get_stt_handler (lines 81–88) according to the default arguments defined in src/speech_to_speech/arguments_classes/module_arguments.py. Recognized text is packaged into STTOutItem objects and queued for the language model stage.
Language Model (LLM) - Responses API
The default LLM backend is ResponsesApiModelHandler, found in src/speech_to_speech/LLM/responses_api_language_model.py. Rather than running a local model, this handler connects to a remote OpenAI-compatible API endpoint (configured via the responses-api backend).
Built by get_llm_handler (lines 88–109), this component consumes text prompts from the STT stage, generates responses, and pushes them downstream as LMOutItem objects. This remote architecture keeps local resource usage low while providing access to large-scale models.
Text-to-Speech (TTS) - Qwen3
The final stage uses Qwen3TTSHandler from src/speech_to_speech/TTS/qwen3_tts_handler.py, leveraging the qwen3 model accelerated via MLX. Before reaching TTS, raw LLM output passes through LMOutputProcessor (instantiated in _build_pipeline_handlers, lines 53–57), which handles speculative-turn detection and optional text events.
The TTS handler is constructed by get_tts_handler (lines 98–115) and outputs synthesized audio that is then transmitted to the client via the selected communication mode.
Pipeline Orchestration and Communication
Queue Types and Typed Communication
All inter-stage communication uses strongly-typed queue items defined in src/speech_to_speech/pipeline/queue_types.py. The four primary data structures are:
AudioInItem– Raw audio from VAD to STTSTTOutItem– Transcribed text from STT to LLMLMOutItem– Generated responses from LLM to TTSTTSInItem– Processed text ready for speech synthesis
This typing ensures that handlers receive expected data formats without runtime ambiguity.
Event Synchronization
Pipeline control relies on synchronization primitives defined in src/speech_to_speech/pipeline/control.py. Key events include stop_event for graceful shutdown and should_listen for managing the listening state. These events coordinate the four handlers across different execution threads or processes.
In realtime mode (the default), the system creates isolated pipeline pools via _build_realtime_pipeline_unit, where each session maintains its own VAD, STT, LLM, and TTS instances to prevent cross-contamination between simultaneous users.
Running the Default Pipeline
To launch the default architecture with all standard components, use the following command:
python -m speech_to_speech.s2s_pipeline \
--mode realtime \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3
For programmatic access, instantiate the pipeline through the Python API:
from speech_to_speech.s2s_pipeline import (
parse_arguments,
build_pipeline,
initialize_queues_and_events,
prepare_all_args
)
# Load default configurations
args = parse_arguments()
# Initialize device-specific and model-specific arguments
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
)
# Create queues and events
queues = initialize_queues_and_events()
# Build and start the pipeline
pipeline_manager = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
queues,
)
pipeline_manager.start()
pipeline_manager.wait()
Summary
- The default architecture implements a VAD → Parakeet-TDT STT → Responses-API LLM → Qwen3 TTS pipeline.
- Component selection is governed by
ModuleArgumentsinsrc/speech_to_speech/arguments_classes/module_arguments.py. - Inter-process communication uses typed queues (
AudioInItem,STTOutItem,LMOutItem,TTSInItem) fromqueue_types.py. - Event synchronization relies on
control.pyfor managing pipeline state across concurrent sessions. - The system supports four execution modes (
local,socket,raw-websocket,realtime), with realtime mode creating isolated pipeline pools for multi-user scenarios.
Frequently Asked Questions
Can I replace the default Parakeet-TDT recognizer with Whisper?
Yes. The pipeline supports multiple STT handlers including Whisper variants. Override the --stt argument with whisper, faster-whisper, or mlx-audio-whisper to switch recognizers. The get_stt_handler factory function in s2s_pipeline.py automatically instantiates the appropriate handler class based on this argument.
Why does the default architecture use a remote LLM instead of a local model?
The default ResponsesApiModelHandler minimizes local GPU/CPU memory requirements by offloading inference to remote API endpoints. This allows the pipeline to run on edge devices while still leveraging large language models. You can switch to local LLM inference by changing --llm_backend to a local-compatible handler, though this requires additional configuration in the language model handler arguments.
How does the pipeline handle multiple simultaneous conversations?
In the default realtime mode, the _build_realtime_pipeline_unit function creates isolated pipeline instances for each connection. Each instance maintains its own VADHandler, ParakeetTDTSTTHandler, ResponsesApiModelHandler, and Qwen3TTSHandler with separate queue sets, ensuring that audio and text from different users never intermingle.
What is the purpose of the LMOutputProcessor between the LLM and TTS?
The LMOutputProcessor (instantiated in _build_pipeline_handlers) performs intermediate text processing before synthesis. It handles speculative-turn detection—identifying when the model has generated a complete thought versus a partial sentence—and manages optional text events. This preprocessing ensures that the TTS handler receives properly segmented text for natural-sounding speech synthesis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →