Speech-to-Speech Model Architecture in the Hugging Face Speech-to-Speech Repository
The huggingface/speech-to-speech repository implements a modular, four-stage pipeline (VAD → STT → LLM → TTS) where each stage runs in its own thread and communicates via typed queues, enabling low-latency, pluggable voice assistants.
The speech-to-speech project provides a production-ready foundation for building real-time voice agents. Its architecture decouples voice activity detection, speech recognition, language generation, and speech synthesis into independent, swappable components. This design allows developers to mix open-source and proprietary models while maintaining thread-safe, low-latency audio streaming.
Four-Stage Pipeline Architecture
The core architecture follows a linear data flow with four sequential handlers. Each stage is implemented as a subclass of BaseHandler and operates in its own thread.
| Stage | Responsibility | Default Backend | Key File |
|---|---|---|---|
| Voice Activity Detection (VAD) | Detects speech boundaries, splits audio into turns, emits turn metadata | Silero VAD v5 | src/speech_to_speech/VAD/vad_handler.py |
| Speech-to-Text (STT) | Transcribes audio turns into text (with optional streaming) | Parakeet TDT | src/speech_to_speech/STT/parakeet_tdt_handler.py |
| Language Model (LLM) | Generates assistant responses, optionally with tool calls | OpenAI-compatible Responses API | src/speech_to_speech/LLM/responses_api_model.py |
| Text-to-Speech (TTS) | Synthesizes text back to audio and streams to client | Qwen3-TTS | src/speech_to_speech/TTS/qwen3_tts_handler.py |
Data Flow Diagram
┌─────────────┐ audio ┌───────┐ text ┌───────┐ text ┌───────┐
│ Audio │ ───────► │ VAD │ ─────► │ STT │ ─────► │ LLM │
│ Source │ │ │ │ │ │
│ (mic/socket│ │ │ │ │ │
│ /file) │ │ │ │ │ │
└─────────────┘ └───────┘ └───────┘ └───────┘
│
▼
audio (synth)
│
▼
┌───────┐
│ TTS │
└───────┘
│
output
Threading and Queue Communication
Inter-stage communication uses typed queues defined in src/speech_to_speech/pipeline/queue_types.py. Each handler reads from its input queue and writes to its output queue, enabling asynchronous, non-blocking data flow.
The build_pipeline function in src/speech_to_speech/s2s_pipeline.py (line 681) orchestrates this wiring:
- Creates queue instances for each stage transition
- Instantiates handler threads with appropriate kwargs
- Returns a
ThreadManagerthat starts and stops all threads
This design ensures that slow stages (like LLM generation) don't block audio capture or playback.
Realtime Mode and Pipeline Pooling
For production deployments, the repository provides realtime mode via RealtimeServer. This exposes an OpenAI Realtime-compatible WebSocket API at /v1/realtime.
Key implementation details from src/speech_to_speech/s2s_pipeline.py:
_build_realtime_pipeline_unit(line 480) creates isolated pipeline instances- Each WebSocket connection gets its own VAD→STT→LLM→TTS chain
- The server pools these units to handle concurrent clients
# Start the realtime server with default backend stack
speech-to-speech \
--mode realtime \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3 \
--enable_live_transcription
Turn-Taking and Speculative Interruptions
The VAD handler implements sophisticated turn management to minimize latency and handle interruptions gracefully.
Progressive Transcription
When enable_realtime_transcription is set, the VAD emits partial audio chunks before the turn completes. This allows the STT to begin transcription early, reducing perceived latency.
Speculative Turn Reopening
The speculative_turns feature (implemented in src/speech_to_speech/pipeline/speculative_turns.py) handles cases where users pause briefly then resume speaking:
SpeculativeTurnTrackermonitors turn state_should_reopen_current_turnevaluates whether to continue the current turn_reopen_current_turnand_begin_pending_reopen_if_neededmanage the transition
This avoids triggering a full LLM inference cycle for brief pauses, with configurable timeout via speculative_reopen_ms.
Backend Flexibility and Device Adaptation
Every stage supports multiple backend implementations. Selection happens at runtime through CLI flags parsed by ParsedArguments in s2s_pipeline.py.
STT Options
The get_stt_handler function (line 500) branches on the --stt flag:
parakeet-tdt— NVIDIA Parakeet TDT (default)whisper— OpenAI Whisperfaster-whisper— Optimized Whisper implementationlightning-whisper-mlx— Apple Silicon optimizedmlx-audio-whisper— MLX-native Whisperparaformer— Alibaba Paraformer
LLM Options
get_llm_handler (line 549) supports:
responses-api— OpenAI-compatible Responses API (default)transformers— Hugging Face Transformersmlx-lm— Apple Silicon optimized inferencechat-completions— OpenAI Chat Completions compatibility
TTS Options
get_tts_handler (line 689) includes:
qwen3— Qwen3-TTS (default)kokoro— Kokoro TTSpocket— Pocket TTS (lightweight)chat-tts— ChatTTSfacebook-mms— Meta MMS
macOS and MLX Integration
The prepare_all_args helper automatically selects MLX backends when --device mps is detected. The --local_mac_optimal_settings flag applies tuned defaults for Apple Silicon.
Building a Custom Pipeline
The Python API provides fine-grained control over speech-to-speech model architecture composition:
from speech_to_speech.s2s_pipeline import (
parse_arguments,
prepare_all_args,
initialize_queues_and_events,
build_pipeline,
)
# Parse arguments programmatically
args = parse_arguments()
# Normalize device settings and backend kwargs
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.qwen3_tts_handler_kwargs,
# ... additional handler kwargs
)
# Initialize communication queues
queues = initialize_queues_and_events()
# Construct the pipeline
pipeline_manager = build_pipeline(
args.module_kwargs,
args.vad_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.qwen3_tts_handler_kwargs,
queues,
)
pipeline_manager.start()
pipeline_manager.wait()
This mirrors the main() entry point but allows dynamic configuration without CLI invocation.
Swapping Backends: Practical Example
Replace the default STT with Faster-Whisper and use a local Transformers LLM:
speech-to-speech \
--mode realtime \
--stt faster-whisper \
--faster_whisper_stt_model_name large-v2 \
--llm_backend transformers \
--model_name meta-llama/Meta-Llama-3.1-8B-Instruct \
--tts qwen3 \
--enable_live_transcription
The --stt faster-whisper branch in get_stt_handler instantiates FasterWhisperSTTHandler from src/speech_to_speech/STT/faster_whisper_handler.py.
Extending the Architecture
Adding new backends requires minimal changes:
- Implement a subclass of
BaseHandlerin the appropriateSTT/,LLM/, orTTS/directory - Register it in the corresponding
get_*_handlerfunction - Add argument dataclass in
src/speech_to_speech/arguments_classes/ - Expose CLI flags through
ParsedArguments
The arguments_classes/ pattern keeps configuration modular—each handler owns its parameter definitions without modifying core pipeline code.
Summary
- The speech-to-speech model architecture uses four pluggable stages: VAD → STT → LLM → TTS
- Each stage runs in its own thread with typed queue communication via
queue_types.py - Realtime mode provides OpenAI-compatible WebSocket API with per-client pipeline isolation
- Speculative turn reopening enables low-latency interruption handling without full pipeline resets
- Multiple backends per stage (6 STT, 4 LLM, 5 TTS) support diverse hardware and latency requirements
- Device adaptation automatically selects MLX backends on Apple Silicon
Frequently Asked Questions
What is the default speech-to-speech pipeline configuration?
The default stack uses Silero VAD v5 for voice detection, Parakeet TDT for speech-to-text, OpenAI Responses API for language modeling, and Qwen3-TTS for text-to-speech. This configuration prioritizes quality and low latency on GPU-equipped servers. Run with speech-to-speech --mode realtime to activate.
How does the pipeline handle user interruptions?
The speculative turn mechanism in VADHandler tracks provisional turn boundaries. If speech resumes within speculative_reopen_ms milliseconds, the turn continues without triggering a new LLM inference. This logic lives in _should_reopen_current_turn and uses SpeculativeTurnTracker from speculative_turns.py to manage state.
Can I run the pipeline on Apple Silicon Macs?
Yes. Set --device mps and optionally --local_mac_optimal_settings. The prepare_all_args helper automatically routes to MLX-compatible backends: mlx-lm for LLM inference, lightning-whisper-mlx or mlx-audio-whisper for STT, and Qwen3-TTS supports MPS directly.
How do I add a custom TTS model to the pipeline?
Implement a BaseHandler subclass in src/speech_to_speech/TTS/, add a factory branch in get_tts_handler (line 689), create an arguments dataclass in arguments_classes/, and register the CLI flag. The handler receives text via its input queue and should emit audio chunks to its output queue in the expected format.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →