How Does the Speech-to-Speech Model Generate Output? A Deep Dive into the Pipeline
The speech-to-speech model generates output through a modular, asynchronous pipeline that converts audio input into spoken responses via six sequential stages: Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM) inference, output processing, Text-to-Speech (TTS), and audio streaming.
The huggingface/speech-to-speech repository implements this architecture as a series of interconnected handlers that communicate through thread-safe queues. Understanding how the speech-to-speech model generates output requires examining the data flow from raw audio capture to final PCM audio playback.
The Six-Stage Generation Pipeline
Stage 1: Voice Activity Detection (VAD)
The pipeline begins with the VADHandler in speech_to_speech/VAD/vad_handler.py, which processes incoming audio streams from microphones, sockets, or local files. This handler trims silence and emits spoken prompts—chunks of voice activity—into the pipeline for downstream processing.
Stage 2: Speech-to-Text Transcription (STT)
Each spoken prompt flows to a selected STT backend following the BaseSTTHandler contract. The concrete implementations include Whisper, Faster-Whisper, Paraformer, and others located in speech_to_speech/STT/whisper_stt_handler.py. These handlers return LLMResponseChunk objects containing transcribed text and optional tool-call metadata.
Stage 3: Language Model Processing
The transcribed text passes to one of three language model handlers: LanguageModelHandler, ResponsesApiModelHandler, or ChatCompletionsApiModelHandler in speech_to_speech/LLM/language_model.py. The LLM streams back LLMResponseChunk objects, TokenUsage reports, and a final EndOfResponse signal to indicate completion.
Stage 4: LM Output Processing
The LMOutputProcessor in speech_to_speech/LLM/lm_output_processor.py intercepts the LLM stream and performs two critical functions:
- Pushes
AssistantTextEventobjects (text + optional tool calls) onto a text-output queue for WebSocket clients - Yields clean
TTSInputobjects downstream only whenresponse_wants_audiois true
The processor uses SpeculativeTurnTracker to drop stale turns and prevent "ghost" generations from outdated responses.
Stage 5: Text-to-Speech Synthesis
The TTSInput flows to concrete TTS handlers such as Qwen3TTSHandler, ChatTTSHandler, or PocketTTSHandler. The implementation in speech_to_speech/TTS/qwen3_tts_handler.py streams int16 PCM audio chunks while handling voice-cloning, custom-speaker, or voice-design modes. These handlers respect session-level voice overrides from OpenAI Realtime responses or runtime configuration.
Stage 6: Audio Output Streaming
Generated audio chunks enter the send-audio queue, consumed by either LocalAudioStreamer or SocketSender. The speech_to_speech/connections/local_audio_streamer.py implementation plays audio on local sound devices or transmits it over the network, completing the generation cycle.
Pipeline Orchestration and Queue Management
Queue Types and Thread Safety
All components communicate via thread-safe Queue objects defined in speech_to_speech/pipeline/queue_types.py. The initialize_queues_and_events helper creates queues for:
- Stop signals
- VAD output
- STT output
- Text prompts
- LLM responses
- TTS input
- Final audio
Speculative Turn Tracking
For real-time usage, the SpeculativeTurnTracker in speech_to_speech/pipeline/speculative_turns.py tracks turn IDs to silently discard late-arriving STT or LLM chunks belonging to older turns. This prevents the system from speaking outdated content when users interrupt or when network latency causes out-of-order processing.
Runtime Configuration
In OpenAI Realtime mode, the system accepts a RealtimeSessionCreateRequest supplied to LLM and TTS handlers via speech_to_speech/api/openai_realtime/runtime_config.py. This enables session-level voice overrides and streaming controls without restarting the pipeline.
Running the Pipeline
Command-Line Execution
Execute the full pipeline using the s2s_pipeline.py module:
python -m speech_to_speech.s2s_pipeline \
--mode realtime \
--stt whisper \
--llm_backend responses-api \
--tts qwen3 \
--ws_host localhost \
--ws_port 8080
Programmatic Python API
Build and start the pipeline programmatically:
from speech_to_speech.s2s_pipeline import (
parse_arguments,
prepare_all_args,
initialize_queues_and_events,
build_pipeline,
)
# Parse CLI-style arguments or construct manually
args = parse_arguments()
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.qwen3_tts_handler_kwargs,
# ... other handler kwargs
)
# Build queues and pipeline
queues = initialize_queues_and_events()
pipeline = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.qwen3_tts_handler_kwargs,
queues,
)
pipeline.start() # Launches all background threads
pipeline.wait() # Blocks until shutdown
Local Audio Testing
For local testing without network configuration:
python scripts/listen_and_play.py \
--send_rate 16000 \
--recv_rate 16000 \
--host localhost \
--send_port 12345 \
--recv_port 12346
This opens socket pairs, streams microphone audio to the pipeline, and plays generated speech in real-time.
Summary
- The speech-to-speech model generates output through a six-stage asynchronous pipeline (VAD → STT → LLM → Processor → TTS → Audio).
- Queue-based architecture in
queue_types.pyensures thread-safe communication between handlers without blocking. - Speculative turn tracking prevents outdated responses from being spoken when users interrupt or network latency occurs.
- Modular handler design allows swapping STT backends (Whisper, Paraformer), LLM providers, and TTS engines (Qwen3, ChatTTS) without changing core logic.
- OpenAI Realtime compatibility enables session-level voice configuration and streaming controls via
runtime_config.py.
Frequently Asked Questions
What prevents the speech-to-speech model from playing outdated responses?
The SpeculativeTurnTracker in speech_to_speech/pipeline/speculative_turns.py tracks unique turn IDs across the pipeline. When new user input arrives, older turns are marked stale, and the LMOutputProcessor silently drops any late-arriving chunks from previous turns before they reach the TTS stage.
How does the pipeline support different TTS backends?
The architecture uses a base handler pattern where concrete implementations like Qwen3TTSHandler and ChatTTSHandler inherit from the base class defined in speech_to_speech/baseHandler.py. Each handler receives TTSInput objects and streams int16 PCM audio, allowing the build_pipeline function in s2s_pipeline.py to inject the appropriate backend based on runtime arguments.
Can the speech-to-speech model operate without speaking responses?
Yes. The LMOutputProcessor checks the response_wants_audio flag before yielding TTSInput downstream. If disabled, the system still pushes AssistantTextEvent objects to the text-output queue for WebSocket clients, enabling text-only responses while maintaining the full LLM inference capability.
What audio format does the LocalAudioStreamer expect?
According to speech_to_speech/connections/local_audio_streamer.py, the system expects int16 PCM audio chunks at the sample rate specified during initialization (typically 16kHz or 24kHz). The TTS handlers convert their native outputs to this format before placing chunks on the send-audio queue.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →