How to Integrate Speech-to-Speech into a Python Application: A Complete Guide
You can integrate speech-to-speech into a Python application by using the huggingface/speech-to-speech pipeline's main() entry point for CLI usage, or by programmatically constructing the four-stage pipeline (VAD → STT → LLM → TTS) using the build_pipeline() function with argument dataclasses.
The huggingface/speech-to-speech repository provides a modular, production-ready framework for building real-time voice conversations. It chains together Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) into a seamless pipeline that runs across multiple threads. This guide shows you how to integrate this system into your Python applications using both command-line and programmatic approaches.
Understand the Four-Stage Pipeline Architecture
The pipeline processes audio through four distinct stages connected by queues and events:
-
VAD (Voice Activity Detection) - Detects when users start/stop speaking. Implemented in
speech_to_speech/VAD/vad_handler.pyand instantiated ins2s_pipeline._build_pipeline_handlers. -
STT (Speech-to-Text) - Converts audio to text using Whisper, Faster-Whisper, Paraformer, or MLX-Audio-Whisper. Selected via
module_kwargs.sttand created ins2s_pipeline.get_stt_handler. -
LLM (Language Model) - Generates responses using Transformers, MLX-LM, or OpenAI-compatible APIs. Built in
s2s_pipeline.get_llm_handler. -
TTS (Text-to-Speech) - Synthesizes speech via ChatTTS, Facebook MMS, Pocket, Kokoro, or Qwen-3. Created in
s2s_pipeline.get_tts_handler.
All stages communicate via thread-safe queues initialized through initialize_queues_and_events() and managed by the ThreadManager in src/speech_to_speech/utils/thread_manager.py.
Method 1: Quick Integration via Command Line
For rapid prototyping, use the built-in CLI entry point.
This approach handles argument parsing via HfArgumentParser, normalization through prepare_all_args, and pipeline construction automatically.
from speech_to_speech.s2s_pipeline import main
if __name__ == "__main__":
main()
Run with command-line flags:
python -m speech_to_speech.s2s_pipeline \
--mode local \
--stt whisper \
--tts qwen3 \
--llm_backend transformers \
--device cpu \
--log_level info
The main() function in s2s_pipeline.py orchestrates the entire flow: parsing arguments, building handlers, and starting the ThreadManager.
Method 2: Programmatic Integration
For embedding within existing applications, construct the pipeline manually using dataclasses from src/speech_to_speech/arguments_classes/.
Step 1: Configure Arguments
Import and instantiate the specific argument classes for your chosen backends:
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelHandlerArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSHandlerArguments
module_args = ModuleArguments(
mode="local", # Options: "local", "websocket", "realtime"
stt="whisper",
tts="qwen3",
llm_backend="transformers",
device="cpu",
log_level="info",
)
whisper_args = WhisperSTTHandlerArguments(model_name="openai/whisper-base")
lm_args = LanguageModelHandlerArguments(model_name="Qwen/Qwen3-4B-Instruct-2507")
tts_args = Qwen3TTSHandlerArguments()
Step 2: Normalize Arguments
Call prepare_all_args() to apply device optimizations and format gen_kwargs:
from speech_to_speech.s2s_pipeline import prepare_all_args
prepare_all_args(
module_args,
whisper_args,
WhisperSTTHandlerArguments(), # Placeholder for paraformer
WhisperSTTHandlerArguments(), # Placeholder for faster-whisper
WhisperSTTHandlerArguments(), # Placeholder for mlx-audio-whisper
WhisperSTTHandlerArguments(), # Placeholder for parakeet-tdt
lm_args,
LanguageModelHandlerArguments(), # Placeholder for responses-api
tts_args,
Qwen3TTSHandlerArguments(), # Placeholder for other TTS handlers
Qwen3TTSHandlerArguments(),
Qwen3TTSHandlerArguments(),
Qwen3TTSHandlerArguments(),
)
Step 3: Initialize Communication Primitives
Create the queues and events that connect pipeline stages:
from speech_to_speech.s2s_pipeline import initialize_queues_and_events
queues = initialize_queues_and_events()
Step 4: Build and Run the Pipeline
Construct the pipeline with build_pipeline() and start execution:
from speech_to_speech.s2s_pipeline import build_pipeline
from speech_to_speech.arguments_classes.socket_receiver_arguments import SocketReceiverArguments
from speech_to_speech.arguments_classes.socket_sender_arguments import SocketSenderArguments
from speech_to_speech.arguments_classes.websocket_streamer_arguments import WebSocketStreamerArguments
from speech_to_speech.arguments_classes.vad_handler_arguments import VADHandlerArguments
from speech_to_speech.arguments_classes.faster_whisper_stt_arguments import FasterWhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.paraformer_stt_arguments import ParaformerSTTHandlerArguments
from speech_to_speech.arguments_classes.mlx_audio_whisper_stt_arguments import MLXAudioWhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.parakeet_tdt_stt_arguments import ParakeetTDTSTTHandlerArguments
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import ResponsesApiLanguageModelHandlerArguments
from speech_to_speech.arguments_classes.chat_tts_arguments import ChatTTSHandlerArguments
from speech_to_speech.arguments_classes.facebook_mms_tts_arguments import FacebookMMSTTSHandlerArguments
from speech_to_speech.arguments_classes.pocket_tts_arguments import PocketTTSHandlerArguments
from speech_to_speech.arguments_classes.kokoro_tts_arguments import KokoroTTSHandlerArguments
pipeline_manager = build_pipeline(
module_args,
SocketReceiverArguments(),
SocketSenderArguments(),
WebSocketStreamerArguments(),
VADHandlerArguments(),
whisper_args,
FasterWhisperSTTHandlerArguments(),
ParaformerSTTHandlerArguments(),
MLXAudioWhisperSTTHandlerArguments(),
ParakeetTDTSTTHandlerArguments(),
lm_args,
ResponsesApiLanguageModelHandlerArguments(),
ChatTTSHandlerArguments(),
FacebookMMSTTSHandlerArguments(),
PocketTTSHandlerArguments(),
KokoroTTSHandlerArguments(),
tts_args,
queues,
)
pipeline_manager.start()
pipeline_manager.wait() # Blocks until interrupted
Configure Audio I/O Modes
The pipeline supports three operational modes controlled by module_args.mode:
- Local mode (
"local"): Direct microphone input and speaker output. Best for standalone applications. - WebSocket mode (
"websocket"): Audio streams over WebSocket connections. Ideal for web applications. - Real-time mode (
"realtime"): OpenAI-compatible realtime API integration.
For low-latency requirements, enable live transcription with module_args.enable_live_transcription=True.
Test with the Socket Demo
The repository includes scripts/listen_and_play.py for testing TCP socket streaming without hardware loopback.
Start the pipeline in WebSocket mode:
python -m speech_to_speech.s2s_pipeline --mode websocket
Then run the demo client:
python scripts/listen_and_play.py \
--host localhost \
--send_port 12345 \
--recv_port 12346
The demo captures raw 16-bit audio, pushes it to the pipeline via TCP, and plays back synthesized responses.
Summary
- Four-stage architecture: The pipeline chains VAD → STT → LLM → TTS through thread-safe queues managed by
ThreadManager. - Two integration paths: Use
main()for CLI-driven deployment orbuild_pipeline()for embedded applications. - Flexible backends: Choose from Whisper, Faster-Whisper, or Paraformer for STT; Qwen-3, ChatTTS, or Kokoro for TTS; and Transformers or MLX-LM for language modeling.
- Multiple modes: Run locally with direct audio hardware, over WebSockets for web apps, or via the OpenAI-compatible realtime API.
- Configuration via dataclasses: All parameters are type-safe arguments defined in
src/speech_to_speech/arguments_classes/.
Frequently Asked Questions
How do I select different STT or TTS backends?
Set the stt and tts parameters in ModuleArguments to your desired backend identifiers (e.g., "whisper", "faster-whisper", "qwen3", "kokoro"), then provide the corresponding handler arguments to prepare_all_args(). The build_pipeline() function instantiates the correct handler classes based on these selections.
Can I run the pipeline on macOS with Apple Silicon?
Yes. The prepare_all_args() function automatically detects macOS and applies optimizations, preferring mlx-lm for the LLM backend and qwen3 for TTS when available. Specify device="mps" or leave it on auto-detect for best performance on Apple Silicon.
What is the difference between local mode and WebSocket mode?
Local mode (mode="local") uses direct system audio through local audio streamers, suitable for desktop applications. WebSocket mode (mode="websocket") accepts audio over WebSocket connections via WebSocketStreamerArguments, enabling remote clients to stream audio to the pipeline over the network.
How do I enable graceful shutdown in my application?
The ThreadManager returned by build_pipeline() registers signal handlers for graceful shutdown. Call pipeline_manager.stop() or send a KeyboardInterrupt (Ctrl+C) to trigger the shutdown sequence, which properly joins all handler threads and clears the inter-thread queues.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →