Python APIs for Speech-to-Speech: Building Low-Latency Pipelines with Hugging Face
The Hugging Face speech-to-speech repository provides a comprehensive Python API for constructing modular, low-latency speech-to-speech pipelines via the s2s_pipeline module, which exposes three core functions—parse_arguments(), initialize_queues_and_events(), and build_pipeline()—to orchestrate VAD → STT → LLM → TTS workflows entirely in Python.
The speech-to-speech repository from Hugging Face offers a fully-featured Python API that enables developers to build real-time speech-to-speech systems without leaving the Python ecosystem. Unlike command-line-only tools, this API allows you to programmatically configure, initialize, and run complete audio processing pipelines using thread-safe handlers and queue-based communication. As implemented in the repository, the API supports both local audio I/O and network-based streaming through a unified interface defined in src/speech_to_speech/s2s_pipeline.py.
Core Python API Components
The primary entry point for the Python API is the speech_to_speech.s2s_pipeline module. According to the source code in src/speech_to_speech/s2s_pipeline.py, this module exposes three high-level functions that handle the complete lifecycle of a speech-to-speech pipeline:
parse_arguments()(line 29): Parses command-line arguments or JSON configuration files into a typedParsedArgumentsdataclass.initialize_queues_and_events()(line 33): Creates thread-safequeue.Queuesubclasses and synchronization events that connect each processing stage.build_pipeline()(line 81): Constructs the full pipeline chain (VAD → STT → LLM → TTS) and returns aThreadManagerinstance for execution control.
These functions work sequentially to transform configuration parameters into a running pipeline managed by the ThreadManager class from src/speech_to_speech/utils/thread_manager.py.
Configuration and Argument Parsing
In src/speech_to_speech/s2s_pipeline.py, the parse_arguments() function serves as the configuration gateway. It accepts command-line arguments or JSON configuration files and returns a ParsedArguments dataclass containing all necessary parameters for the pipeline stages. This includes model selections (e.g., --stt whisper, --tts qwen3), device configurations (CPU, CUDA, or Apple Silicon), and backend-specific options for each handler.
Queue Initialization and Thread Safety
The initialize_queues_and_events() function (line 33) establishes the inter-process communication infrastructure. It creates the thread-safe queues and threading.Event objects that allow each stage to communicate asynchronously. This architecture ensures that voice activity detection, speech recognition, language model inference, and speech synthesis can run concurrently without blocking, with each handler waiting on its respective input queue and signaling completion via shared events.
Pipeline Construction
The build_pipeline() function (line 81) acts as the factory that instantiates and connects all pipeline components. It accepts keyword arguments for every supported handler—including vad_handler, whisper_stt_handler, language_model_handler, and qwen3_tts_handler—and wires them together according to the architecture: VADHandler → STTHandler → TranscriptionNotifier → LLMHandler → LMOutputProcessor → TTSHandler. The function returns a ThreadManager that provides start(), stop(), and wait() methods for lifecycle management.
Modular Pipeline Architecture
The Python API implements a handler-based architecture where each processing stage is encapsulated in a specialized handler class. All handlers inherit from the abstract base class defined in src/speech_to_speech/baseHandler.py and communicate via the queues initialized earlier. You can swap implementations by changing the arguments passed to build_pipeline().
Voice Activity Detection (VAD)
The pipeline begins with voice activity detection handled by the VADHandler class in src/speech_to_speech/VAD/vad_handler.py. This component monitors audio input and triggers speech processing only when voice activity is detected, reducing computational load and latency for silent periods.
Speech-to-Text (STT) Backends
The STT stage supports multiple backends through handler classes in src/speech_to_speech/STT/:
WhisperSTTHandler: Uses OpenAI Whisper models (viawhisper_stt_handler.py)FasterWhisperSTTHandler: Optimized Whisper implementation via faster-whisperParaformerSTTHandler: Alibaba Paraformer supportParakeetTDTSTTHandler: NVIDIA Parakeet TDT implementationMLXAudioWhisperSTTHandler: Apple Silicon optimized Whisper via mlx-audio
Each handler implements the same interface and can be selected via the stt argument in the configuration.
Language Model (LLM) Integration
The LLM stage in src/speech_to_speech/LLM/ supports multiple inference backends:
LanguageModelHandler: Hugging Face Transformers and MLX-LM support (vialanguage_model.py)ResponsesAPILanguageModelHandler: OpenAI-compatible API integration
These handlers process transcriptions and generate text responses that feed into the TTS stage.
Text-to-Speech (TTS) Synthesis
The final stage converts text to speech using handlers in src/speech_to_speech/TTS/:
Qwen3TTSHandler: Supports bothmlx-audioandfaster-qwen3-ttsbackends (detailed inqwen3_tts_handler.py)ChatTTSHandler: ChatTTS integrationFacebookMMSTTSHandler: Facebook MMS TTS modelsPocketTTSHandler: Pocket TTS lightweight implementationKokoroTTSHandler: Kokoro TTS support
The Qwen3TTSHandler specifically supports the Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice model and outputs NumPy int16 arrays suitable for direct audio playback or WAV file writing.
Practical Implementation Examples
The following examples demonstrate how to use the Python APIs for speech-to-speech in different scenarios, from simple command-line execution to embedded server deployment.
Command-Line Interface Usage
For quick testing or scripting, you can invoke the pipeline through the main() function:
# example.py
from speech_to_speech.s2s_pipeline import main
if __name__ == "__main__":
main()
Run from the terminal with specific backend selections:
python example.py \
--mode local \
--stt whisper \
--tts qwen3 \
--llm_backend transformers \
--device cpu
This approach uses parse_arguments() internally and supports all command-line options defined in the repository.
Programmatic Pipeline Control
For integration into existing Python applications, instantiate the pipeline components directly:
from speech_to_speech.s2s_pipeline import (
parse_arguments,
initialize_queues_and_events,
build_pipeline,
)
# Parse configuration (can construct args manually instead of CLI)
args = parse_arguments()
# Initialize communication infrastructure
queues_and_events = initialize_queues_and_events()
# Build the complete pipeline
pipeline_manager = build_pipeline(
module_kwargs=args.module_kwargs,
socket_receiver_kwargs=args.socket_receiver_kwargs,
socket_sender_kwargs=args.socket_sender_kwargs,
websocket_streamer_kwargs=args.websocket_streamer_kwargs,
vad_handler_kwargs=args.vad_handler_kwargs,
whisper_stt_handler_kwargs=args.whisper_stt_handler_kwargs,
faster_whisper_stt_handler_kwargs=args.faster_whisper_stt_handler_kwargs,
paraformer_stt_handler_kwargs=args.paraformer_stt_handler_kwargs,
mlx_audio_whisper_stt_handler_kwargs=args.mlx_audio_whisper_stt_handler_kwargs,
parakeet_tdt_stt_handler_kwargs=args.parakeet_tdt_stt_handler_kwargs,
language_model_handler_kwargs=args.language_model_handler_kwargs,
responses_api_language_model_handler_kwargs=args.responses_api_language_model_handler_kwargs,
chat_tts_handler_kwargs=args.chat_tts_handler_kwargs,
facebook_mms_tts_handler_kwargs=args.facebook_mms_tts_handler_kwargs,
pocket_tts_handler_kwargs=args.pocket_tts_handler_kwargs,
kokoro_tts_handler_kwargs=args.kokoro_tts_handler_kwargs,
qwen3_tts_handler_kwargs=args.qwen3_tts_handler_kwargs,
queues_and_events=queues_and_events,
)
# Execute the pipeline
pipeline_manager.start()
pipeline_manager.wait() # Blocks until shutdown (Ctrl-C)
Direct Handler Access
For specialized use cases, instantiate individual handlers directly. This example uses the Qwen-3 TTS handler:
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from threading import Event
# Setup dummy input structure
class TTSInput:
def __init__(self, text):
self.text = text
self.language_code = None
self.turn_id = "turn0"
self.turn_revision = 0
self.runtime_config = None
self.response = None
# Instantiate handler
handler = Qwen3TTSHandler()
handler.setup(
should_listen=Event(),
model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
device="cpu",
)
# Process text to audio
tts_input = TTSInput("Hello, this demonstrates the Qwen-3 TTS handler.")
for audio_chunk in handler.process(tts_input):
# audio_chunk is NumPy int16 array
print(f"Generated {len(audio_chunk)} samples")
# Save to file: sf.write("output.wav", audio_chunk, 16000, subtype="PCM_16")
OpenAI Realtime Server Deployment
Deploy the pipeline as an OpenAI-compatible realtime server using the WebSocket implementation in src/speech_to_speech/api/openai_realtime/server.py:
from speech_to_speech.api.openai_realtime.server import RealtimeServer
from speech_to_speech.s2s_pipeline import (
parse_arguments,
initialize_queues_and_events,
build_pipeline
)
args = parse_arguments()
queues = initialize_queues_and_events()
pipeline_manager = build_pipeline(
# ... include all required kwargs as shown above ...
queues_and_events=queues,
)
# Start the realtime server embedded in the pipeline
pipeline_manager.start()
pipeline_manager.wait()
When running with --mode realtime, the pipeline exposes the OpenAI Realtime protocol over WebSocket, allowing client applications to stream audio in and receive synthesized speech out.
Summary
The Hugging Face speech-to-speech repository provides a production-ready Python API for building speech-to-speech pipelines:
- Three core functions (
parse_arguments(),initialize_queues_and_events(),build_pipeline()) insrc/speech_to_speech/s2s_pipeline.pyprovide the high-level interface for pipeline construction. - Modular handler architecture supports interchangeable VAD, STT, LLM, and TTS components via classes in
src/speech_to_speech/VAD/,STT/,LLM/, andTTS/directories. - Thread-safe execution is managed by the
ThreadManagerclass insrc/speech_to_speech/utils/thread_manager.py, coordinating handlers through queue-based communication. - Multiple backends are supported including Whisper, Faster-Whisper, Qwen-3 TTS, ChatTTS, and MLX-Audio for Apple Silicon optimization.
- OpenAI Realtime compatibility is available through the server implementation in
src/speech_to_speech/api/openai_realtime/server.py.
Frequently Asked Questions
What Python version is required for the speech-to-speech API?
The repository requires Python 3.9 or higher due to its use of modern typing features and asynchronous programming patterns. The API relies heavily on threading and queue modules from the standard library, along with specific versions of PyTorch, Transformers, and optional dependencies like mlx-audio for Apple Silicon or faster-whisper for optimized inference.
Can I use individual handlers without the full pipeline?
Yes, you can import and instantiate individual handlers directly from their respective modules. For example, import Qwen3TTSHandler from src/speech_to_speech/TTS/qwen3_tts_handler.py or WhisperSTTHandler from src/speech_to_speech/STT/whisper_stt_handler.py. Each handler implements a standard interface with setup() and process() methods, allowing you to use them in isolation or integrate them into custom pipeline architectures outside of the provided ThreadManager.
How do I switch between different STT or TTS backends?
You specify the backend through the arguments passed to build_pipeline() or via command-line flags when using parse_arguments(). For STT, use arguments like --stt whisper, --stt faster_whisper, or --stt paraformer, which correspond to the whisper_stt_handler_kwargs, faster_whisper_stt_handler_kwargs, and other parameters. For TTS, options include --tts qwen3, --tts chat_tts, or --tts kokoro, which route to their respective handler keyword arguments in the build function.
Is the pipeline suitable for real-time production deployment?
Yes, the architecture is designed for low-latency real-time processing. The queue-based communication system in src/speech_to_speech/s2s_pipeline.py uses queue.Queue subclasses that enable concurrent processing across threads without blocking. For production web deployment, use the RealtimeServer class in src/speech_to_speech/api/openai_realtime/server.py with --mode realtime, which exposes the pipeline via WebSocket following the OpenAI Realtime API specification. The ThreadManager provides robust lifecycle management with proper start/stop semantics for long-running services.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →