How to Use the speech-to-speech Python API: A Complete Guide
The speech-to-speech Python API exposes a modular voice-assistant pipeline through the s2s_pipeline.py module, allowing developers to programmatically configure VAD, STT, LLM, and TTS components using factory functions and the ThreadManager class.
The huggingface/speech-to-speech repository implements a fully modular voice-assistant architecture that can be invoked from Python code or the provided CLI. Unlike monolithic speech systems, this framework separates concerns into four interchangeable handlers that communicate through typed queues, enabling low-latency streaming and easy backend swapping. Whether building a local offline assistant or an OpenAI Realtime-compatible server, the speech-to-speech Python API provides granular control over every pipeline stage.
Architecture Overview
The system consists of four interchangeable components that run in parallel threads and exchange data through typed queues:
- Voice-Activity Detection (VAD) – Detects speech boundaries using Silero VAD and optionally streams live transcription events.
- Speech-to-Text (STT) – Converts spoken input into text using backends like Parakeet TDT, Whisper, Faster-Whisper, MLX-Audio-Whisper, or Paraformer.
- Large Language Model (LLM) – Generates responses via local models (Transformers, mlx-lm) or OpenAI-compatible APIs (
responses-apiorchat-completions). - Text-to-Speech (TTS) – Synthesizes audio using Qwen3-TTS (default), Pocket TTS, ChatTTS, Kokoro-82M, or Facebook MMS.
These handlers communicate through typed queues (AudioInItem, STTOutItem, LMOutItem, TTSInItem) and synchronize using Event flags (stop_event, should_listen, response_playing). The orchestration logic in src/speech_to_speech/s2s_pipeline.py uses HfArgumentParser to parse CLI arguments or JSON configs, normalizes argument prefixes (e.g., --stt_*, --tts_*) via rename_args, and builds concrete handler instances through factory functions.
Transport Modes
The pipeline supports four transport modes controlled via the --mode argument:
| Mode | Transport | Use Case |
|---|---|---|
realtime |
OpenAI Realtime API over WebSocket/WebRTC | Build voice assistants compatible with OpenAI clients |
local |
Direct microphone & speakers | Quick prototyping on a single machine |
raw-websocket |
Raw PCM over WebSocket | Custom clients streaming 16 kHz int16 PCM |
socket |
Raw PCM over TCP | Simple LAN streaming without Realtime features |
Programmatic API Usage
To start the pipeline programmatically instead of via CLI, import the core functions from src/speech_to_speech/s2s_pipeline.py and manage the lifecycle through ThreadManager:
from speech_to_speech.s2s_pipeline import (
parse_arguments,
prepare_all_args,
initialize_queues_and_events,
build_pipeline,
)
# Parse CLI arguments or construct dataclasses manually
args = parse_arguments()
# Normalize argument prefixes and apply device defaults
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
)
# Initialize shared queues and synchronization primitives
queues_and_events = initialize_queues_and_events()
# Build handlers via factory functions (get_stt_handler, get_llm_handler, get_tts_handler)
pipeline_manager = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
queues_and_events,
)
# Start and block until SIGINT/SIGTERM
pipeline_manager.start()
pipeline_manager.wait()
The build_pipeline function internally uses factory functions (get_stt_handler, get_llm_handler, get_tts_handler) to instantiate concrete handlers based on your configuration, then wires them together under a ThreadManager that handles graceful shutdown.
CLI Configuration Examples
While the Python API provides full control, you can also leverage the CLI entry point which calls main() in s2s_pipeline.py:
Start a realtime server (default VAD → Parakeet TDT → Responses-API → Qwen3-TTS):
export OPENAI_API_KEY=sk-...
speech-to-speech
Swap STT to Whisper:
speech-to-speech --stt whisper --stt_model_name openai/whisper-base
Run locally with MLX LLM on Apple Silicon:
speech-to-speech --mode local --llm_backend mlx-lm --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16
Use Pocket TTS:
speech-to-speech --tts pocket --pocket_tts_voice jean --pocket_tts_device cpu
Enable the LLM proxy with --enable_llm_proxy to expose the configured LLM as a stand-alone OpenAI-compatible endpoint, letting other services call the model while the voice pipeline runs uninterrupted.
Key Source Files
Understanding the repository structure helps when extending the API:
src/speech_to_speech/s2s_pipeline.py– Core orchestration containingparse_arguments(),prepare_all_args(),initialize_queues_and_events(), andbuild_pipeline().src/speech_to_speech/utils/thread_manager.py– ImplementsThreadManagerfor handler lifecycle management.src/speech_to_speech/arguments_classes/*.py– Dataclasses defining CLI/JSON arguments for each component.src/speech_to_speech/STT/*_handler.py– Concrete STT implementations (Whisper, Parakeet TDT, etc.).src/speech_to_speech/LLM/*_handler.py– LLM backends including local and API variants.src/speech_to_speech/TTS/*_handler.py– TTS synthesizers (Qwen3-TTS, Kokoro, Pocket, etc.).src/speech_to_speech/pipeline/*.py– Queue definitions and speculative turn tracking.
Summary
- The speech-to-speech Python API in
s2s_pipeline.pyprovides a modular interface for building voice assistants through four core handlers: VAD, STT, LLM, and TTS. - Factory functions (
get_stt_handler,get_llm_handler,get_tts_handler) instantiate backends based on configuration dataclasses, whileThreadManagercoordinates parallel execution. - Typed queues (
AudioInItem,STTOutItem,LMOutItem,TTSInItem) and Event flags (stop_event,should_listen) enable safe thread communication and graceful shutdown. - The API supports four transport modes (
realtime,local,raw-websocket,socket) and can expose an LLM proxy for external API access. - Configuration occurs through
HfArgumentParserwith prefixed arguments (--stt_*,--tts_*) normalized byprepare_all_args().
Frequently Asked Questions
How do I switch between different STT backends in the Python API?
Pass the specific handler kwargs to prepare_all_args() and build_pipeline(). The factory function get_stt_handler selects the concrete implementation based on the stt argument value (e.g., "whisper", "parakeet_tdt", "faster_whisper"). For example, to use Whisper, ensure args.whisper_stt_handler_kwargs contains model_name="openai/whisper-base" and the STT selector is set to "whisper".
What is the difference between realtime and local transport modes?
The realtime mode exposes an OpenAI Realtime API-compatible WebSocket endpoint at ws://localhost:8765/v1/realtime, allowing external clients to connect using standard OpenAI SDKs. The local mode bypasses network transport and connects directly to system microphone and speakers via the local audio handler, making it ideal for single-machine prototyping without network overhead.
How does the pipeline handle graceful shutdown?
The ThreadManager class (in src/speech_to_speech/utils/thread_manager.py) monitors a stop_event threading.Event shared across all handlers. When the process receives SIGINT or SIGTERM, the event is set, causing each handler's run() loop to exit cleanly. The main thread then calls join() on each handler thread via pipeline_manager.wait() to ensure all audio buffers are flushed before termination.
Can I use the speech-to-speech pipeline with a custom LLM server?
Yes. Set --enable_llm_proxy to expose the configured LLM as an OpenAI-compatible HTTP endpoint, or use --llm_backend with an OpenAI-compatible base URL. The LanguageModelHandler supports both responses-api and chat-completions formats, allowing you to proxy requests to custom servers while maintaining the voice pipeline's streaming capabilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →