How to Integrate Hugging Face Speech‑to‑Speech into a Python Application
Import the orchestration functions from speech_to_speech.s2s_pipeline, prepare your argument dataclasses, initialize queues and events, build the pipeline with build_pipeline(), and start the returned ThreadManager to run the full VAD→STT→LLM→TTS voice agent inside any Python service.
Integrating Hugging Face speech‑to‑speech into a Python application lets you embed a complete modular voice agent directly into web servers, background workers, or desktop apps. The pipeline runs each stage—Voice Activity Detection (VAD), Speech‑to‑Text (STT), Language Model (LLM), and Text‑to‑Speech (TTS)—in its own thread, communicating through thread‑safe queues. This article walks through the exact source code paths and Python APIs you need to integrate the pipeline programmatically.
Four Run Modes for Different Deployment Scenarios
The repository supports four transport modes, each suited to different integration patterns:
| Mode | Transport | Best For |
|---|---|---|
realtime |
OpenAI Realtime protocol over WebSocket/WebRTC | Building voice‑assistant APIs compatible with OpenAI clients |
local |
Direct microphone and speaker access | Quick prototyping on a single machine |
raw‑websocket |
Plain PCM over WebSocket | Minimal custom clients that stream raw audio |
socket |
Raw TCP socket | Simple remote audio streaming setups |
Select your mode via ModuleArguments.mode in src/speech_to_speech/arguments_classes/module_arguments.py. The core orchestration that branches on this mode lives in src/speech_to_speech/s2s_pipeline.py.
Core Integration Steps
Follow these five steps to embed the pipeline in any Python codebase:
-
Create argument dataclasses — Use
ParsedArgumentsor parse CLI‑style flags withparse_arguments(). The top‑level container groups handler‑specific classes likeWhisperSTTHandlerArgumentsandChatTTSHandlerArguments. -
Prepare arguments — Call
prepare_all_args()to expand prefixed flags, apply macOS defaults, and isolate parameters per handler. -
Initialize infrastructure — Use
initialize_queues_and_events()to create thread‑safe queues and synchronization events. -
Build the pipeline —
build_pipeline()returns aThreadManagerthat wires all handlers together with your chosen transport. -
Run the manager — Call
start()thenwait(); handle shutdown via signals or explicitstop().
Minimal Library Usage: Running Locally in a Script
This snippet demonstrates the simplest embed pattern—running the full pipeline with local audio I/O:
# --------------------------------------------------
# 1️⃣ Import the public API
# --------------------------------------------------
from speech_to_speech.s2s_pipeline import (
parse_arguments,
prepare_all_args,
initialize_queues_and_events,
build_pipeline,
)
# --------------------------------------------------
# 2️⃣ Parse CLI‑style arguments
# --------------------------------------------------
import sys
sys.argv = [
"script_name",
"--mode", "local",
"--stt", "whisper",
"--tts", "qwen3",
"--model_name", "gpt-4o-mini",
"--responses_api_api_key", "sk-dummy",
]
args = parse_arguments()
# --------------------------------------------------
# 3️⃣ Prepare and normalize all arguments
# --------------------------------------------------
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
)
# --------------------------------------------------
# 4️⃣ Allocate shared queues and events
# --------------------------------------------------
queues_and_events = initialize_queues_and_events()
# --------------------------------------------------
# 5️⃣ Build the pipeline (returns ThreadManager)
# --------------------------------------------------
pipeline = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
queues_and_events,
)
# --------------------------------------------------
# 6️⃣ Start the pipeline
# --------------------------------------------------
pipeline.start()
pipeline.wait() # blocks until interrupted
The parse_arguments() helper uses HfArgumentParser (the same validator as the CLI), ensuring correctness. The ThreadManager abstracts thread lifecycle—no manual thread management required.
Embedding the Realtime Server in FastAPI
For production APIs, run the realtime mode inside an async framework. This pattern launches the pipeline in a background executor without blocking your event loop:
import asyncio
from speech_to_speech.s2s_pipeline import (
parse_arguments,
prepare_all_args,
initialize_queues_and_events,
build_pipeline,
)
async def start_realtime():
import sys
sys.argv = [
"script",
"--mode", "realtime",
"--stt", "parakeet-tdt",
"--tts", "qwen3",
"--llm_backend", "responses-api",
"--model_name", "gpt-4o-mini",
"--responses_api_api_key", "sk-dummy",
"--enable_llm_proxy", "true",
]
args = parse_arguments()
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
)
queues_and_events = initialize_queues_and_events()
manager = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
queues_and_events,
)
loop = asyncio.get_running_loop()
await loop.run_in_executor(None, manager.start)
# Server now listening on ws://localhost:8765/v1/realtime
asyncio.create_task(start_realtime())
The realtime server is OpenAI‑compatible—any standard client can connect. Enable --enable_llm_proxy to expose /v1/chat/completions or /v1/responses for downstream services.
Isolated LLM Usage: Bypassing Audio Components
For testing or batch text generation, invoke the LLM handler directly without audio I/O:
from speech_to_speech.s2s_pipeline import (
parse_arguments,
prepare_all_args,
get_llm_handler,
)
from speech_to_speech.pipeline.handler_types import LLMIn, LLMOut
from queue import Queue
from threading import Event
# Parse minimal LLM‑focused args
import sys
sys.argv = [
"script",
"--llm_backend", "responses-api",
"--model_name", "gpt-4o-mini",
"--responses_api_api_key", "sk-dummy",
]
args = parse_arguments()
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
)
# Create minimal handler infrastructure
stop_event = Event()
input_q: Queue[LLMIn] = Queue()
output_q: Queue[LLMOut] = Queue()
# Build LLM handler
llm_handler = get_llm_handler(
module_kwargs=args.module_kwargs,
stop_event=stop_event,
text_prompt_queue=input_q,
lm_response_queue=output_q,
language_model_handler_kwargs=args.language_model_handler_kwargs,
responses_api_language_model_handler_kwargs=args.responses_api_language_model_handler_kwargs,
)
# Push prompt and collect streaming response
input_q.put({"type": "text", "content": "What is the capital of France?"})
while True:
item = output_q.get()
if item.get("type") == "final":
print("LLM answer:", item["content"])
break
print("partial:", item["content"])
This uses the same queue‑based contract as the full pipeline, making it trivial to test LLM integrations in isolation.
Key Source Files Reference
| File | Purpose |
|---|---|
src/speech_to_speech/s2s_pipeline.py |
Top‑level orchestration, parse_arguments(), build_pipeline(), ThreadManager |
src/speech_to_speech/arguments_classes/module_arguments.py |
ModuleArguments dataclass—global mode, device, backend selections |
src/speech_to_speech/arguments_classes/*.py |
Handler‑specific options (STT, LLM, TTS, VAD, sockets) |
src/speech_to_speech/VAD/vad_handler.py |
Voice activity detection and turn‑taking logic |
src/speech_to_speech/STT/* |
STT backends: whisper, parakeet-tdt, faster-whisper, mlx-audio-whisper, paraformer |
src/speech_to_speech/LLM/* |
LLM backends: responses-api, chat-completions, transformers, mlx-lm |
src/speech_to_speech/TTS/* |
TTS backends: qwen3, kokoro, pocket, chat-tts, facebook-mms |
src/speech_to_speech/api/openai_realtime/README.md |
OpenAI Realtime protocol specification |
Summary
-
Import from
speech_to_speech.s2s_pipelineto accessparse_arguments(),prepare_all_args(),initialize_queues_and_events(), andbuild_pipeline(). -
Choose your transport mode (
local,realtime,raw‑websocket,socket) viaModuleArguments.mode. -
Build and start a
ThreadManager—it handles all thread lifecycle and graceful shutdown. -
Swap components freely by changing argument values for STT, LLM, or TTS backends without code changes.
-
Embed anywhere: Flask, FastAPI, Celery workers, Jupyter notebooks, or desktop applications all work identically.
Frequently Asked Questions
Can I run the speech‑to‑speech pipeline without installing the CLI tool?
Yes. The pipeline is pure Python—install the package with pip install -e . from the repository, then import speech_to_speech.s2s_pipeline directly. No CLI invocation is required at runtime.
How do I switch between different STT or TTS backends programmatically?
Change the corresponding argument in your sys.argv or dataclass construction. For example, set --stt parakeet-tdt or --tts kokoro before calling parse_arguments(). The build_pipeline() function instantiates the correct handler class based on these values.
Is the Realtime server compatible with the OpenAI Python SDK?
Yes. The realtime mode implements the OpenAI Realtime protocol as documented in src/speech_to_speech/api/openai_realtime/README.md. Any client that speaks this protocol—including the official openai SDK or WebRTC browsers—can connect to ws://localhost:8765/v1/realtime.
Can I use the LLM handler for text‑only applications without audio?
Absolutely. Import get_llm_handler() from s2s_pipeline, create your own Queue objects for text_prompt_queue and lm_response_queue, and push text prompts directly. This bypasses VAD, STT, and TTS entirely while using the same handler implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →