How to Integrate Hugging Face Speech‑to‑Speech into a Python Application

Import the orchestration functions from speech_to_speech.s2s_pipeline, prepare your argument dataclasses, initialize queues and events, build the pipeline with build_pipeline(), and start the returned ThreadManager to run the full VAD→STT→LLM→TTS voice agent inside any Python service.

Integrating Hugging Face speech‑to‑speech into a Python application lets you embed a complete modular voice agent directly into web servers, background workers, or desktop apps. The pipeline runs each stage—Voice Activity Detection (VAD), Speech‑to‑Text (STT), Language Model (LLM), and Text‑to‑Speech (TTS)—in its own thread, communicating through thread‑safe queues. This article walks through the exact source code paths and Python APIs you need to integrate the pipeline programmatically.

Four Run Modes for Different Deployment Scenarios

The repository supports four transport modes, each suited to different integration patterns:

Mode Transport Best For
realtime OpenAI Realtime protocol over WebSocket/WebRTC Building voice‑assistant APIs compatible with OpenAI clients
local Direct microphone and speaker access Quick prototyping on a single machine
raw‑websocket Plain PCM over WebSocket Minimal custom clients that stream raw audio
socket Raw TCP socket Simple remote audio streaming setups

Select your mode via ModuleArguments.mode in src/speech_to_speech/arguments_classes/module_arguments.py. The core orchestration that branches on this mode lives in src/speech_to_speech/s2s_pipeline.py.

Core Integration Steps

Follow these five steps to embed the pipeline in any Python codebase:

  1. Create argument dataclasses — Use ParsedArguments or parse CLI‑style flags with parse_arguments(). The top‑level container groups handler‑specific classes like WhisperSTTHandlerArguments and ChatTTSHandlerArguments.

  2. Prepare arguments — Call prepare_all_args() to expand prefixed flags, apply macOS defaults, and isolate parameters per handler.

  3. Initialize infrastructure — Use initialize_queues_and_events() to create thread‑safe queues and synchronization events.

  4. Build the pipeline — build_pipeline() returns a ThreadManager that wires all handlers together with your chosen transport.

  5. Run the manager — Call start() then wait(); handle shutdown via signals or explicit stop().

Minimal Library Usage: Running Locally in a Script

This snippet demonstrates the simplest embed pattern—running the full pipeline with local audio I/O:


# --------------------------------------------------

# 1️⃣ Import the public API

# --------------------------------------------------

from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# --------------------------------------------------

# 2️⃣ Parse CLI‑style arguments

# --------------------------------------------------

import sys
sys.argv = [
    "script_name",
    "--mode", "local",
    "--stt", "whisper",
    "--tts", "qwen3",
    "--model_name", "gpt-4o-mini",
    "--responses_api_api_key", "sk-dummy",
]

args = parse_arguments()

# --------------------------------------------------

# 3️⃣ Prepare and normalize all arguments

# --------------------------------------------------

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# --------------------------------------------------

# 4️⃣ Allocate shared queues and events

# --------------------------------------------------

queues_and_events = initialize_queues_and_events()

# --------------------------------------------------

# 5️⃣ Build the pipeline (returns ThreadManager)

# --------------------------------------------------

pipeline = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues_and_events,
)

# --------------------------------------------------

# 6️⃣ Start the pipeline

# --------------------------------------------------

pipeline.start()
pipeline.wait()  # blocks until interrupted

The parse_arguments() helper uses HfArgumentParser (the same validator as the CLI), ensuring correctness. The ThreadManager abstracts thread lifecycle—no manual thread management required.

Embedding the Realtime Server in FastAPI

For production APIs, run the realtime mode inside an async framework. This pattern launches the pipeline in a background executor without blocking your event loop:

import asyncio
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

async def start_realtime():
    import sys
    sys.argv = [
        "script",
        "--mode", "realtime",
        "--stt", "parakeet-tdt",
        "--tts", "qwen3",
        "--llm_backend", "responses-api",
        "--model_name", "gpt-4o-mini",
        "--responses_api_api_key", "sk-dummy",
        "--enable_llm_proxy", "true",
    ]

    args = parse_arguments()
    prepare_all_args(
        args.module_kwargs,
        args.whisper_stt_handler_kwargs,
        args.paraformer_stt_handler_kwargs,
        args.faster_whisper_stt_handler_kwargs,
        args.mlx_audio_whisper_stt_handler_kwargs,
        args.parakeet_tdt_stt_handler_kwargs,
        args.language_model_handler_kwargs,
        args.responses_api_language_model_handler_kwargs,
        args.chat_tts_handler_kwargs,
        args.facebook_mms_tts_handler_kwargs,
        args.pocket_tts_handler_kwargs,
        args.kokoro_tts_handler_kwargs,
        args.qwen3_tts_handler_kwargs,
    )
    queues_and_events = initialize_queues_and_events()
    manager = build_pipeline(
        args.module_kwargs,
        args.socket_receiver_kwargs,
        args.socket_sender_kwargs,
        args.websocket_streamer_kwargs,
        args.vad_handler_kwargs,
        args.whisper_stt_handler_kwargs,
        args.faster_whisper_stt_handler_kwargs,
        args.paraformer_stt_handler_kwargs,
        args.mlx_audio_whisper_stt_handler_kwargs,
        args.parakeet_tdt_stt_handler_kwargs,
        args.language_model_handler_kwargs,
        args.responses_api_language_model_handler_kwargs,
        args.chat_tts_handler_kwargs,
        args.facebook_mms_tts_handler_kwargs,
        args.pocket_tts_handler_kwargs,
        args.kokoro_tts_handler_kwargs,
        args.qwen3_tts_handler_kwargs,
        queues_and_events,
    )

    loop = asyncio.get_running_loop()
    await loop.run_in_executor(None, manager.start)
    # Server now listening on ws://localhost:8765/v1/realtime

asyncio.create_task(start_realtime())

The realtime server is OpenAI‑compatible—any standard client can connect. Enable --enable_llm_proxy to expose /v1/chat/completions or /v1/responses for downstream services.

Isolated LLM Usage: Bypassing Audio Components

For testing or batch text generation, invoke the LLM handler directly without audio I/O:

from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    get_llm_handler,
)
from speech_to_speech.pipeline.handler_types import LLMIn, LLMOut
from queue import Queue
from threading import Event

# Parse minimal LLM‑focused args

import sys
sys.argv = [
    "script",
    "--llm_backend", "responses-api",
    "--model_name", "gpt-4o-mini",
    "--responses_api_api_key", "sk-dummy",
]
args = parse_arguments()
prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Create minimal handler infrastructure

stop_event = Event()
input_q: Queue[LLMIn] = Queue()
output_q: Queue[LLMOut] = Queue()

# Build LLM handler

llm_handler = get_llm_handler(
    module_kwargs=args.module_kwargs,
    stop_event=stop_event,
    text_prompt_queue=input_q,
    lm_response_queue=output_q,
    language_model_handler_kwargs=args.language_model_handler_kwargs,
    responses_api_language_model_handler_kwargs=args.responses_api_language_model_handler_kwargs,
)

# Push prompt and collect streaming response

input_q.put({"type": "text", "content": "What is the capital of France?"})
while True:
    item = output_q.get()
    if item.get("type") == "final":
        print("LLM answer:", item["content"])
        break
    print("partial:", item["content"])

This uses the same queue‑based contract as the full pipeline, making it trivial to test LLM integrations in isolation.

Key Source Files Reference

File Purpose
src/speech_to_speech/s2s_pipeline.py Top‑level orchestration, parse_arguments(), build_pipeline(), ThreadManager
src/speech_to_speech/arguments_classes/module_arguments.py ModuleArguments dataclass—global mode, device, backend selections
src/speech_to_speech/arguments_classes/*.py Handler‑specific options (STT, LLM, TTS, VAD, sockets)
src/speech_to_speech/VAD/vad_handler.py Voice activity detection and turn‑taking logic
src/speech_to_speech/STT/* STT backends: whisper, parakeet-tdt, faster-whisper, mlx-audio-whisper, paraformer
src/speech_to_speech/LLM/* LLM backends: responses-api, chat-completions, transformers, mlx-lm
src/speech_to_speech/TTS/* TTS backends: qwen3, kokoro, pocket, chat-tts, facebook-mms
src/speech_to_speech/api/openai_realtime/README.md OpenAI Realtime protocol specification

Summary

  • Import from speech_to_speech.s2s_pipeline to access parse_arguments(), prepare_all_args(), initialize_queues_and_events(), and build_pipeline().

  • Choose your transport mode (local, realtime, raw‑websocket, socket) via ModuleArguments.mode.

  • Build and start a ThreadManager—it handles all thread lifecycle and graceful shutdown.

  • Swap components freely by changing argument values for STT, LLM, or TTS backends without code changes.

  • Embed anywhere: Flask, FastAPI, Celery workers, Jupyter notebooks, or desktop applications all work identically.

Frequently Asked Questions

Can I run the speech‑to‑speech pipeline without installing the CLI tool?

Yes. The pipeline is pure Python—install the package with pip install -e . from the repository, then import speech_to_speech.s2s_pipeline directly. No CLI invocation is required at runtime.

How do I switch between different STT or TTS backends programmatically?

Change the corresponding argument in your sys.argv or dataclass construction. For example, set --stt parakeet-tdt or --tts kokoro before calling parse_arguments(). The build_pipeline() function instantiates the correct handler class based on these values.

Is the Realtime server compatible with the OpenAI Python SDK?

Yes. The realtime mode implements the OpenAI Realtime protocol as documented in src/speech_to_speech/api/openai_realtime/README.md. Any client that speaks this protocol—including the official openai SDK or WebRTC browsers—can connect to ws://localhost:8765/v1/realtime.

Can I use the LLM handler for text‑only applications without audio?

Absolutely. Import get_llm_handler() from s2s_pipeline, create your own Queue objects for text_prompt_queue and lm_response_queue, and push text prompts directly. This bypasses VAD, STT, and TTS entirely while using the same handler implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →