# How to Integrate Hugging Face Speech‑to‑Speech into a Python Application

> Integrate Hugging Face speech-to-speech into your Python app. Learn to build and run a voice agent pipeline including VAD, STT, LLM, and TTS. Get started today.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-02

---

**Import the orchestration functions from `speech_to_speech.s2s_pipeline`, prepare your argument dataclasses, initialize queues and events, build the pipeline with `build_pipeline()`, and start the returned `ThreadManager` to run the full VAD→STT→LLM→TTS voice agent inside any Python service.**

Integrating Hugging Face **speech‑to‑speech** into a Python application lets you embed a complete modular voice agent directly into web servers, background workers, or desktop apps. The pipeline runs each stage—**Voice Activity Detection (VAD)**, **Speech‑to‑Text (STT)**, **Language Model (LLM)**, and **Text‑to‑Speech (TTS)**—in its own thread, communicating through thread‑safe queues. This article walks through the exact source code paths and Python APIs you need to integrate the pipeline programmatically.

## Four Run Modes for Different Deployment Scenarios

The repository supports four transport modes, each suited to different integration patterns:

| Mode | Transport | Best For |
|------|-----------|----------|
| `realtime` | OpenAI Realtime protocol over WebSocket/WebRTC | Building voice‑assistant APIs compatible with OpenAI clients |
| `local` | Direct microphone and speaker access | Quick prototyping on a single machine |
| `raw‑websocket` | Plain PCM over WebSocket | Minimal custom clients that stream raw audio |
| `socket` | Raw TCP socket | Simple remote audio streaming setups |

Select your mode via `ModuleArguments.mode` in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py). The core orchestration that branches on this mode lives in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

## Core Integration Steps

Follow these five steps to embed the pipeline in any Python codebase:

1. **Create argument dataclasses** — Use `ParsedArguments` or parse CLI‑style flags with `parse_arguments()`. The top‑level container groups handler‑specific classes like `WhisperSTTHandlerArguments` and `ChatTTSHandlerArguments`.

2. **Prepare arguments** — Call `prepare_all_args()` to expand prefixed flags, apply macOS defaults, and isolate parameters per handler.

3. **Initialize infrastructure** — Use `initialize_queues_and_events()` to create thread‑safe queues and synchronization events.

4. **Build the pipeline** — `build_pipeline()` returns a `ThreadManager` that wires all handlers together with your chosen transport.

5. **Run the manager** — Call `start()` then `wait()`; handle shutdown via signals or explicit `stop()`.

## Minimal Library Usage: Running Locally in a Script

This snippet demonstrates the simplest embed pattern—running the full pipeline with local audio I/O:

```python

# --------------------------------------------------

# 1️⃣ Import the public API

# --------------------------------------------------

from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# --------------------------------------------------

# 2️⃣ Parse CLI‑style arguments

# --------------------------------------------------

import sys
sys.argv = [
    "script_name",
    "--mode", "local",
    "--stt", "whisper",
    "--tts", "qwen3",
    "--model_name", "gpt-4o-mini",
    "--responses_api_api_key", "sk-dummy",
]

args = parse_arguments()

# --------------------------------------------------

# 3️⃣ Prepare and normalize all arguments

# --------------------------------------------------

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# --------------------------------------------------

# 4️⃣ Allocate shared queues and events

# --------------------------------------------------

queues_and_events = initialize_queues_and_events()

# --------------------------------------------------

# 5️⃣ Build the pipeline (returns ThreadManager)

# --------------------------------------------------

pipeline = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues_and_events,
)

# --------------------------------------------------

# 6️⃣ Start the pipeline

# --------------------------------------------------

pipeline.start()
pipeline.wait()  # blocks until interrupted

```

The `parse_arguments()` helper uses `HfArgumentParser` (the same validator as the CLI), ensuring correctness. The `ThreadManager` abstracts thread lifecycle—no manual thread management required.

## Embedding the Realtime Server in FastAPI

For production APIs, run the **realtime** mode inside an async framework. This pattern launches the pipeline in a background executor without blocking your event loop:

```python
import asyncio
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

async def start_realtime():
    import sys
    sys.argv = [
        "script",
        "--mode", "realtime",
        "--stt", "parakeet-tdt",
        "--tts", "qwen3",
        "--llm_backend", "responses-api",
        "--model_name", "gpt-4o-mini",
        "--responses_api_api_key", "sk-dummy",
        "--enable_llm_proxy", "true",
    ]

    args = parse_arguments()
    prepare_all_args(
        args.module_kwargs,
        args.whisper_stt_handler_kwargs,
        args.paraformer_stt_handler_kwargs,
        args.faster_whisper_stt_handler_kwargs,
        args.mlx_audio_whisper_stt_handler_kwargs,
        args.parakeet_tdt_stt_handler_kwargs,
        args.language_model_handler_kwargs,
        args.responses_api_language_model_handler_kwargs,
        args.chat_tts_handler_kwargs,
        args.facebook_mms_tts_handler_kwargs,
        args.pocket_tts_handler_kwargs,
        args.kokoro_tts_handler_kwargs,
        args.qwen3_tts_handler_kwargs,
    )
    queues_and_events = initialize_queues_and_events()
    manager = build_pipeline(
        args.module_kwargs,
        args.socket_receiver_kwargs,
        args.socket_sender_kwargs,
        args.websocket_streamer_kwargs,
        args.vad_handler_kwargs,
        args.whisper_stt_handler_kwargs,
        args.faster_whisper_stt_handler_kwargs,
        args.paraformer_stt_handler_kwargs,
        args.mlx_audio_whisper_stt_handler_kwargs,
        args.parakeet_tdt_stt_handler_kwargs,
        args.language_model_handler_kwargs,
        args.responses_api_language_model_handler_kwargs,
        args.chat_tts_handler_kwargs,
        args.facebook_mms_tts_handler_kwargs,
        args.pocket_tts_handler_kwargs,
        args.kokoro_tts_handler_kwargs,
        args.qwen3_tts_handler_kwargs,
        queues_and_events,
    )

    loop = asyncio.get_running_loop()
    await loop.run_in_executor(None, manager.start)
    # Server now listening on ws://localhost:8765/v1/realtime

asyncio.create_task(start_realtime())

```

The realtime server is **OpenAI‑compatible**—any standard client can connect. Enable `--enable_llm_proxy` to expose `/v1/chat/completions` or `/v1/responses` for downstream services.

## Isolated LLM Usage: Bypassing Audio Components

For testing or batch text generation, invoke the LLM handler directly without audio I/O:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    get_llm_handler,
)
from speech_to_speech.pipeline.handler_types import LLMIn, LLMOut
from queue import Queue
from threading import Event

# Parse minimal LLM‑focused args

import sys
sys.argv = [
    "script",
    "--llm_backend", "responses-api",
    "--model_name", "gpt-4o-mini",
    "--responses_api_api_key", "sk-dummy",
]
args = parse_arguments()
prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Create minimal handler infrastructure

stop_event = Event()
input_q: Queue[LLMIn] = Queue()
output_q: Queue[LLMOut] = Queue()

# Build LLM handler

llm_handler = get_llm_handler(
    module_kwargs=args.module_kwargs,
    stop_event=stop_event,
    text_prompt_queue=input_q,
    lm_response_queue=output_q,
    language_model_handler_kwargs=args.language_model_handler_kwargs,
    responses_api_language_model_handler_kwargs=args.responses_api_language_model_handler_kwargs,
)

# Push prompt and collect streaming response

input_q.put({"type": "text", "content": "What is the capital of France?"})
while True:
    item = output_q.get()
    if item.get("type") == "final":
        print("LLM answer:", item["content"])
        break
    print("partial:", item["content"])

```

This uses the same **queue‑based contract** as the full pipeline, making it trivial to test LLM integrations in isolation.

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) | Top‑level orchestration, `parse_arguments()`, `build_pipeline()`, `ThreadManager` |
| [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) | `ModuleArguments` dataclass—global mode, device, backend selections |
| `src/speech_to_speech/arguments_classes/*.py` | Handler‑specific options (STT, LLM, TTS, VAD, sockets) |
| [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) | Voice activity detection and turn‑taking logic |
| `src/speech_to_speech/STT/*` | STT backends: `whisper`, `parakeet-tdt`, `faster-whisper`, `mlx-audio-whisper`, `paraformer` |
| `src/speech_to_speech/LLM/*` | LLM backends: `responses-api`, `chat-completions`, `transformers`, `mlx-lm` |
| `src/speech_to_speech/TTS/*` | TTS backends: `qwen3`, `kokoro`, `pocket`, `chat-tts`, `facebook-mms` |
| [`src/speech_to_speech/api/openai_realtime/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/README.md) | OpenAI Realtime protocol specification |

## Summary

- **Import from `speech_to_speech.s2s_pipeline`** to access `parse_arguments()`, `prepare_all_args()`, `initialize_queues_and_events()`, and `build_pipeline()`.

- **Choose your transport mode** (`local`, `realtime`, `raw‑websocket`, `socket`) via `ModuleArguments.mode`.

- **Build and start a `ThreadManager`**—it handles all thread lifecycle and graceful shutdown.

- **Swap components freely** by changing argument values for STT, LLM, or TTS backends without code changes.

- **Embed anywhere**: Flask, FastAPI, Celery workers, Jupyter notebooks, or desktop applications all work identically.

## Frequently Asked Questions

### Can I run the speech‑to‑speech pipeline without installing the CLI tool?

Yes. The pipeline is pure Python—install the package with `pip install -e .` from the repository, then import `speech_to_speech.s2s_pipeline` directly. No CLI invocation is required at runtime.

### How do I switch between different STT or TTS backends programmatically?

Change the corresponding argument in your `sys.argv` or dataclass construction. For example, set `--stt parakeet-tdt` or `--tts kokoro` before calling `parse_arguments()`. The `build_pipeline()` function instantiates the correct handler class based on these values.

### Is the Realtime server compatible with the OpenAI Python SDK?

Yes. The `realtime` mode implements the OpenAI Realtime protocol as documented in [`src/speech_to_speech/api/openai_realtime/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/README.md). Any client that speaks this protocol—including the official `openai` SDK or WebRTC browsers—can connect to `ws://localhost:8765/v1/realtime`.

### Can I use the LLM handler for text‑only applications without audio?

Absolutely. Import `get_llm_handler()` from `s2s_pipeline`, create your own `Queue` objects for `text_prompt_queue` and `lm_response_queue`, and push text prompts directly. This bypasses VAD, STT, and TTS entirely while using the same handler implementation.