# How to Integrate Speech-to-Speech into a Python Application: A Complete Guide

> Integrate speech-to-speech into your Python app using the Hugging Face pipeline. Follow our guide for a four-stage process: VAD, STT, LLM, and TTS for seamless voice interaction.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-07

---

**You can integrate speech-to-speech into a Python application by using the `huggingface/speech-to-speech` pipeline's `main()` entry point for CLI usage, or by programmatically constructing the four-stage pipeline (VAD → STT → LLM → TTS) using the `build_pipeline()` function with argument dataclasses.**

The `huggingface/speech-to-speech` repository provides a modular, production-ready framework for building real-time voice conversations. It chains together Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) into a seamless pipeline that runs across multiple threads. This guide shows you how to integrate this system into your Python applications using both command-line and programmatic approaches.

## Understand the Four-Stage Pipeline Architecture

The pipeline processes audio through four distinct stages connected by queues and events:

1. **VAD (Voice Activity Detection)** - Detects when users start/stop speaking. Implemented in [`speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/VAD/vad_handler.py) and instantiated in `s2s_pipeline._build_pipeline_handlers`.

2. **STT (Speech-to-Text)** - Converts audio to text using Whisper, Faster-Whisper, Paraformer, or MLX-Audio-Whisper. Selected via `module_kwargs.stt` and created in `s2s_pipeline.get_stt_handler`.

3. **LLM (Language Model)** - Generates responses using Transformers, MLX-LM, or OpenAI-compatible APIs. Built in `s2s_pipeline.get_llm_handler`.

4. **TTS (Text-to-Speech)** - Synthesizes speech via ChatTTS, Facebook MMS, Pocket, Kokoro, or Qwen-3. Created in `s2s_pipeline.get_tts_handler`.

All stages communicate via thread-safe queues initialized through `initialize_queues_and_events()` and managed by the `ThreadManager` in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py).

## Method 1: Quick Integration via Command Line

For rapid prototyping, use the built-in CLI entry point.

This approach handles argument parsing via `HfArgumentParser`, normalization through `prepare_all_args`, and pipeline construction automatically.

```python
from speech_to_speech.s2s_pipeline import main

if __name__ == "__main__":
    main()

```

Run with command-line flags:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode local \
    --stt whisper \
    --tts qwen3 \
    --llm_backend transformers \
    --device cpu \
    --log_level info

```

The `main()` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) orchestrates the entire flow: parsing arguments, building handlers, and starting the `ThreadManager`.

## Method 2: Programmatic Integration

For embedding within existing applications, construct the pipeline manually using dataclasses from `src/speech_to_speech/arguments_classes/`.

### Step 1: Configure Arguments

Import and instantiate the specific argument classes for your chosen backends:

```python
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelHandlerArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSHandlerArguments

module_args = ModuleArguments(
    mode="local",          # Options: "local", "websocket", "realtime"

    stt="whisper",
    tts="qwen3",
    llm_backend="transformers",
    device="cpu",
    log_level="info",
)

whisper_args = WhisperSTTHandlerArguments(model_name="openai/whisper-base")
lm_args = LanguageModelHandlerArguments(model_name="Qwen/Qwen3-4B-Instruct-2507")
tts_args = Qwen3TTSHandlerArguments()

```

### Step 2: Normalize Arguments

Call `prepare_all_args()` to apply device optimizations and format `gen_kwargs`:

```python
from speech_to_speech.s2s_pipeline import prepare_all_args

prepare_all_args(
    module_args,
    whisper_args,
    WhisperSTTHandlerArguments(),  # Placeholder for paraformer

    WhisperSTTHandlerArguments(),  # Placeholder for faster-whisper

    WhisperSTTHandlerArguments(),  # Placeholder for mlx-audio-whisper

    WhisperSTTHandlerArguments(),  # Placeholder for parakeet-tdt

    lm_args,
    LanguageModelHandlerArguments(),  # Placeholder for responses-api

    tts_args,
    Qwen3TTSHandlerArguments(),  # Placeholder for other TTS handlers

    Qwen3TTSHandlerArguments(),
    Qwen3TTSHandlerArguments(),
    Qwen3TTSHandlerArguments(),
)

```

### Step 3: Initialize Communication Primitives

Create the queues and events that connect pipeline stages:

```python
from speech_to_speech.s2s_pipeline import initialize_queues_and_events

queues = initialize_queues_and_events()

```

### Step 4: Build and Run the Pipeline

Construct the pipeline with `build_pipeline()` and start execution:

```python
from speech_to_speech.s2s_pipeline import build_pipeline
from speech_to_speech.arguments_classes.socket_receiver_arguments import SocketReceiverArguments
from speech_to_speech.arguments_classes.socket_sender_arguments import SocketSenderArguments
from speech_to_speech.arguments_classes.websocket_streamer_arguments import WebSocketStreamerArguments
from speech_to_speech.arguments_classes.vad_handler_arguments import VADHandlerArguments
from speech_to_speech.arguments_classes.faster_whisper_stt_arguments import FasterWhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.paraformer_stt_arguments import ParaformerSTTHandlerArguments
from speech_to_speech.arguments_classes.mlx_audio_whisper_stt_arguments import MLXAudioWhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.parakeet_tdt_stt_arguments import ParakeetTDTSTTHandlerArguments
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import ResponsesApiLanguageModelHandlerArguments
from speech_to_speech.arguments_classes.chat_tts_arguments import ChatTTSHandlerArguments
from speech_to_speech.arguments_classes.facebook_mms_tts_arguments import FacebookMMSTTSHandlerArguments
from speech_to_speech.arguments_classes.pocket_tts_arguments import PocketTTSHandlerArguments
from speech_to_speech.arguments_classes.kokoro_tts_arguments import KokoroTTSHandlerArguments

pipeline_manager = build_pipeline(
    module_args,
    SocketReceiverArguments(),
    SocketSenderArguments(),
    WebSocketStreamerArguments(),
    VADHandlerArguments(),
    whisper_args,
    FasterWhisperSTTHandlerArguments(),
    ParaformerSTTHandlerArguments(),
    MLXAudioWhisperSTTHandlerArguments(),
    ParakeetTDTSTTHandlerArguments(),
    lm_args,
    ResponsesApiLanguageModelHandlerArguments(),
    ChatTTSHandlerArguments(),
    FacebookMMSTTSHandlerArguments(),
    PocketTTSHandlerArguments(),
    KokoroTTSHandlerArguments(),
    tts_args,
    queues,
)

pipeline_manager.start()
pipeline_manager.wait()  # Blocks until interrupted

```

## Configure Audio I/O Modes

The pipeline supports three operational modes controlled by `module_args.mode`:

- **Local mode** (`"local"`): Direct microphone input and speaker output. Best for standalone applications.
- **WebSocket mode** (`"websocket"`): Audio streams over WebSocket connections. Ideal for web applications.
- **Real-time mode** (`"realtime"`): OpenAI-compatible realtime API integration.

For low-latency requirements, enable live transcription with `module_args.enable_live_transcription=True`.

## Test with the Socket Demo

The repository includes [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py) for testing TCP socket streaming without hardware loopback.

Start the pipeline in WebSocket mode:

```bash
python -m speech_to_speech.s2s_pipeline --mode websocket

```

Then run the demo client:

```bash
python scripts/listen_and_play.py \
    --host localhost \
    --send_port 12345 \
    --recv_port 12346

```

The demo captures raw 16-bit audio, pushes it to the pipeline via TCP, and plays back synthesized responses.

## Summary

- **Four-stage architecture**: The pipeline chains VAD → STT → LLM → TTS through thread-safe queues managed by `ThreadManager`.
- **Two integration paths**: Use `main()` for CLI-driven deployment or `build_pipeline()` for embedded applications.
- **Flexible backends**: Choose from Whisper, Faster-Whisper, or Paraformer for STT; Qwen-3, ChatTTS, or Kokoro for TTS; and Transformers or MLX-LM for language modeling.
- **Multiple modes**: Run locally with direct audio hardware, over WebSockets for web apps, or via the OpenAI-compatible realtime API.
- **Configuration via dataclasses**: All parameters are type-safe arguments defined in `src/speech_to_speech/arguments_classes/`.

## Frequently Asked Questions

### How do I select different STT or TTS backends?

Set the `stt` and `tts` parameters in `ModuleArguments` to your desired backend identifiers (e.g., `"whisper"`, `"faster-whisper"`, `"qwen3"`, `"kokoro"`), then provide the corresponding handler arguments to `prepare_all_args()`. The `build_pipeline()` function instantiates the correct handler classes based on these selections.

### Can I run the pipeline on macOS with Apple Silicon?

Yes. The `prepare_all_args()` function automatically detects macOS and applies optimizations, preferring `mlx-lm` for the LLM backend and `qwen3` for TTS when available. Specify `device="mps"` or leave it on auto-detect for best performance on Apple Silicon.

### What is the difference between local mode and WebSocket mode?

**Local mode** (`mode="local"`) uses direct system audio through local audio streamers, suitable for desktop applications. **WebSocket mode** (`mode="websocket"`) accepts audio over WebSocket connections via `WebSocketStreamerArguments`, enabling remote clients to stream audio to the pipeline over the network.

### How do I enable graceful shutdown in my application?

The `ThreadManager` returned by `build_pipeline()` registers signal handlers for graceful shutdown. Call `pipeline_manager.stop()` or send a KeyboardInterrupt (Ctrl+C) to trigger the shutdown sequence, which properly joins all handler threads and clears the inter-thread queues.