# How to Integrate Speech-to-Speech into a Python Project: A Complete Developer Guide

> Integrate speech-to-speech into your Python project with Hugging Face. Follow this developer guide to build a four-stage pipeline for seamless voice interaction.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**You can integrate the Hugging Face speech-to-speech pipeline into any Python project by importing the `main()` function from `s2s_pipeline` for CLI usage, or by programmatically constructing the four-stage pipeline (VAD → STT → LLM → TTS) using the `build_pipeline()` function with argument dataclasses.**

The **huggingface/speech-to-speech** repository provides a modular, real-time voice conversation system that processes audio through four distinct stages. Whether you need a local microphone-to-speaker loop or a networked WebSocket service, you can embed this pipeline directly into existing Python applications without managing complex audio threading yourself.

## Understanding the Four-Stage Pipeline Architecture

The repository implements a threaded pipeline where each stage communicates via **queues and events** managed by `initialize_queues_and_events()` in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). Understanding these components helps you configure the integration correctly:

- **VAD (Voice Activity Detection)**: Detects when users start and stop speaking. Implemented in [`speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/VAD/vad_handler.py) and instantiated during pipeline construction.
- **STT (Speech-to-Text)**: Converts audio to text using handlers for Whisper, Faster-Whisper, Paraformer, or MLX-Audio-Whisper. Selected via `module_kwargs.stt` and created in `s2s_pipeline.get_stt_handler`.
- **LLM (Language Model)**: Generates text responses using Transformers, MLX-LM, or OpenAI-compatible APIs. Built via `s2s_pipeline.get_llm_handler`.
- **TTS (Text-to-Speech)**: Synthesizes audio from text using ChatTTS, Facebook MMS, Pocket, Kokoro, or Qwen-3 handlers. Created in `s2s_pipeline.get_tts_handler`.

All handlers run in separate threads coordinated by the `ThreadManager` utility in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py).

## Method 1: CLI Integration (Fastest Setup)

For rapid prototyping or standalone deployment, invoke the pipeline directly from the command line. The entry point uses `HfArgumentParser` to process configuration from JSON files or command-line flags defined in `src/speech_to_speech/arguments_classes/`.

```python
from speech_to_speech.s2s_pipeline import main

if __name__ == "__main__":
    # Reads CLI flags or JSON config automatically

    main()

```

Run directly via terminal:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode local \
    --stt whisper \
    --tts qwen3 \
    --llm_backend transformers \
    --device cpu \
    --log_level info

```

This approach handles all argument normalization, queue initialization, and thread management automatically through the `prepare_all_args()` function.

## Method 2: Programmatic Integration (Embedded Usage)

To integrate speech-to-speech into an existing Python application, manually construct the pipeline using the lower-level API. This gives you fine-grained control over model selection, device placement, and communication modes.

```python
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelHandlerArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSHandlerArguments
from speech_to_speech.arguments_classes.socket_receiver_arguments import SocketReceiverArguments
from speech_to_speech.arguments_classes.socket_sender_arguments import SocketSenderArguments
from speech_to_speech.arguments_classes.websocket_streamer_arguments import WebSocketStreamerArguments
from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments
from speech_to_speech.s2s_pipeline import (
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# 1. Configure module-level settings

module_args = ModuleArguments(
    mode="local",          # Options: "local", "websocket", "realtime"

    stt="whisper",
    tts="qwen3",
    llm_backend="transformers",
    device="cpu",          # or "cuda", "mps" for Apple Silicon

    log_level="info",
)

# 2. Configure individual handler arguments

whisper_args = WhisperSTTHandlerArguments(model_name="openai/whisper-base")
lm_args = LanguageModelHandlerArguments(model_name="Qwen/Qwen3-4B-Instruct-2507")
tts_args = Qwen3TTSHandlerArguments()

# 3. Normalize arguments (applies device optimizations and macOS-specific settings)

prepare_all_args(
    module_args,
    whisper_args,
    WhisperSTTHandlerArguments(),  # Placeholder for paraformer

    WhisperSTTHandlerArguments(),  # Placeholder for faster-whisper

    WhisperSTTHandlerArguments(),  # Placeholder for mlx-audio-whisper

    WhisperSTTHandlerArguments(),  # Placeholder for parakeet-tdt

    lm_args,
    LanguageModelHandlerArguments(),  # Placeholder for responses-api

    tts_args,
    Qwen3TTSHandlerArguments(),  # Placeholder for other TTS handlers

    Qwen3TTSHandlerArguments(),
    Qwen3TTSHandlerArguments(),
    Qwen3TTSHandlerArguments(),
)

# 4. Initialize inter-thread communication queues

queues = initialize_queues_and_events()

# 5. Build the pipeline with all handler configurations

pipeline_manager = build_pipeline(
    module_args,
    socket_receiver_kwargs=SocketReceiverArguments(),
    socket_sender_kwargs=SocketSenderArguments(),
    websocket_streamer_kwargs=WebSocketStreamerArguments(),
    vad_handler_kwargs=VADHandlerArguments(),
    whisper_stt_handler_kwargs=whisper_args,
    faster_whisper_stt_handler_kwargs=FasterWhisperSTTHandlerArguments(),
    paraformer_stt_handler_kwargs=ParaformerSTTHandlerArguments(),
    mlx_audio_whisper_stt_handler_kwargs=MLXAudioWhisperSTTHandlerArguments(),
    parakeet_tdt_stt_handler_kwargs=ParakeetTDTSTTHandlerArguments(),
    language_model_handler_kwargs=lm_args,
    responses_api_language_model_handler_kwargs=ResponsesApiLanguageModelHandlerArguments(),
    chat_tts_handler_kwargs=ChatTTSHandlerArguments(),
    facebook_mms_tts_handler_kwargs=FacebookMMSTTSHandlerArguments(),
    pocket_tts_handler_kwargs=PocketTTSHandlerArguments(),
    kokoro_tts_handler_kwargs=KokoroTTSHandlerArguments(),
    qwen3_tts_handler_kwargs=tts_args,
    queues_and_events=queues,
)

# 6. Start processing (blocks until interrupted)

pipeline_manager.start()
pipeline_manager.wait()

```

**Key implementation details**: The `prepare_all_args()` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) handles device detection and argument validation. On macOS, it automatically optimizes for MLX-based backends when available.

## Method 3: Networked Deployment (WebSocket and Socket Modes)

For client-server architectures, change `module_args.mode` to `"websocket"` or use the socket-based sender/receiver arguments. The repository includes [`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py) as a reference implementation for TCP audio streaming.

Start the pipeline in WebSocket mode:

```bash
python -m speech_to_speech.s2s_pipeline --mode websocket --stt whisper --tts qwen3

```

Then connect a client using the demo script:

```bash
python scripts/listen_and_play.py \
    --host localhost \
    --send_port 12345 \
    --recv_port 12346

```

The demo script captures raw 16-bit audio, transmits it to the pipeline, and plays back synthesized responses. This is useful for testing when direct microphone access isn't available or when building distributed voice applications.

## Configuring Handler Arguments

All configurable parameters reside as dataclasses in `src/speech_to_speech/arguments_classes/`:

- **STT configuration**: [`whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_arguments.py), [`faster_whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/faster_whisper_stt_arguments.py), [`paraformer_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/paraformer_stt_arguments.py)
- **LLM configuration**: [`language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model_arguments.py), [`responses_api_language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/responses_api_language_model_arguments.py)
- **TTS configuration**: [`qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_arguments.py), [`chat_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_tts_arguments.py), [`facebook_mms_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/facebook_mms_tts_arguments.py), [`kokoro_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/kokoro_tts_arguments.py)
- **Infrastructure**: [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) (selects mode and backend), [`vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_arguments.py), [`socket_receiver_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/socket_receiver_arguments.py)

Pass these dataclass instances to `build_pipeline()` to override defaults without modifying source code.

## Summary

- **Import the entry point**: Use `from speech_to_speech.s2s_pipeline import main` for CLI-style integration that handles all setup automatically.
- **Construct manually**: Call `prepare_all_args()`, `initialize_queues_and_events()`, and `build_pipeline()` to embed the four-stage pipeline (VAD → STT → LLM → TTS) directly into your application.
- **Choose your mode**: Select `"local"` for direct audio hardware access, `"websocket"` for browser clients, or `"realtime"` for OpenAI-compatible API streaming.
- **Leverage argument dataclasses**: All configuration lives in `src/speech_to_speech/arguments_classes/` with specific handlers for Whisper, Qwen3, Transformers, and socket communication.
- **Enable low-latency**: Set `module_args.enable_live_transcription=True` for real-time transcription feedback during processing.

## Frequently Asked Questions

### How do I reduce latency when integrating speech-to-speech into my Python application?

Set `module_args.enable_live_transcription=True` to enable live transcription mode, which streams partial results before the VAD detects speech completion. Additionally, specify `device="mps"` on Apple Silicon or `device="cuda"` on NVIDIA GPUs to leverage hardware acceleration in the handlers defined in `src/speech_to_speech/LLM/` and `src/speech_to_speech/TTS/`.

### Can I swap the default Whisper STT handler for Faster-Whisper or Paraformer?

Yes. Change `module_args.stt` to `"faster_whisper"` or `"paraformer"`, then import and pass the corresponding argument dataclass (e.g., `FasterWhisperSTTHandlerArguments` or `ParaformerSTTHandlerArguments`) to the `prepare_all_args()` and `build_pipeline()` functions in your integration code.

### What is the difference between "local", "websocket", and "realtime" modes in ModuleArguments?

**Local mode** reads from your system microphone and outputs to speakers directly. **WebSocket mode** accepts audio via WebSocket connections and returns synthesized audio over the same channel, suitable for web applications. **Realtime mode** uses the OpenAI Realtime API-compatible handler for cloud-based low-latency processing. Configure this in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py).

### How do I gracefully shut down the pipeline when embedding it in a larger application?

The `build_pipeline()` function returns a `ThreadManager` instance (from [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py)) that registers signal handlers for graceful shutdown. Call `pipeline_manager.stop()` from your application's shutdown hook, or rely on the automatic Ctrl-C handling implemented in the `main()` function's `try/finally` block.