# How to Use the speech-to-speech Python API: A Complete Guide

> Learn how to use the speech-to-speech Python API to build modular voice assistants. This guide covers VAD, STT, LLM, and TTS components with the ThreadManager.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**The speech-to-speech Python API exposes a modular voice-assistant pipeline through the [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) module, allowing developers to programmatically configure VAD, STT, LLM, and TTS components using factory functions and the `ThreadManager` class.**

The `huggingface/speech-to-speech` repository implements a fully modular voice-assistant architecture that can be invoked from Python code or the provided CLI. Unlike monolithic speech systems, this framework separates concerns into four interchangeable handlers that communicate through typed queues, enabling low-latency streaming and easy backend swapping. Whether building a local offline assistant or an OpenAI Realtime-compatible server, the speech-to-speech Python API provides granular control over every pipeline stage.

## Architecture Overview

The system consists of four interchangeable components that run in parallel threads and exchange data through typed queues:

- **Voice-Activity Detection (VAD)** – Detects speech boundaries using Silero VAD and optionally streams live transcription events.
- **Speech-to-Text (STT)** – Converts spoken input into text using backends like Parakeet TDT, Whisper, Faster-Whisper, MLX-Audio-Whisper, or Paraformer.
- **Large Language Model (LLM)** – Generates responses via local models (Transformers, mlx-lm) or OpenAI-compatible APIs (`responses-api` or `chat-completions`).
- **Text-to-Speech (TTS)** – Synthesizes audio using Qwen3-TTS (default), Pocket TTS, ChatTTS, Kokoro-82M, or Facebook MMS.

These handlers communicate through typed queues (`AudioInItem`, `STTOutItem`, `LMOutItem`, `TTSInItem`) and synchronize using `Event` flags (`stop_event`, `should_listen`, `response_playing`). The orchestration logic in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) uses `HfArgumentParser` to parse CLI arguments or JSON configs, normalizes argument prefixes (e.g., `--stt_*`, `--tts_*`) via `rename_args`, and builds concrete handler instances through factory functions.

## Transport Modes

The pipeline supports four transport modes controlled via the `--mode` argument:

| Mode | Transport | Use Case |
|------|-----------|----------|
| `realtime` | OpenAI Realtime API over WebSocket/WebRTC | Build voice assistants compatible with OpenAI clients |
| `local` | Direct microphone & speakers | Quick prototyping on a single machine |
| `raw-websocket` | Raw PCM over WebSocket | Custom clients streaming 16 kHz int16 PCM |
| `socket` | Raw PCM over TCP | Simple LAN streaming without Realtime features |

## Programmatic API Usage

To start the pipeline programmatically instead of via CLI, import the core functions from [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) and manage the lifecycle through `ThreadManager`:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse CLI arguments or construct dataclasses manually

args = parse_arguments()

# Normalize argument prefixes and apply device defaults

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Initialize shared queues and synchronization primitives

queues_and_events = initialize_queues_and_events()

# Build handlers via factory functions (get_stt_handler, get_llm_handler, get_tts_handler)

pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues_and_events,
)

# Start and block until SIGINT/SIGTERM

pipeline_manager.start()
pipeline_manager.wait()

```

The `build_pipeline` function internally uses factory functions (`get_stt_handler`, `get_llm_handler`, `get_tts_handler`) to instantiate concrete handlers based on your configuration, then wires them together under a `ThreadManager` that handles graceful shutdown.

## CLI Configuration Examples

While the Python API provides full control, you can also leverage the CLI entry point which calls `main()` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py):

**Start a realtime server** (default VAD → Parakeet TDT → Responses-API → Qwen3-TTS):

```bash
export OPENAI_API_KEY=sk-...
speech-to-speech

```

**Swap STT to Whisper**:

```bash
speech-to-speech --stt whisper --stt_model_name openai/whisper-base

```

**Run locally with MLX LLM on Apple Silicon**:

```bash
speech-to-speech --mode local --llm_backend mlx-lm --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16

```

**Use Pocket TTS**:

```bash
speech-to-speech --tts pocket --pocket_tts_voice jean --pocket_tts_device cpu

```

Enable the **LLM proxy** with `--enable_llm_proxy` to expose the configured LLM as a stand-alone OpenAI-compatible endpoint, letting other services call the model while the voice pipeline runs uninterrupted.

## Key Source Files

Understanding the repository structure helps when extending the API:

- **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)** – Core orchestration containing `parse_arguments()`, `prepare_all_args()`, `initialize_queues_and_events()`, and `build_pipeline()`.
- **[`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py)** – Implements `ThreadManager` for handler lifecycle management.
- **`src/speech_to_speech/arguments_classes/*.py`** – Dataclasses defining CLI/JSON arguments for each component.
- **`src/speech_to_speech/STT/*_handler.py`** – Concrete STT implementations (Whisper, Parakeet TDT, etc.).
- **`src/speech_to_speech/LLM/*_handler.py`** – LLM backends including local and API variants.
- **`src/speech_to_speech/TTS/*_handler.py`** – TTS synthesizers (Qwen3-TTS, Kokoro, Pocket, etc.).
- **`src/speech_to_speech/pipeline/*.py`** – Queue definitions and speculative turn tracking.

## Summary

- The **speech-to-speech Python API** in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) provides a modular interface for building voice assistants through four core handlers: VAD, STT, LLM, and TTS.
- **Factory functions** (`get_stt_handler`, `get_llm_handler`, `get_tts_handler`) instantiate backends based on configuration dataclasses, while `ThreadManager` coordinates parallel execution.
- **Typed queues** (`AudioInItem`, `STTOutItem`, `LMOutItem`, `TTSInItem`) and **Event flags** (`stop_event`, `should_listen`) enable safe thread communication and graceful shutdown.
- The API supports **four transport modes** (`realtime`, `local`, `raw-websocket`, `socket`) and can expose an **LLM proxy** for external API access.
- Configuration occurs through `HfArgumentParser` with prefixed arguments (`--stt_*`, `--tts_*`) normalized by `prepare_all_args()`.

## Frequently Asked Questions

### How do I switch between different STT backends in the Python API?

Pass the specific handler kwargs to `prepare_all_args()` and `build_pipeline()`. The factory function `get_stt_handler` selects the concrete implementation based on the `stt` argument value (e.g., `"whisper"`, `"parakeet_tdt"`, `"faster_whisper"`). For example, to use Whisper, ensure `args.whisper_stt_handler_kwargs` contains `model_name="openai/whisper-base"` and the STT selector is set to `"whisper"`.

### What is the difference between `realtime` and `local` transport modes?

The `realtime` mode exposes an OpenAI Realtime API-compatible WebSocket endpoint at `ws://localhost:8765/v1/realtime`, allowing external clients to connect using standard OpenAI SDKs. The `local` mode bypasses network transport and connects directly to system microphone and speakers via the local audio handler, making it ideal for single-machine prototyping without network overhead.

### How does the pipeline handle graceful shutdown?

The `ThreadManager` class (in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py)) monitors a `stop_event` threading.Event shared across all handlers. When the process receives SIGINT or SIGTERM, the event is set, causing each handler's `run()` loop to exit cleanly. The main thread then calls `join()` on each handler thread via `pipeline_manager.wait()` to ensure all audio buffers are flushed before termination.

### Can I use the speech-to-speech pipeline with a custom LLM server?

Yes. Set `--enable_llm_proxy` to expose the configured LLM as an OpenAI-compatible HTTP endpoint, or use `--llm_backend` with an OpenAI-compatible base URL. The `LanguageModelHandler` supports both `responses-api` and `chat-completions` formats, allowing you to proxy requests to custom servers while maintaining the voice pipeline's streaming capabilities.