# Python APIs for Speech-to-Speech: Building Low-Latency Pipelines with Hugging Face

> Explore Python APIs for speech-to-speech with Hugging Face. Build low-latency VAD, STT, LLM, and TTS pipelines entirely in Python using the speech-to-speech repository.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-07

---

**The Hugging Face `speech-to-speech` repository provides a comprehensive Python API for constructing modular, low-latency speech-to-speech pipelines via the `s2s_pipeline` module, which exposes three core functions—`parse_arguments()`, `initialize_queues_and_events()`, and `build_pipeline()`—to orchestrate VAD → STT → LLM → TTS workflows entirely in Python.**

The `speech-to-speech` repository from Hugging Face offers a fully-featured Python API that enables developers to build real-time speech-to-speech systems without leaving the Python ecosystem. Unlike command-line-only tools, this API allows you to programmatically configure, initialize, and run complete audio processing pipelines using thread-safe handlers and queue-based communication. As implemented in the repository, the API supports both local audio I/O and network-based streaming through a unified interface defined in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

## Core Python API Components

The primary entry point for the Python API is the `speech_to_speech.s2s_pipeline` module. According to the source code in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), this module exposes three high-level functions that handle the complete lifecycle of a speech-to-speech pipeline:

- **`parse_arguments()`** (line 29): Parses command-line arguments or JSON configuration files into a typed `ParsedArguments` dataclass.
- **`initialize_queues_and_events()`** (line 33): Creates thread-safe `queue.Queue` subclasses and synchronization events that connect each processing stage.
- **`build_pipeline()`** (line 81): Constructs the full pipeline chain (VAD → STT → LLM → TTS) and returns a `ThreadManager` instance for execution control.

These functions work sequentially to transform configuration parameters into a running pipeline managed by the `ThreadManager` class from [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py).

### Configuration and Argument Parsing

In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the `parse_arguments()` function serves as the configuration gateway. It accepts command-line arguments or JSON configuration files and returns a `ParsedArguments` dataclass containing all necessary parameters for the pipeline stages. This includes model selections (e.g., `--stt whisper`, `--tts qwen3`), device configurations (CPU, CUDA, or Apple Silicon), and backend-specific options for each handler.

### Queue Initialization and Thread Safety

The `initialize_queues_and_events()` function (line 33) establishes the inter-process communication infrastructure. It creates the thread-safe queues and `threading.Event` objects that allow each stage to communicate asynchronously. This architecture ensures that voice activity detection, speech recognition, language model inference, and speech synthesis can run concurrently without blocking, with each handler waiting on its respective input queue and signaling completion via shared events.

### Pipeline Construction

The `build_pipeline()` function (line 81) acts as the factory that instantiates and connects all pipeline components. It accepts keyword arguments for every supported handler—including `vad_handler`, `whisper_stt_handler`, `language_model_handler`, and `qwen3_tts_handler`—and wires them together according to the architecture: VADHandler → STTHandler → TranscriptionNotifier → LLMHandler → LMOutputProcessor → TTSHandler. The function returns a `ThreadManager` that provides `start()`, `stop()`, and `wait()` methods for lifecycle management.

## Modular Pipeline Architecture

The Python API implements a handler-based architecture where each processing stage is encapsulated in a specialized handler class. All handlers inherit from the abstract base class defined in [`src/speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/baseHandler.py) and communicate via the queues initialized earlier. You can swap implementations by changing the arguments passed to `build_pipeline()`.

### Voice Activity Detection (VAD)

The pipeline begins with voice activity detection handled by the `VADHandler` class in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py). This component monitors audio input and triggers speech processing only when voice activity is detected, reducing computational load and latency for silent periods.

### Speech-to-Text (STT) Backends

The STT stage supports multiple backends through handler classes in `src/speech_to_speech/STT/`:

- **`WhisperSTTHandler`**: Uses OpenAI Whisper models (via [`whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_handler.py))
- **`FasterWhisperSTTHandler`**: Optimized Whisper implementation via faster-whisper
- **`ParaformerSTTHandler`**: Alibaba Paraformer support
- **`ParakeetTDTSTTHandler`**: NVIDIA Parakeet TDT implementation
- **`MLXAudioWhisperSTTHandler`**: Apple Silicon optimized Whisper via mlx-audio

Each handler implements the same interface and can be selected via the `stt` argument in the configuration.

### Language Model (LLM) Integration

The LLM stage in `src/speech_to_speech/LLM/` supports multiple inference backends:

- **`LanguageModelHandler`**: Hugging Face Transformers and MLX-LM support (via [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py))
- **`ResponsesAPILanguageModelHandler`**: OpenAI-compatible API integration

These handlers process transcriptions and generate text responses that feed into the TTS stage.

### Text-to-Speech (TTS) Synthesis

The final stage converts text to speech using handlers in `src/speech_to_speech/TTS/`:

- **`Qwen3TTSHandler`**: Supports both `mlx-audio` and `faster-qwen3-tts` backends (detailed in [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py))
- **`ChatTTSHandler`**: ChatTTS integration
- **`FacebookMMSTTSHandler`**: Facebook MMS TTS models
- **`PocketTTSHandler`**: Pocket TTS lightweight implementation
- **`KokoroTTSHandler`**: Kokoro TTS support

The `Qwen3TTSHandler` specifically supports the `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice` model and outputs NumPy int16 arrays suitable for direct audio playback or WAV file writing.

## Practical Implementation Examples

The following examples demonstrate how to use the Python APIs for speech-to-speech in different scenarios, from simple command-line execution to embedded server deployment.

### Command-Line Interface Usage

For quick testing or scripting, you can invoke the pipeline through the `main()` function:

```python

# example.py

from speech_to_speech.s2s_pipeline import main

if __name__ == "__main__":
    main()

```

Run from the terminal with specific backend selections:

```bash
python example.py \
    --mode local \
    --stt whisper \
    --tts qwen3 \
    --llm_backend transformers \
    --device cpu

```

This approach uses `parse_arguments()` internally and supports all command-line options defined in the repository.

### Programmatic Pipeline Control

For integration into existing Python applications, instantiate the pipeline components directly:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse configuration (can construct args manually instead of CLI)

args = parse_arguments()

# Initialize communication infrastructure

queues_and_events = initialize_queues_and_events()

# Build the complete pipeline

pipeline_manager = build_pipeline(
    module_kwargs=args.module_kwargs,
    socket_receiver_kwargs=args.socket_receiver_kwargs,
    socket_sender_kwargs=args.socket_sender_kwargs,
    websocket_streamer_kwargs=args.websocket_streamer_kwargs,
    vad_handler_kwargs=args.vad_handler_kwargs,
    whisper_stt_handler_kwargs=args.whisper_stt_handler_kwargs,
    faster_whisper_stt_handler_kwargs=args.faster_whisper_stt_handler_kwargs,
    paraformer_stt_handler_kwargs=args.paraformer_stt_handler_kwargs,
    mlx_audio_whisper_stt_handler_kwargs=args.mlx_audio_whisper_stt_handler_kwargs,
    parakeet_tdt_stt_handler_kwargs=args.parakeet_tdt_stt_handler_kwargs,
    language_model_handler_kwargs=args.language_model_handler_kwargs,
    responses_api_language_model_handler_kwargs=args.responses_api_language_model_handler_kwargs,
    chat_tts_handler_kwargs=args.chat_tts_handler_kwargs,
    facebook_mms_tts_handler_kwargs=args.facebook_mms_tts_handler_kwargs,
    pocket_tts_handler_kwargs=args.pocket_tts_handler_kwargs,
    kokoro_tts_handler_kwargs=args.kokoro_tts_handler_kwargs,
    qwen3_tts_handler_kwargs=args.qwen3_tts_handler_kwargs,
    queues_and_events=queues_and_events,
)

# Execute the pipeline

pipeline_manager.start()
pipeline_manager.wait()  # Blocks until shutdown (Ctrl-C)

```

### Direct Handler Access

For specialized use cases, instantiate individual handlers directly. This example uses the Qwen-3 TTS handler:

```python
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from threading import Event

# Setup dummy input structure

class TTSInput:
    def __init__(self, text):
        self.text = text
        self.language_code = None
        self.turn_id = "turn0"
        self.turn_revision = 0
        self.runtime_config = None
        self.response = None

# Instantiate handler

handler = Qwen3TTSHandler()
handler.setup(
    should_listen=Event(),
    model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    device="cpu",
)

# Process text to audio

tts_input = TTSInput("Hello, this demonstrates the Qwen-3 TTS handler.")
for audio_chunk in handler.process(tts_input):
    # audio_chunk is NumPy int16 array

    print(f"Generated {len(audio_chunk)} samples")
    # Save to file: sf.write("output.wav", audio_chunk, 16000, subtype="PCM_16")

```

### OpenAI Realtime Server Deployment

Deploy the pipeline as an OpenAI-compatible realtime server using the WebSocket implementation in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py):

```python
from speech_to_speech.api.openai_realtime.server import RealtimeServer
from speech_to_speech.s2s_pipeline import (
    parse_arguments, 
    initialize_queues_and_events, 
    build_pipeline
)

args = parse_arguments()
queues = initialize_queues_and_events()
pipeline_manager = build_pipeline(
    # ... include all required kwargs as shown above ...

    queues_and_events=queues,
)

# Start the realtime server embedded in the pipeline

pipeline_manager.start()
pipeline_manager.wait()

```

When running with `--mode realtime`, the pipeline exposes the OpenAI Realtime protocol over WebSocket, allowing client applications to stream audio in and receive synthesized speech out.

## Summary

The Hugging Face `speech-to-speech` repository provides a production-ready Python API for building speech-to-speech pipelines:

- **Three core functions** (`parse_arguments()`, `initialize_queues_and_events()`, `build_pipeline()`) in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) provide the high-level interface for pipeline construction.
- **Modular handler architecture** supports interchangeable VAD, STT, LLM, and TTS components via classes in `src/speech_to_speech/VAD/`, `STT/`, `LLM/`, and `TTS/` directories.
- **Thread-safe execution** is managed by the `ThreadManager` class in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py), coordinating handlers through queue-based communication.
- **Multiple backends** are supported including Whisper, Faster-Whisper, Qwen-3 TTS, ChatTTS, and MLX-Audio for Apple Silicon optimization.
- **OpenAI Realtime compatibility** is available through the server implementation in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py).

## Frequently Asked Questions

### What Python version is required for the speech-to-speech API?

The repository requires Python 3.9 or higher due to its use of modern typing features and asynchronous programming patterns. The API relies heavily on `threading` and `queue` modules from the standard library, along with specific versions of PyTorch, Transformers, and optional dependencies like `mlx-audio` for Apple Silicon or `faster-whisper` for optimized inference.

### Can I use individual handlers without the full pipeline?

Yes, you can import and instantiate individual handlers directly from their respective modules. For example, import `Qwen3TTSHandler` from [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) or `WhisperSTTHandler` from [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py). Each handler implements a standard interface with `setup()` and `process()` methods, allowing you to use them in isolation or integrate them into custom pipeline architectures outside of the provided `ThreadManager`.

### How do I switch between different STT or TTS backends?

You specify the backend through the arguments passed to `build_pipeline()` or via command-line flags when using `parse_arguments()`. For STT, use arguments like `--stt whisper`, `--stt faster_whisper`, or `--stt paraformer`, which correspond to the `whisper_stt_handler_kwargs`, `faster_whisper_stt_handler_kwargs`, and other parameters. For TTS, options include `--tts qwen3`, `--tts chat_tts`, or `--tts kokoro`, which route to their respective handler keyword arguments in the build function.

### Is the pipeline suitable for real-time production deployment?

Yes, the architecture is designed for low-latency real-time processing. The queue-based communication system in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) uses `queue.Queue` subclasses that enable concurrent processing across threads without blocking. For production web deployment, use the `RealtimeServer` class in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) with `--mode realtime`, which exposes the pipeline via WebSocket following the OpenAI Realtime API specification. The `ThreadManager` provides robust lifecycle management with proper start/stop semantics for long-running services.