# What Is the Hugging Face speech-to-speech Repository? A Complete Guide to the Modular Voice Agent Pipeline

> Explore the huggingface speech-to-speech repository, a modular pipeline for real-time voice agents. Connect VAD, STT, LLM, and TTS with OpenAI Realtime compatibility.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: tutorial
- Published: 2026-07-07

---

**The huggingface speech-to-speech repository provides a low-latency, fully modular voice-agent pipeline that connects Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS) stages via thread-safe queues, exposing an OpenAI Realtime-compatible WebSocket server.**

The huggingface speech-to-speech framework enables developers to build real-time conversational AI systems using interchangeable backends. Each processing stage runs in an independent thread, communicating through `asyncio` queues to minimize latency between speech input and audio output. The system automatically selects optimal CPU, GPU, or Apple Silicon backends based on the host environment.

## Architecture Overview: The Four-Stage Pipeline

The repository implements a streaming pipeline that processes audio through four distinct stages orchestrated by [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). Each stage pushes results to the next via thread-safe queues, allowing parallel processing and sub-second response times.

- **Voice Activity Detection (VAD):** Detects speech boundaries and manages turn-taking using Silero VAD v5 by default.
- **Speech-to-Text (STT):** Transcribes user audio with optional partial streaming support.
- **Language Model (LLM):** Generates textual responses with streaming output and tool-calling capabilities.
- **Text-to-Speech (TTS):** Synthesizes spoken output from generated text.

The pipeline runs in **four independent threads**, ensuring that audio capture, transcription, inference, and synthesis occur concurrently without blocking.

## Core Components and Swappable Backends

Each stage supports multiple backend implementations configurable via command-line flags defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py).

### Voice Activity Detection (VAD)

The default **Silero VAD v5** implementation in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) handles speech boundary detection and manages conversation turn-taking. This component determines when the user has finished speaking and triggers the transcription phase.

### Speech-to-Text (STT)

By default, the system uses **Parakeet TDT** (implemented in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py)) for fast transcription with live streaming support. Alternative backends include:

- **Whisper** (OpenAI and Faster Whisper variants)
- **MLX Whisper** (optimized for Apple Silicon)
- **Paraformer**

### Language Model (LLM)

The default implementation in [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py) connects to OpenAI-compatible Responses APIs. Supported alternatives include:

- **Transformers** (Hugging Face models)
- **mlx-lm** (Apple Silicon optimization)
- **Self-hosted vLLM or llama.cpp**
- **HF Inference Providers**

### Text-to-Speech (TTS)

**Qwen-3 TTS** serves as the default backend in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), supporting both GGML and MLX inference. Additional options include:

- **Kokoro-82M**
- **Pocket TTS**
- **ChatTTS**
- **Facebook MMS**

## Installation and Quick Start

Install the package via pip and launch the realtime server with default configuration:

```bash
pip install speech-to-speech
export OPENAI_API_KEY=...
speech-to-speech

```

This command starts a WebSocket server at `ws://localhost:8765/v1/realtime` using the default backend stack (Silero VAD → Parakeet TDT → OpenAI Responses API → Qwen-3 TTS).

## OpenAI Realtime API Compatibility

The repository exposes an OpenAI Realtime-compatible WebSocket endpoint at `/v1/realtime`, implemented in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py). This allows any OpenAI Realtime client to connect without modification.

Connect using the official Python client:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed",
)

with client.realtime.connect(model="local") as conn:
    conn.send(
        {
            "type": "session.update",
            "session": {
                "type": "realtime",
                "instructions": "You are a helpful assistant.",
                "audio": {"input": {"turn_detection": {"type": "server_vad", "interrupt_response": True}}},
            },
        }
    )
    for event in conn:
        print(event.type)

```

The server also supports plain WebSocket connections for raw PCM audio and TCP sockets for minimal custom clients.

## Customizing the Pipeline with Alternative Backends

Mix and match components using CLI flags to integrate local models or specific hardware optimizations. For example, to use a local **llama.cpp** server as the LLM backend:

```bash
speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "ggml-org/gemma-4-E4B-it-GGUF" \
    --responses_api_base_url "http://127.0.0.1:8080/v1" \
    --responses_api_api_key "" \
    --responses_api_stream \
    --enable_live_transcription

```

This configuration maintains the default VAD, STT, and TTS components while routing LLM inference to a local endpoint.

## Key Source Files and Implementation Details

Understanding the codebase structure helps when extending functionality:

- **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)**: Core orchestrator that initializes the four-stage pipeline and launches the selected server mode (local, WebSocket, or TCP).
- **[`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py)**: Implements the WebSocket server handling the OpenAI Realtime protocol, including session management and audio streaming.
- **[`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py)**: Central argument parser exposing all configuration options for devices, model selections, and server parameters.
- **[`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py)**: Manages voice activity detection logic and queue coordination for speech segments.
- **[`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py)**: Default speech recognition implementation with support for live transcription streaming.
- **[`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)**: Text-to-speech handler supporting GGML and MLX inference backends.
- **[`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py)**: Wrapper for OpenAI-compatible chat completion and responses APIs.

## Summary

- The huggingface speech-to-speech repository provides a **modular, four-stage voice agent pipeline** (VAD → STT → LLM → TTS) with thread-safe queue communication.
- Each stage supports **multiple interchangeable backends** (Silero, Whisper, Parakeet, Qwen-3, Kokoro, etc.) selectable via CLI flags.
- The system exposes an **OpenAI Realtime-compatible WebSocket API** at `/v1/realtime`, enabling drop-in compatibility with existing clients.
- Four independent threads process audio concurrently, achieving **low-latency** suitable for real-time conversation.
- The codebase is **platform-agnostic**, automatically optimizing for CPU, CUDA, or Apple Silicon hardware.

## Frequently Asked Questions

### What hardware platforms does huggingface speech-to-speech support?

The framework automatically detects the host environment and selects appropriate backends for CPU, NVIDIA GPU (CUDA), or Apple Silicon (MLX). The [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) file handles device selection, and specific handlers like [`parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/parakeet_tdt_handler.py) and [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) include optimized paths for GPU and MLX acceleration.

### Can I run the pipeline without a WebSocket server?

Yes. The repository supports **local mode** for direct microphone-to-speaker interaction without requiring a client connection. Use the `--mode local` flag to run the pipeline interactively on your machine, processing audio directly from system input devices.

### How does this compare to using OpenAI's official Realtime API?

While OpenAI's Realtime API is a hosted service, the huggingface speech-to-speech repository allows you to **self-host the entire pipeline** with custom models. You retain control over data privacy, can swap individual components (e.g., using a local Llama model instead of GPT-4), and avoid per-token API costs by running open-source models locally or on your own infrastructure.

### Is Docker deployment supported?

Yes. The repository includes **Docker Compose** configuration for end-to-end local deployment. This encapsulates the entire pipeline—including optional GPU support—making it suitable for production deployments or development environments where dependency isolation is required.