What Is the Hugging Face speech-to-speech Repository? A Complete Guide to the Modular Voice Agent Pipeline

The huggingface speech-to-speech repository provides a low-latency, fully modular voice-agent pipeline that connects Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS) stages via thread-safe queues, exposing an OpenAI Realtime-compatible WebSocket server.

The huggingface speech-to-speech framework enables developers to build real-time conversational AI systems using interchangeable backends. Each processing stage runs in an independent thread, communicating through asyncio queues to minimize latency between speech input and audio output. The system automatically selects optimal CPU, GPU, or Apple Silicon backends based on the host environment.

Architecture Overview: The Four-Stage Pipeline

The repository implements a streaming pipeline that processes audio through four distinct stages orchestrated by src/speech_to_speech/s2s_pipeline.py. Each stage pushes results to the next via thread-safe queues, allowing parallel processing and sub-second response times.

  • Voice Activity Detection (VAD): Detects speech boundaries and manages turn-taking using Silero VAD v5 by default.
  • Speech-to-Text (STT): Transcribes user audio with optional partial streaming support.
  • Language Model (LLM): Generates textual responses with streaming output and tool-calling capabilities.
  • Text-to-Speech (TTS): Synthesizes spoken output from generated text.

The pipeline runs in four independent threads, ensuring that audio capture, transcription, inference, and synthesis occur concurrently without blocking.

Core Components and Swappable Backends

Each stage supports multiple backend implementations configurable via command-line flags defined in src/speech_to_speech/arguments_classes/module_arguments.py.

Voice Activity Detection (VAD)

The default Silero VAD v5 implementation in src/speech_to_speech/VAD/vad_handler.py handles speech boundary detection and manages conversation turn-taking. This component determines when the user has finished speaking and triggers the transcription phase.

Speech-to-Text (STT)

By default, the system uses Parakeet TDT (implemented in src/speech_to_speech/STT/parakeet_tdt_handler.py) for fast transcription with live streaming support. Alternative backends include:

  • Whisper (OpenAI and Faster Whisper variants)
  • MLX Whisper (optimized for Apple Silicon)
  • Paraformer

Language Model (LLM)

The default implementation in src/speech_to_speech/LLM/responses_api_language_model.py connects to OpenAI-compatible Responses APIs. Supported alternatives include:

  • Transformers (Hugging Face models)
  • mlx-lm (Apple Silicon optimization)
  • Self-hosted vLLM or llama.cpp
  • HF Inference Providers

Text-to-Speech (TTS)

Qwen-3 TTS serves as the default backend in src/speech_to_speech/TTS/qwen3_tts_handler.py, supporting both GGML and MLX inference. Additional options include:

  • Kokoro-82M
  • Pocket TTS
  • ChatTTS
  • Facebook MMS

Installation and Quick Start

Install the package via pip and launch the realtime server with default configuration:

pip install speech-to-speech
export OPENAI_API_KEY=...
speech-to-speech

This command starts a WebSocket server at ws://localhost:8765/v1/realtime using the default backend stack (Silero VAD → Parakeet TDT → OpenAI Responses API → Qwen-3 TTS).

OpenAI Realtime API Compatibility

The repository exposes an OpenAI Realtime-compatible WebSocket endpoint at /v1/realtime, implemented in src/speech_to_speech/api/openai_realtime/server.py. This allows any OpenAI Realtime client to connect without modification.

Connect using the official Python client:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed",
)

with client.realtime.connect(model="local") as conn:
    conn.send(
        {
            "type": "session.update",
            "session": {
                "type": "realtime",
                "instructions": "You are a helpful assistant.",
                "audio": {"input": {"turn_detection": {"type": "server_vad", "interrupt_response": True}}},
            },
        }
    )
    for event in conn:
        print(event.type)

The server also supports plain WebSocket connections for raw PCM audio and TCP sockets for minimal custom clients.

Customizing the Pipeline with Alternative Backends

Mix and match components using CLI flags to integrate local models or specific hardware optimizations. For example, to use a local llama.cpp server as the LLM backend:

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "ggml-org/gemma-4-E4B-it-GGUF" \
    --responses_api_base_url "http://127.0.0.1:8080/v1" \
    --responses_api_api_key "" \
    --responses_api_stream \
    --enable_live_transcription

This configuration maintains the default VAD, STT, and TTS components while routing LLM inference to a local endpoint.

Key Source Files and Implementation Details

Understanding the codebase structure helps when extending functionality:

Summary

  • The huggingface speech-to-speech repository provides a modular, four-stage voice agent pipeline (VAD → STT → LLM → TTS) with thread-safe queue communication.
  • Each stage supports multiple interchangeable backends (Silero, Whisper, Parakeet, Qwen-3, Kokoro, etc.) selectable via CLI flags.
  • The system exposes an OpenAI Realtime-compatible WebSocket API at /v1/realtime, enabling drop-in compatibility with existing clients.
  • Four independent threads process audio concurrently, achieving low-latency suitable for real-time conversation.
  • The codebase is platform-agnostic, automatically optimizing for CPU, CUDA, or Apple Silicon hardware.

Frequently Asked Questions

What hardware platforms does huggingface speech-to-speech support?

The framework automatically detects the host environment and selects appropriate backends for CPU, NVIDIA GPU (CUDA), or Apple Silicon (MLX). The module_arguments.py file handles device selection, and specific handlers like parakeet_tdt_handler.py and qwen3_tts_handler.py include optimized paths for GPU and MLX acceleration.

Can I run the pipeline without a WebSocket server?

Yes. The repository supports local mode for direct microphone-to-speaker interaction without requiring a client connection. Use the --mode local flag to run the pipeline interactively on your machine, processing audio directly from system input devices.

How does this compare to using OpenAI's official Realtime API?

While OpenAI's Realtime API is a hosted service, the huggingface speech-to-speech repository allows you to self-host the entire pipeline with custom models. You retain control over data privacy, can swap individual components (e.g., using a local Llama model instead of GPT-4), and avoid per-token API costs by running open-source models locally or on your own infrastructure.

Is Docker deployment supported?

Yes. The repository includes Docker Compose configuration for end-to-end local deployment. This encapsulates the entire pipeline—including optional GPU support—making it suitable for production deployments or development environments where dependency isolation is required.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →