How to Deploy a Speech-to-Speech Model with the Hugging Face Speech-to-Speech Repository

Deploy a speech-to-speech model by installing the speech-to-speech package, selecting your VAD, STT, LLM, and TTS backends, and running the OpenAI Realtime-compatible WebSocket server.

This guide covers deploying a modular, low-latency voice-agent pipeline from the huggingface/speech-to-speech repository. The system chains Voice Activity Detection (VAD) → Speech-to-Text (STT) → Large Language Model (LLM) → Text-to-Speech (TTS) into a production-ready WebSocket service.

Installation and Component Selection

The first step to deploy a speech-to-speech model is installing the package and choosing your pipeline components.

Install the Library

pip install speech-to-speech

Optional extras for alternative backends:

pip install "speech-to-speech[pocket]"         # Pocket TTS

pip install "speech-to-speech[chattts]"        # ChatTTS

pip install "speech-to-speech[faster-whisper]" # Faster Whisper STT

Select Your Backends

The CLI in src/speech_to_speech/s2s_pipeline.py accepts flags to swap any pipeline stage. Default components are:

  • VAD: Silero VAD
  • STT: Parakeet TDT
  • LLM: OpenAI-compatible Responses API
  • TTS: Qwen3-TTS

These flags are defined in src/speech_to_speech/arguments_classes/module_arguments.py and specialized files like stt_arguments.py, llm_arguments.py, and tts_arguments.py.

Running the Speech-to-Speech Server

Basic Server Launch

Start the realtime WebSocket server with a single command:

speech-to-speech

The server binds to ws://0.0.0.0:8765/v1/realtime by default.

Customized Deployment Example

For a fully local speech-to-speech deployment:

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-5.4-mini" \
    --responses_api_stream \
    --enable_live_transcription

Flag explanations:

  • --mode realtime — Activates the OpenAI-Realtime WebSocket API
  • --stt parakeet-tdt — Selects Parakeet TDT for speech-to-text
  • --llm_backend responses-api — Uses an OpenAI-compatible LLM endpoint
  • --tts qwen3 — Enables Qwen3-TTS for speech synthesis
  • --responses_api_stream — Streams LLM responses for lower latency
  • --enable_live_transcription — Returns transcription events to the client

Server Architecture and Pipeline Units

When you deploy a speech-to-speech model, src/speech_to_speech/api/openai_realtime/server.py creates a RealtimeServer that:

  1. Launches a uvicorn-powered FastAPI application
  2. Manages a pool of PipelineUnit objects from src/speech_to_speech/api/openai_realtime/pipeline_unit.py

Each PipelineUnit encapsulates an isolated VAD → STT → LLM → TTS thread chain, ensuring concurrent client connections do not interfere with each other.

The WebSocket routing layer in src/speech_to_speech/api/openai_realtime/websocket_router.py parses incoming JSON events and dispatches them to per-connection RealtimeService instances.

Docker Deployment for Production

The repository includes a docker-compose.yml for containerized deployment:

docker compose up

This configuration:

  • Pulls a llama.cpp image for local LLM inference
  • Builds the speech-to-speech pipeline container
  • Runs in socket mode for inter-container communication

Connecting a Client to Your Deployed Model

Any OpenAI-Realtime-compatible client can connect to your deployed speech-to-speech model.

Python SDK Example

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed"  # Server does not enforce API keys

)

with client.realtime.connect(model="local") as conn:
    # Configure session

    conn.send({
        "type": "session.update",
        "session": {
            "type": "realtime",
            "instructions": "You are a helpful assistant.",
            "audio": {"input": {"turn_detection": {"type": "server_vad", "interrupt_response": True}}}
        },
    })

    # Stream audio and receive events

    for event in conn:
        print(event.type)  # "response.output_audio.delta", "conversation.item.created", etc.

Reference Client Implementation

The repository provides scripts/listen_and_play_realtime.py, a complete reference client that:

  • Captures microphone audio
  • Encodes to 16 kHz PCM
  • Streams to the WebSocket endpoint
  • Plays back synthesized speech responses

Handling User Interruptions with CancelScope

A critical feature when you deploy a speech-to-speech model is barge-in support—allowing users to interrupt ongoing generation.

The src/speech_to_speech/pipeline/cancel_scope.py module implements generation-aware cancellation:

  • Tracks active LLM and TTS generation scopes
  • Gracefully aborts in-flight requests when new user speech is detected
  • Prevents audio artifacts and partial responses

This mechanism is integrated into the PipelineUnit thread chain through TranscriptionNotifier and LMOutputProcessor stages.

Key Source Files for Deployment

File Purpose
src/speech_to_speech/s2s_pipeline.py Entry point, argument parsing, mode selection
src/speech_to_speech/api/openai_realtime/server.py FastAPI/uvicorn server, PipelineUnit pool management
src/speech_to_speech/api/openai_realtime/pipeline_unit.py Single VAD-STT-LLM-TTS thread chain encapsulation
src/speech_to_speech/api/openai_realtime/websocket_router.py WebSocket event routing and RealtimeService dispatch
src/speech_to_speech/pipeline/cancel_scope.py Generation-aware cancellation for interruption handling
docker-compose.yml Container orchestration with local LLM
scripts/listen_and_play_realtime.py Production-ready reference client

Protocol and Audio Specifications

When deploying a speech-to-speech model, clients must adhere to:

  • Transport: WebSocket at /v1/realtime
  • Input audio: 16 kHz PCM, 16-bit, mono
  • Output audio: 16 kHz PCM (synthesized speech)
  • Event format: OpenAI Realtime API compatible JSON

The RealtimeService in the pipeline handles resampling and format conversion automatically.

Summary

  • Install with pip install speech-to-speech and optional extras for alternate backends
  • Configure components via CLI flags defined in arguments_classes/ modules
  • Launch the WebSocket server with speech-to-speech or Docker Compose
  • Scale through PipelineUnit pools, where each unit runs an isolated VAD-STT-LLM-TTS chain
  • Connect any OpenAI-Realtime client to ws://localhost:8765/v1/realtime
  • Handle interruptions via CancelScope for production-grade barge-in support

Frequently Asked Questions

What hardware requirements are needed to deploy a speech-to-speech model?

GPU acceleration is recommended but not required. The default Parakeet TDT STT and Qwen3 TTS models run efficiently on modern NVIDIA GPUs with CUDA support. CPU-only deployment is possible with slower inference speeds. The LLM backend can be offloaded to external APIs (OpenAI, Together, etc.) or run locally via llama.cpp depending on your latency and privacy requirements.

Can I replace individual pipeline components without modifying source code?

Yes. The CLI flags --stt, --llm_backend, and --tts accept string identifiers that map to handler classes in src/speech_to_speech/handlers/. Adding a custom backend requires: (1) implementing the handler interface, (2) registering it in src/speech_to_speech/handlers/__init__.py, and (3) referencing it by name in your deployment command. No changes to core pipeline logic are needed.

How does the server handle multiple concurrent connections?

The RealtimeServer in server.py maintains a pool of PipelineUnit instances. When a WebSocket client connects, websocket_router.py attaches it to an available unit. Each unit's isolated thread chain ensures one client's audio processing never blocks another. Pool sizing defaults are tuned for typical GPU memory constraints but can be adjusted via CLI arguments.

Is the deployed server compatible with OpenAI's official Realtime API?

The WebSocket protocol and event schema follow OpenAI Realtime API conventions, making clients like the official Python SDK interchangeable. Minor differences exist in authentication (the server ignores API keys) and available model identifiers (use "local" or custom strings). The compatibility target is documented in src/speech_to_speech/api/openai_realtime/README.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →