How to Deploy a Speech-to-Speech Model with the Hugging Face Speech-to-Speech Repository
Deploy a speech-to-speech model by installing the speech-to-speech package, selecting your VAD, STT, LLM, and TTS backends, and running the OpenAI Realtime-compatible WebSocket server.
This guide covers deploying a modular, low-latency voice-agent pipeline from the huggingface/speech-to-speech repository. The system chains Voice Activity Detection (VAD) → Speech-to-Text (STT) → Large Language Model (LLM) → Text-to-Speech (TTS) into a production-ready WebSocket service.
Installation and Component Selection
The first step to deploy a speech-to-speech model is installing the package and choosing your pipeline components.
Install the Library
pip install speech-to-speech
Optional extras for alternative backends:
pip install "speech-to-speech[pocket]" # Pocket TTS
pip install "speech-to-speech[chattts]" # ChatTTS
pip install "speech-to-speech[faster-whisper]" # Faster Whisper STT
Select Your Backends
The CLI in src/speech_to_speech/s2s_pipeline.py accepts flags to swap any pipeline stage. Default components are:
- VAD: Silero VAD
- STT: Parakeet TDT
- LLM: OpenAI-compatible Responses API
- TTS: Qwen3-TTS
These flags are defined in src/speech_to_speech/arguments_classes/module_arguments.py and specialized files like stt_arguments.py, llm_arguments.py, and tts_arguments.py.
Running the Speech-to-Speech Server
Basic Server Launch
Start the realtime WebSocket server with a single command:
speech-to-speech
The server binds to ws://0.0.0.0:8765/v1/realtime by default.
Customized Deployment Example
For a fully local speech-to-speech deployment:
speech-to-speech \
--mode realtime \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3 \
--model_name "gpt-5.4-mini" \
--responses_api_stream \
--enable_live_transcription
Flag explanations:
--mode realtime— Activates the OpenAI-Realtime WebSocket API--stt parakeet-tdt— Selects Parakeet TDT for speech-to-text--llm_backend responses-api— Uses an OpenAI-compatible LLM endpoint--tts qwen3— Enables Qwen3-TTS for speech synthesis--responses_api_stream— Streams LLM responses for lower latency--enable_live_transcription— Returns transcription events to the client
Server Architecture and Pipeline Units
When you deploy a speech-to-speech model, src/speech_to_speech/api/openai_realtime/server.py creates a RealtimeServer that:
- Launches a uvicorn-powered FastAPI application
- Manages a pool of
PipelineUnitobjects fromsrc/speech_to_speech/api/openai_realtime/pipeline_unit.py
Each PipelineUnit encapsulates an isolated VAD → STT → LLM → TTS thread chain, ensuring concurrent client connections do not interfere with each other.
The WebSocket routing layer in src/speech_to_speech/api/openai_realtime/websocket_router.py parses incoming JSON events and dispatches them to per-connection RealtimeService instances.
Docker Deployment for Production
The repository includes a docker-compose.yml for containerized deployment:
docker compose up
This configuration:
- Pulls a llama.cpp image for local LLM inference
- Builds the speech-to-speech pipeline container
- Runs in socket mode for inter-container communication
Connecting a Client to Your Deployed Model
Any OpenAI-Realtime-compatible client can connect to your deployed speech-to-speech model.
Python SDK Example
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8765/v1",
websocket_base_url="ws://localhost:8765/v1",
api_key="not-needed" # Server does not enforce API keys
)
with client.realtime.connect(model="local") as conn:
# Configure session
conn.send({
"type": "session.update",
"session": {
"type": "realtime",
"instructions": "You are a helpful assistant.",
"audio": {"input": {"turn_detection": {"type": "server_vad", "interrupt_response": True}}}
},
})
# Stream audio and receive events
for event in conn:
print(event.type) # "response.output_audio.delta", "conversation.item.created", etc.
Reference Client Implementation
The repository provides scripts/listen_and_play_realtime.py, a complete reference client that:
- Captures microphone audio
- Encodes to 16 kHz PCM
- Streams to the WebSocket endpoint
- Plays back synthesized speech responses
Handling User Interruptions with CancelScope
A critical feature when you deploy a speech-to-speech model is barge-in support—allowing users to interrupt ongoing generation.
The src/speech_to_speech/pipeline/cancel_scope.py module implements generation-aware cancellation:
- Tracks active LLM and TTS generation scopes
- Gracefully aborts in-flight requests when new user speech is detected
- Prevents audio artifacts and partial responses
This mechanism is integrated into the PipelineUnit thread chain through TranscriptionNotifier and LMOutputProcessor stages.
Key Source Files for Deployment
| File | Purpose |
|---|---|
src/speech_to_speech/s2s_pipeline.py |
Entry point, argument parsing, mode selection |
src/speech_to_speech/api/openai_realtime/server.py |
FastAPI/uvicorn server, PipelineUnit pool management |
src/speech_to_speech/api/openai_realtime/pipeline_unit.py |
Single VAD-STT-LLM-TTS thread chain encapsulation |
src/speech_to_speech/api/openai_realtime/websocket_router.py |
WebSocket event routing and RealtimeService dispatch |
src/speech_to_speech/pipeline/cancel_scope.py |
Generation-aware cancellation for interruption handling |
docker-compose.yml |
Container orchestration with local LLM |
scripts/listen_and_play_realtime.py |
Production-ready reference client |
Protocol and Audio Specifications
When deploying a speech-to-speech model, clients must adhere to:
- Transport: WebSocket at
/v1/realtime - Input audio: 16 kHz PCM, 16-bit, mono
- Output audio: 16 kHz PCM (synthesized speech)
- Event format: OpenAI Realtime API compatible JSON
The RealtimeService in the pipeline handles resampling and format conversion automatically.
Summary
- Install with
pip install speech-to-speechand optional extras for alternate backends - Configure components via CLI flags defined in
arguments_classes/modules - Launch the WebSocket server with
speech-to-speechor Docker Compose - Scale through
PipelineUnitpools, where each unit runs an isolated VAD-STT-LLM-TTS chain - Connect any OpenAI-Realtime client to
ws://localhost:8765/v1/realtime - Handle interruptions via
CancelScopefor production-grade barge-in support
Frequently Asked Questions
What hardware requirements are needed to deploy a speech-to-speech model?
GPU acceleration is recommended but not required. The default Parakeet TDT STT and Qwen3 TTS models run efficiently on modern NVIDIA GPUs with CUDA support. CPU-only deployment is possible with slower inference speeds. The LLM backend can be offloaded to external APIs (OpenAI, Together, etc.) or run locally via llama.cpp depending on your latency and privacy requirements.
Can I replace individual pipeline components without modifying source code?
Yes. The CLI flags --stt, --llm_backend, and --tts accept string identifiers that map to handler classes in src/speech_to_speech/handlers/. Adding a custom backend requires: (1) implementing the handler interface, (2) registering it in src/speech_to_speech/handlers/__init__.py, and (3) referencing it by name in your deployment command. No changes to core pipeline logic are needed.
How does the server handle multiple concurrent connections?
The RealtimeServer in server.py maintains a pool of PipelineUnit instances. When a WebSocket client connects, websocket_router.py attaches it to an available unit. Each unit's isolated thread chain ensures one client's audio processing never blocks another. Pool sizing defaults are tuned for typical GPU memory constraints but can be adjusted via CLI arguments.
Is the deployed server compatible with OpenAI's official Realtime API?
The WebSocket protocol and event schema follow OpenAI Realtime API conventions, making clients like the official Python SDK interchangeable. Minor differences exist in authentication (the server ignores API keys) and available model identifiers (use "local" or custom strings). The compatibility target is documented in src/speech_to_speech/api/openai_realtime/README.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →