# How to Deploy a Speech-to-Speech Model with the Hugging Face Speech-to-Speech Repository

> Deploy a speech-to-speech model easily with the Hugging Face repository. Select your backends and run the OpenAI Realtime-compatible WebSocket server for seamless integration.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-02

---

**Deploy a speech-to-speech model by installing the `speech-to-speech` package, selecting your VAD, STT, LLM, and TTS backends, and running the OpenAI Realtime-compatible WebSocket server.**

This guide covers deploying a modular, low-latency voice-agent pipeline from the `huggingface/speech-to-speech` repository. The system chains Voice Activity Detection (VAD) → Speech-to-Text (STT) → Large Language Model (LLM) → Text-to-Speech (TTS) into a production-ready WebSocket service.

## Installation and Component Selection

The first step to deploy a speech-to-speech model is installing the package and choosing your pipeline components.

### Install the Library

```bash
pip install speech-to-speech

```

Optional extras for alternative backends:

```bash
pip install "speech-to-speech[pocket]"         # Pocket TTS

pip install "speech-to-speech[chattts]"        # ChatTTS

pip install "speech-to-speech[faster-whisper]" # Faster Whisper STT

```

### Select Your Backends

The CLI in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) accepts flags to swap any pipeline stage. Default components are:

- **VAD**: Silero VAD
- **STT**: Parakeet TDT
- **LLM**: OpenAI-compatible Responses API
- **TTS**: Qwen3-TTS

These flags are defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) and specialized files like [`stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/stt_arguments.py), [`llm_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/llm_arguments.py), and [`tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/tts_arguments.py).

## Running the Speech-to-Speech Server

### Basic Server Launch

Start the realtime WebSocket server with a single command:

```bash
speech-to-speech

```

The server binds to `ws://0.0.0.0:8765/v1/realtime` by default.

### Customized Deployment Example

For a fully local speech-to-speech deployment:

```bash
speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-5.4-mini" \
    --responses_api_stream \
    --enable_live_transcription

```

Flag explanations:

- `--mode realtime` — Activates the OpenAI-Realtime WebSocket API
- `--stt parakeet-tdt` — Selects Parakeet TDT for speech-to-text
- `--llm_backend responses-api` — Uses an OpenAI-compatible LLM endpoint
- `--tts qwen3` — Enables Qwen3-TTS for speech synthesis
- `--responses_api_stream` — Streams LLM responses for lower latency
- `--enable_live_transcription` — Returns transcription events to the client

## Server Architecture and Pipeline Units

When you deploy a speech-to-speech model, [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) creates a `RealtimeServer` that:

1. Launches a uvicorn-powered FastAPI application
2. Manages a pool of `PipelineUnit` objects from [`src/speech_to_speech/api/openai_realtime/pipeline_unit.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/pipeline_unit.py)

Each `PipelineUnit` encapsulates an isolated **VAD → STT → LLM → TTS thread chain**, ensuring concurrent client connections do not interfere with each other.

The WebSocket routing layer in [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py) parses incoming JSON events and dispatches them to per-connection `RealtimeService` instances.

## Docker Deployment for Production

The repository includes a [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) for containerized deployment:

```bash
docker compose up

```

This configuration:
- Pulls a llama.cpp image for local LLM inference
- Builds the speech-to-speech pipeline container
- Runs in socket mode for inter-container communication

## Connecting a Client to Your Deployed Model

Any OpenAI-Realtime-compatible client can connect to your deployed speech-to-speech model.

### Python SDK Example

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed"  # Server does not enforce API keys

)

with client.realtime.connect(model="local") as conn:
    # Configure session

    conn.send({
        "type": "session.update",
        "session": {
            "type": "realtime",
            "instructions": "You are a helpful assistant.",
            "audio": {"input": {"turn_detection": {"type": "server_vad", "interrupt_response": True}}}
        },
    })

    # Stream audio and receive events

    for event in conn:
        print(event.type)  # "response.output_audio.delta", "conversation.item.created", etc.

```

### Reference Client Implementation

The repository provides [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py), a complete reference client that:

- Captures microphone audio
- Encodes to 16 kHz PCM
- Streams to the WebSocket endpoint
- Plays back synthesized speech responses

## Handling User Interruptions with CancelScope

A critical feature when you deploy a speech-to-speech model is **barge-in support**—allowing users to interrupt ongoing generation.

The [`src/speech_to_speech/pipeline/cancel_scope.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/cancel_scope.py) module implements generation-aware cancellation:

- Tracks active LLM and TTS generation scopes
- Gracefully aborts in-flight requests when new user speech is detected
- Prevents audio artifacts and partial responses

This mechanism is integrated into the `PipelineUnit` thread chain through `TranscriptionNotifier` and `LMOutputProcessor` stages.

## Key Source Files for Deployment

| File | Purpose |
|------|---------|
| [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) | Entry point, argument parsing, mode selection |
| [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) | FastAPI/uvicorn server, `PipelineUnit` pool management |
| [`src/speech_to_speech/api/openai_realtime/pipeline_unit.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/pipeline_unit.py) | Single VAD-STT-LLM-TTS thread chain encapsulation |
| [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py) | WebSocket event routing and `RealtimeService` dispatch |
| [`src/speech_to_speech/pipeline/cancel_scope.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/cancel_scope.py) | Generation-aware cancellation for interruption handling |
| [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) | Container orchestration with local LLM |
| [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) | Production-ready reference client |

## Protocol and Audio Specifications

When deploying a speech-to-speech model, clients must adhere to:

- **Transport**: WebSocket at `/v1/realtime`
- **Input audio**: 16 kHz PCM, 16-bit, mono
- **Output audio**: 16 kHz PCM (synthesized speech)
- **Event format**: OpenAI Realtime API compatible JSON

The `RealtimeService` in the pipeline handles resampling and format conversion automatically.

## Summary

- **Install** with `pip install speech-to-speech` and optional extras for alternate backends
- **Configure** components via CLI flags defined in `arguments_classes/` modules
- **Launch** the WebSocket server with `speech-to-speech` or Docker Compose
- **Scale** through `PipelineUnit` pools, where each unit runs an isolated VAD-STT-LLM-TTS chain
- **Connect** any OpenAI-Realtime client to `ws://localhost:8765/v1/realtime`
- **Handle interruptions** via `CancelScope` for production-grade barge-in support

## Frequently Asked Questions

### What hardware requirements are needed to deploy a speech-to-speech model?

GPU acceleration is recommended but not required. The default Parakeet TDT STT and Qwen3 TTS models run efficiently on modern NVIDIA GPUs with CUDA support. CPU-only deployment is possible with slower inference speeds. The LLM backend can be offloaded to external APIs (OpenAI, Together, etc.) or run locally via llama.cpp depending on your latency and privacy requirements.

### Can I replace individual pipeline components without modifying source code?

Yes. The CLI flags `--stt`, `--llm_backend`, and `--tts` accept string identifiers that map to handler classes in `src/speech_to_speech/handlers/`. Adding a custom backend requires: (1) implementing the handler interface, (2) registering it in [`src/speech_to_speech/handlers/__init__.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/handlers/__init__.py), and (3) referencing it by name in your deployment command. No changes to core pipeline logic are needed.

### How does the server handle multiple concurrent connections?

The `RealtimeServer` in [`server.py`](https://github.com/huggingface/speech-to-speech/blob/main/server.py) maintains a pool of `PipelineUnit` instances. When a WebSocket client connects, [`websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_router.py) attaches it to an available unit. Each unit's isolated thread chain ensures one client's audio processing never blocks another. Pool sizing defaults are tuned for typical GPU memory constraints but can be adjusted via CLI arguments.

### Is the deployed server compatible with OpenAI's official Realtime API?

The WebSocket protocol and event schema follow OpenAI Realtime API conventions, making clients like the official Python SDK interchangeable. Minor differences exist in authentication (the server ignores API keys) and available model identifiers (use `"local"` or custom strings). The compatibility target is documented in [`src/speech_to_speech/api/openai_realtime/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/README.md).