# Deploying Speech-to-Speech in Production: Key Considerations for the Hugging Face Pipeline

> Learn key considerations for speech-to-speech production deployment. Optimize hardware, manage latency with CancelScope, and choose the right deployment mode for your Hugging Face pipeline.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: best-practices
- Published: 2026-07-10

---

**Deploying speech-to-speech in production requires optimizing hardware-accelerated backends, managing low-latency interruption handling through the `CancelScope` mechanism, and selecting the appropriate deployment mode from the Realtime WebSocket API to raw TCP sockets.**

The [huggingface/speech-to-speech](https://github.com/huggingface/speech-to-speech) repository implements a fully modular **VAD → STT → LLM → TTS** pipeline designed for low-latency conversational AI. When transitioning this system from development to production, architects must evaluate quantization strategies, backend compatibility, and concurrency models to ensure reliable performance under load.

## Hardware Optimization and Backend Selection

Production deployments must align compute resources with the specific requirements of each pipeline stage. The repository supports heterogeneous hardware configurations across CUDA, CPU, and Apple Silicon environments.

### CUDA Quantization and Wheel Management

For **Qwen3-TTS**, the default GGML backend automatically selects a CUDA 12 wheel. However, production environments running different driver versions must install matching wheels before the main package to avoid runtime errors.

Install the specific CUDA wheel for your runtime:

```bash
pip install "qwentts-cpp-python==0.3.0+cu130" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu130

```

Configuration for hardware selection resides in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py), which handles device-specific argument parsing for the TTS component.

### Apple Silicon and CPU Fallbacks

On macOS, the system automatically utilizes the `mlx-audio` backend for accelerated inference. For CPU-only deployments, each component supports explicit device selection via CLI flags such as `--pocket_tts_device cpu`, ensuring the pipeline operates within constrained environments.

## Pipeline Architecture and Latency Management

The speech-to-speech system achieves conversational responsiveness through sophisticated turn-taking mechanisms and preemptive generation cancellation.

### Interruptible Turn-Taking with CancelScope

The **Voice Activity Detection (VAD)** runs continuously, emitting `speech_started` and `speech_stopped` events that trigger **interruption handling**. When a user speaks during assistant generation, the `CancelScope` object in [`src/speech_to_speech/pipeline/cancel_scope.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/cancel_scope.py) guarantees that stale generations are discarded and only the current response reaches the client.

This mechanism prevents latency buildup from overlapping speech generations and enables truly interactive experiences through live transcription events emitted while the assistant speaks.

### Thread Management and Concurrency

The pipeline spawns dedicated threads per component and uses queue-driven communication. For multi-user services, the `--num_pipelines` flag allows running several parallel pipelines on a single machine.

Thread lifecycle management is encapsulated in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py), which handles safe shutdown sequences and resource cleanup. This architecture supports horizontal scaling within a single instance before requiring additional container orchestration.

## Deployment Modes and API Compatibility

The repository exposes four distinct deployment modes selectable via the `--mode` parameter in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

### OpenAI Realtime WebSocket API

The **Realtime** mode (default) exposes an OpenAI Realtime-compatible WebSocket endpoint at `/v1/realtime`. This mode speaks the standard OpenAI protocol, emitting events such as `response.output_audio.delta` and `response.done`, making it compatible with existing OpenAI client libraries.

### Alternative Deployment Architectures

- **Local**: Direct microphone and speaker interaction without network overhead
- **WebSocket**: Raw PCM streaming without the OpenAI protocol wrapper
- **Socket**: Minimal TCP PCM streaming for remote server integration

Mode definitions and protocol handlers are located in [`src/speech_to_speech/pipeline/handler_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/handler_types.py).

## Configuration and Observability

Production systems require dynamic configuration updates and comprehensive monitoring capabilities.

### Runtime Configuration and Session Updates

All components consume a shared `RuntimeConfig` model defined in [`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py). This configuration can be updated on-the-fly via `session.update` events, allowing adjustment of turn-detection thresholds without restarting the service.

Tune VAD parameters for your specific acoustic environment:

```bash
speech-to-speech \
  --thresh 0.6 \
  --min_speech_ms 384 \
  --min_speech_continuation_ms 192

```

### Dependency Isolation and Security

Install optional backend dependencies via extras to minimize container size:

```bash
pip install "speech-to-speech[pocket]" "speech-to-speech[kokoro]"

```

API keys are read exclusively from environment variables (`OPENAI_API_KEY`, `HF_TOKEN`), with enforcement logic preventing hard-coded secrets in the repository.

### Monitoring and Logging

The server emits structured OpenAI Realtime events suitable for monitoring agents. Logging routes through the standard library `logging` module and the `rich` console for human-readable output during debugging sessions.

## Production Deployment Examples

### Docker Compose for Reproducible Infrastructure

Deploy the complete stack with GPU isolation and dependency management:

```yaml
services:
  speech-to-speech:
    image: huggingface/speech-to-speech:latest
    command: >
      speech-to-speech
      --mode realtime
      --stt parakeet-tdt
      --llm_backend responses-api
      --responses_api_base_url http://llm:8000/v1
      --tts qwen3
      --enable_live_transcription
    environment:
      - OPENAI_API_KEY=${OPENAI_API_KEY}
    ports:
      - "8765:8765"
  llm:
    image: ghcr.io/vllm-project/vllm:latest
    command: >
      python -m vllm.entrypoints.openai.api_server
      --model ggml-org/gemma-4-E4B-it-GGUF
      --port 8000
    ports:
      - "8000:8000"

```

Execute with `docker compose up -d`. The configuration automatically handles GPU selection via the `NVIDIA_CONTAINER_RUNTIME` driver and isolates the self-hosted LLM from the speech pipeline.

### Client Integration Example

Connect to the Realtime endpoint using the OpenAI client library:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed"
)

with client.realtime.connect(model="local") as conn:
    conn.send({
        "type": "session.update",
        "session": {
            "instructions": "You are a helpful assistant.",
            "audio": {"input": {"turn_detection": {"type": "server_vad"}}}
        }
    })
    
    for ev in conn:
        print(ev.type, ev)

```

The client receives `conversation.item.input_audio_transcription.delta` events for live transcription and `response.output_audio.delta` for generated speech chunks.

## Summary

- **Hardware alignment**: Install CUDA-specific wheels for Qwen3-TTS and select appropriate backends for CPU, CUDA, or Apple Silicon environments.
- **Latency control**: Leverage the `CancelScope` mechanism in [`src/speech_to_speech/pipeline/cancel_scope.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/cancel_scope.py) to handle interruptions and prevent stale audio generation.
- **Scalability**: Use `--num_pipelines` and the `ThreadManager` in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py) to run multiple concurrent instances.
- **API compatibility**: The Realtime mode exposes an OpenAI-compatible WebSocket API at `/v1/realtime`, supporting standard client libraries.
- **Dynamic configuration**: Update VAD thresholds and session parameters via `session.update` events without service restarts.
- **Security**: Configure API keys through environment variables only, never hard-coding secrets in configuration files.

## Frequently Asked Questions

### How do I select the correct CUDA wheel for Qwen3-TTS in production?

Match the wheel version to your CUDA driver version before installing the main package. If running CUDA 13.0, install `qwentts-cpp-python==0.3.0+cu130` from the Hugging Face wheels repository. The default wheel targets CUDA 12, which may cause runtime errors on systems with different driver versions.

### What is the difference between Realtime mode and WebSocket mode?

**Realtime mode** implements the full OpenAI Realtime API protocol with structured events like `response.done` and `session.update`, compatible with official OpenAI clients. **WebSocket mode** streams raw PCM audio without protocol overhead, suitable for custom client implementations that do not require OpenAI compatibility.

### How does the system handle user interruptions during speech generation?

The VAD detects new speech through `speech_started` events, triggering the `CancelScope` object to abort stale LLM and TTS generations. This mechanism ensures only the current user turn receives audio output, preventing overlapping responses and reducing latency.

### Can I update VAD sensitivity without restarting the server?

Yes. The `RuntimeConfig` model in [`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py) supports hot-reloading via `session.update` events. Send a WebSocket message with updated `turn_detection` parameters to adjust thresholds like `thresh` or `min_speech_ms` dynamically.