# Speech-to-Speech Pipeline Production Deployment: Best Practices for High-Performance Voice AI

> Deploy huggingface/speech-to-speech in production with best practices for high-performance voice AI. Learn to minimize latency using realtime mode GPU acceleration pooled pipelines and self-hosted LLM backends.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: best-practices
- Published: 2026-08-05

---

**Deploy the `huggingface/speech-to-speech` pipeline in production using `realtime` mode with GPU acceleration, pooled pipelines for concurrency, and self-hosted LLM backends to minimize latency and external dependencies.**

The `huggingface/speech-to-speech` repository implements a modular, low-latency voice agent pipeline (VAD → STT → LLM → TTS) designed for production voice AI applications. This guide covers concrete best practices for configuring the pipeline for production deployment, with direct references to the source code implementation in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) and supporting modules.

---

## Choose the Correct Run Mode for Production

The pipeline supports four distinct run modes selectable via `--mode`. For production deployment, **only `realtime` mode** provides the protocol compatibility and scalability required for external client integration.

| Mode | Transport | Production Suitability |
|------|-----------|------------------------|
| `realtime` | OpenAI Realtime-compatible WebSocket / WebRTC | **Recommended** — standard protocol for apps, browsers, robots |
| `local` | Direct microphone / speaker | Debugging only — single-user, no network interface |
| `raw-websocket` | Raw PCM over WebSocket | Minimal custom clients, lacks full protocol features |
| `socket` | Raw PCM over TCP | Legacy setups — avoid for new deployments |

The mode selection logic in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (lines 76–90) instantiates communication handlers based on `module_kwargs.mode`:

```python

# From s2s_pipeline.py - mode handler selection

if module_kwargs.mode == "realtime":
    communication_handler = RealtimeCommunicationHandler(...)
elif module_kwargs.mode == "local":
    communication_handler = LocalCommunicationHandler(...)

# ... additional modes

```

**Production command:**

```bash
speech-to-speech --mode realtime [additional flags]

```

---

## Enable Concurrency with Pipeline Pools

Production workloads require handling multiple simultaneous conversations. The pipeline supports pooling via `--num_pipelines N`, where each pipeline maintains isolated VAD/STT/LLM/TTS state.

The pool construction occurs in `_build_realtime_pipeline_unit` (lines 528–564 in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)):

```python
def _build_realtime_pipeline_unit(self, unit_id: int, ...):
    # Each unit: independent handlers + dedicated queues

    vad_handler = VADHandler(...)
    stt_handler = ParakeetTDTSTTHandler(...)  # or selected backend

    llm_handler = ResponsesAPILLMHandler(...)
    tts_handler = Qwen3TTSHandler(...)

```

**Configuration constraints** (enforced in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py), lines 94–100):

- `num_pipelines >= 1` (required)
- `num_pipelines > 1` only valid with `mode=realtime`

**Best practice:** Set `--num_pipelines` to 1.5× expected peak concurrent sessions. Example for GPU server:

```bash
speech-to-speech \
    --mode realtime \
    --num_pipelines 8 \
    --device cuda

```

---

## Select Hardware-Optimized Backends

Component backend selection critically impacts latency and throughput. The `overwrite_device_argument` helper (lines 304–318) propagates a common device setting to all handlers.

### Recommended Production Stack

| Component | Backend | CLI Flag | Rationale |
|-----------|---------|----------|-----------|
| **Device** | CUDA (Linux) / MPS (macOS) | `--device cuda` / `--device mps` | Maximum GPU utilization |
| **STT** | Parakeet-TDT | `--stt parakeet-tdt` | Low latency, partial transcripts, CPU/GPU agnostic |
| **LLM** | Self-hosted vLLM/llama.cpp | `--llm_backend responses-api --responses_api_base_url <url>` | Predictable latency, no API key exposure |
| **TTS** | Qwen3 | `--tts qwen3` | Streaming-optimized, GGML/MLX variants |

**Full production GPU command:**

```bash
export OPENAI_API_KEY="sk-..."  # for LLM proxy if enabled

speech-to-speech \
    --mode realtime \
    --device cuda \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --responses_api_base_url http://localhost:8000/v1 \
    --responses_api_api_key "$OPENAI_API_KEY" \
    --tts qwen3 \
    --num_pipelines 4 \
    --enable_live_transcription

```

CLI arguments are normalized via `rename_args` (lines 222–236), mapping prefixed options to handler-specific configurations.

---

## Apply macOS-Specific Optimizations Correctly

Apple Silicon deployments use a dedicated optimization path. The `--local_mac_optimal_settings` flag triggers `optimal_mac_settings` (lines 71–84), which:

- Forces `device=mps`
- Selects Parakeet-TDT (STT), MLX-LM (LLM), Qwen3-TTS (TTS)
- Switches mode to `local`

**Critical:** This flag validates platform in `check_mac_settings` (lines 90–97) and raises `ValueError` on non-macOS systems.

```bash

# Apple Silicon single-pipeline deployment

speech-to-speech \
    --local_mac_optimal_settings \
    --mode local \
    --log_level info

```

---

## Secure the LLM Proxy

The optional LLM proxy exposes a configured remote LLM as an OpenAI-compatible HTTP endpoint. Built in `build_llm_proxy_config` (lines 38–64), it launches only when `--enable_llm_proxy` is set.

**Security warning:** The proxy performs **no authentication**. Production deployments must:

- Place behind firewall or API gateway with authentication
- Restrict to trusted internal networks only
- Never expose directly to public internet

```bash

# LLM proxy behind internal firewall

speech-to-speech \
    --enable_llm_proxy \
    --llm_backend responses-api \
    --responses_api_base_url http://internal-llm:8000/v1

```

---

## Configure Logging and Graceful Shutdown

Production observability requires structured logging. The `setup_logger` function (lines 46–61) configures pipeline-wide formatting with `PipelineLogFilter` for per-handler log prefixes.

```bash
--log_level info    # Normal operations

--log_level debug   # Deep troubleshooting

```

Graceful shutdown is implemented via signal handler (lines 122–128), ensuring clean thread termination:

```python
def signal_handler(signum, frame):
    stop_event.set()  # Signals all pipeline threads

    pipeline.join()   # Waits for clean exit

```

---

## Deploy with Docker Compose

The repository includes [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) for reproducible production deployment. It orchestrates:

- LLM server (llama.cpp or vLLM)
- Speech-to-speech pipeline in `realtime` mode
- Exposed ports: `8080` (Realtime API), `12345`, `12346`

**Prerequisites:** NVIDIA Container Toolkit for GPU acceleration.

```bash

# Production container deployment

docker compose up -d

```

---

## Prevent Memory Issues in Long-Running Services

Memory management differs by mode. Only `realtime` and `raw-websocket` allocate `text_output_queue` for streaming transcription events (lines 672–674). Other modes set it to `None`, preventing unbounded queue growth in long-running deployments.

---

## Manage Secrets and Environment Variables

The pipeline reads API keys from environment variables:

| Variable | Purpose |
|----------|---------|
| `OPENAI_API_KEY` | LLM proxy authentication |
| `HF_TOKEN` | Hugging Face model access |

**Never hard-code secrets.** Use deployment-time environment files or secret managers. The codebase never logs these values.

---

## Validate Configuration Early

The `main()` entry point enforces constraints before launch:

| Check | Location | Failure Mode |
|-------|----------|--------------|
| `num_pipelines >= 1` | Line 554 | `ValueError` |
| `num_pipelines > 1` requires `mode=realtime` | Line 577 | `ValueError` |
| macOS settings on valid platform | Line 90 | `ValueError` |

These validations enable automated CI/CD gates for production deployments.

---

## Summary

- **Use `realtime` mode** for production API deployments with external clients
- **Scale with `--num_pipelines`** — set to expected peak concurrent sessions
- **Prefer self-hosted LLMs** via `responses-api` or `chat-completions` backends
- **Deploy on CUDA (Linux)** or **MPS (macOS)** with appropriate backend selection
- **Secure the LLM proxy** behind firewalls — it has no built-in authentication
- **Containerize with Docker Compose** for reproducible, GPU-accelerated deployments
- **Validate configuration** through built-in constraints before production launch

---

## Frequently Asked Questions

### What is the difference between `realtime` and `local` mode in the speech-to-speech pipeline?

`realtime` mode implements the OpenAI Realtime-compatible protocol over WebSocket/WebRTC, enabling external clients like browsers and mobile apps to connect. `local` mode binds directly to microphone and speaker hardware for single-user debugging. Production deployments must use `realtime` mode for network accessibility and scalability.

### How many pipeline instances should I run for production workloads?

Set `--num_pipelines` to your expected peak concurrent sessions, typically 4–8 on a GPU server. Each pipeline maintains independent state, so this directly determines your connection capacity. The server rejects connections beyond this limit. Only `realtime` mode supports `num_pipelines > 1`.

### Can I run the speech-to-speech pipeline on Apple Silicon in production?

Apple Silicon supports production deployment but requires different optimization. Use `--local_mac_optimal_settings` to enable `mps` device with MLX-optimized backends (Parakeet-TDT, MLX-LM, Qwen3-TTS). For multi-user production on Apple Silicon, containerize with appropriate resource limits rather than using `local` mode.

### Is the built-in LLM proxy secure for production use?

No. The LLM proxy (enabled with `--enable_llm_proxy`) provides no authentication. Production deployments must place it behind a firewall, API gateway, or reverse proxy with authentication. Exposing the proxy directly to the internet creates a security vulnerability where anyone can access your configured LLM backend.