# Best Practices for the Hugging Face Speech-to-Speech Library

> Master the Hugging Face Speech-to-Speech library. Learn to optimize deployment, reduce latency with efficient back-ends, and enhance your real-time applications. Get started now.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: best-practices
- Published: 2026-07-07

---

**Optimize your speech-to-speech deployment by selecting the appropriate mode (`local`, `socket`, `websocket`, or `realtime`), enabling Apple Silicon optimizations with `--local_mac_optimal_settings`, and pairing fast back-ends like `parakeet-tdt` for STT and `qwen3` for TTS to minimize latency while maintaining quality.**

The `huggingface/speech-to-speech` repository provides an end-to-end pipeline that converts spoken input into spoken output through a sequence of Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM) processing, and Text-to-Speech (TTS) synthesis. Understanding the architecture and configuration options available in the source code allows you to build everything from local prototypes to production-grade realtime services.

## Understanding the Pipeline Architecture

The pipeline follows a strict data flow implemented in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). Audio enters through a **LocalAudioStreamer**, socket, or WebSocket connection, then flows through specialized handlers:

1. **VADHandler** ([`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py)) detects speech activity and manages speculative turn IDs
2. **STT Handlers** convert audio to text (Whisper, Faster-Whisper, Paraformer, Parakeet-TDT, or MLX-Audio-Whisper)
3. **LLM Handlers** generate responses via OpenAI-compatible APIs, Transformers, or MLX-LM
4. **TTS Handlers** synthesize speech (ChatTTS, Facebook MMS, Pocket, Kokoro, or Qwen3)
5. **LMOutputProcessor** optionally inserts text-only messages for realtime UI feedback

The `build_pipeline` function (lines 581-661 in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)) wires these components together using thread-safe queues and returns a `ThreadManager` instance that monitors all handler threads.

## Selecting the Right Deployment Mode

The library supports four distinct operating modes controlled via the `--mode` argument:

- **`local`** – Development on a single machine with microphone and speaker access
- **`socket`** – TCP socket communication for custom client/server implementations
- **`websocket`** – Browser-based interfaces or WebSocket clients
- **`realtime`** – Production-grade multi-client deployment with OpenAI-compatible Realtime API support

Only `realtime` mode supports parallel pipelines via `--num_pipelines > 1`, as enforced by the guard in `main()` at lines 1109-1115. If you need speculative turn handling or live transcription for multiple concurrent users, you must use `realtime` mode.

## Optimizing Device and Backend Performance

### Device Configuration

Set the global device using `--device cuda` for NVIDIA GPUs, `--device mps` for Apple Silicon, or `--device cpu` for CPU inference. The `overwrite_device_argument` function propagates this setting to every handler in the pipeline.

For Apple Silicon users, enable `--local_mac_optimal_settings` to automatically configure:
- `stt=parakeet-tdt`
- `llm_backend=mlx-lm`
- `tts=qwen3`
- `device=mps` and `mode=local`

This optimization is defined in `optimal_mac_settings` at lines 31-46 of [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py). The library warns if you run on macOS without these settings via `check_mac_settings` (lines 49-61).

### Backend Selection

Choose handlers based on your latency and quality requirements:

**STT Options:**
- **`parakeet-tdt`** – Fastest option with good quality and live transcription support (recommended default)
- **`whisper`** – Use for maximum accuracy when latency is less critical
- **`whisper-mlx`** – Optimized for Apple Silicon

**LLM Options:**
- **`responses-api`** – OpenAI-compatible chat completions (recommended for production)
- **`transformers`** – Local Hugging Face models
- **`mlx-lm`** – Apple-optimized local inference

**TTS Options:**
- **`qwen3`** – High quality, runs on both CPU and GPU (recommended default)
- **`pocket`** – Lowest latency for CPU-only deployments
- **`kokoro`** – Experimental style-transfer capabilities

All back-ends expose handler-specific arguments via the `--<handler>_handler_kwargs` namespace. The `rename_args` helper (lines 20-26) automatically strips prefixes to build generation-specific keyword arguments.

## Handling Live Transcription and Speculative Turns

Enable **live transcription** with `--enable_live_transcription true`. The VAD handler emits progressive audio chunks via `_process_realtime` (lines 69-102 in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)), allowing your UI to display partial transcripts before the turn completes.

**Speculative turns** automatically activate when running in `realtime` mode with live transcription enabled. This feature, implemented via `SpeculativeTurnTracker` in [`pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline/speculative_turns.py), allows the VAD to "re-open" a turn if the user speaks again before the assistant finishes responding, significantly reducing perceived latency.

Note that when running multiple pipelines on macOS (`--num_pipelines > 1`), live transcription is automatically disabled because the MLX lock would generate excessive warnings (see `main()` lines 1020-1028).

## Implementation Examples

### Local Development with CLI

Run a complete local pipeline with debug logging:

```bash
python -m speech_to_speech.s2s_pipeline \
  --mode local \
  --device cuda \
  --stt whisper \
  --llm_backend transformers \
  --tts qwen3 \
  --log_level debug

```

### Embedding in Python Applications

Import the pipeline builder functions for programmatic control:

```python
from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse arguments or build ParsedArguments manually

args = parse_arguments()

# Apply defaults and optimizations

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Initialize queues and events

queues_and_events = initialize_queues_and_events()

# Build the pipeline

pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues_and_events,
)

# Execute

pipeline_manager.start()
pipeline_manager.wait()  # Blocks until stop() is called

```

### Production Realtime Server

Deploy a WebSocket server supporting multiple concurrent conversations:

```bash
python -m speech_to_speech.s2s_pipeline \
  --mode realtime \
  --ws_host 0.0.0.0 \
  --ws_port 8000 \
  --stt parakeet-tdt \
  --tts qwen3 \
  --llm_backend responses-api \
  --num_pipelines 4

```

Clients connect to `ws://<host>:8000` streaming raw 16-kHz PCM audio. The server automatically manages speculative turns and live transcription for each connected client.

## Graceful Shutdown and Debugging

The CLI registers SIGINT/SIGTERM handlers that invoke `pipeline_manager.stop()` (see `main()` lines 1056-1064). When embedding the pipeline, always call `pipeline_manager.stop()` before exiting to ensure all threads terminate cleanly.

Enable detailed logging with `--log_level debug` to monitor queue activity, turn management, and audio chunk processing. The VAD handler throttles logging to once per second to prevent console flooding (lines 48-56 in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)).

## Summary

- **Select mode based on scale**: Use `local` for development, `realtime` for production multi-client deployments
- **Optimize for hardware**: Enable `--local_mac_optimal_settings` on Apple Silicon; set `--device cuda` for NVIDIA GPUs
- **Balance speed and quality**: Default to `parakeet-tdt` for STT and `qwen3` for TTS; switch to `whisper` or `transformers` when accuracy matters more than latency
- **Enable progressive feedback**: Turn on `--enable_live_transcription` for realtime UIs (disabled automatically with multiple pipelines on macOS)
- **Clean shutdown**: Always call `pipeline_manager.stop()` or rely on the CLI's signal handlers to prevent zombie threads

## Frequently Asked Questions

### What is the difference between `websocket` and `realtime` modes?

The `websocket` mode runs a single pipeline served over WebSocket, suitable for one-to-one browser interactions. The `realtime` mode creates a pool of isolated pipelines (controlled by `--num_pipelines`) and provides OpenAI-compatible Realtime API support, speculative turn handling, and live transcription features required for production multi-client deployments.

### How do I optimize performance on Apple Silicon Macs?

Enable `--local_mac_optimal_settings` to automatically configure `parakeet-tdt` for STT, `mlx-lm` for LLM inference, `qwen3` for TTS, and `mps` as the device. This setting, defined in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) lines 31-46, ensures all back-ends use Apple-optimized code paths. Avoid running multiple pipelines (`--num_pipelines > 1`) with live transcription on macOS, as the MLX framework's global lock forces automatic disabling of progressive audio emission.

### What are speculative turns and when should I use them?

Speculative turns allow the Voice Activity Detection system to re-open a conversation turn if the user speaks again before the assistant finishes responding. This feature, implemented via `SpeculativeTurnTracker` in [`pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline/speculative_turns.py), reduces perceived latency in natural conversations. It activates automatically in `realtime` mode when live transcription is enabled, requiring no manual configuration.

### How do I properly shut down the pipeline when embedding it in my application?

Always call `pipeline_manager.stop()` before exiting your application. The CLI handles this automatically via SIGINT/SIGTERM handlers registered in `main()` (lines 1056-1064), but programmatic users must invoke this method to ensure all handler threads terminate cleanly and system resources are released.