Best Practices for the Hugging Face Speech-to-Speech Library

Optimize your speech-to-speech deployment by selecting the appropriate mode (local, socket, websocket, or realtime), enabling Apple Silicon optimizations with --local_mac_optimal_settings, and pairing fast back-ends like parakeet-tdt for STT and qwen3 for TTS to minimize latency while maintaining quality.

The huggingface/speech-to-speech repository provides an end-to-end pipeline that converts spoken input into spoken output through a sequence of Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM) processing, and Text-to-Speech (TTS) synthesis. Understanding the architecture and configuration options available in the source code allows you to build everything from local prototypes to production-grade realtime services.

Understanding the Pipeline Architecture

The pipeline follows a strict data flow implemented in src/speech_to_speech/s2s_pipeline.py. Audio enters through a LocalAudioStreamer, socket, or WebSocket connection, then flows through specialized handlers:

  1. VADHandler (src/speech_to_speech/VAD/vad_handler.py) detects speech activity and manages speculative turn IDs
  2. STT Handlers convert audio to text (Whisper, Faster-Whisper, Paraformer, Parakeet-TDT, or MLX-Audio-Whisper)
  3. LLM Handlers generate responses via OpenAI-compatible APIs, Transformers, or MLX-LM
  4. TTS Handlers synthesize speech (ChatTTS, Facebook MMS, Pocket, Kokoro, or Qwen3)
  5. LMOutputProcessor optionally inserts text-only messages for realtime UI feedback

The build_pipeline function (lines 581-661 in s2s_pipeline.py) wires these components together using thread-safe queues and returns a ThreadManager instance that monitors all handler threads.

Selecting the Right Deployment Mode

The library supports four distinct operating modes controlled via the --mode argument:

  • local – Development on a single machine with microphone and speaker access
  • socket – TCP socket communication for custom client/server implementations
  • websocket – Browser-based interfaces or WebSocket clients
  • realtime – Production-grade multi-client deployment with OpenAI-compatible Realtime API support

Only realtime mode supports parallel pipelines via --num_pipelines > 1, as enforced by the guard in main() at lines 1109-1115. If you need speculative turn handling or live transcription for multiple concurrent users, you must use realtime mode.

Optimizing Device and Backend Performance

Device Configuration

Set the global device using --device cuda for NVIDIA GPUs, --device mps for Apple Silicon, or --device cpu for CPU inference. The overwrite_device_argument function propagates this setting to every handler in the pipeline.

For Apple Silicon users, enable --local_mac_optimal_settings to automatically configure:

  • stt=parakeet-tdt
  • llm_backend=mlx-lm
  • tts=qwen3
  • device=mps and mode=local

This optimization is defined in optimal_mac_settings at lines 31-46 of s2s_pipeline.py. The library warns if you run on macOS without these settings via check_mac_settings (lines 49-61).

Backend Selection

Choose handlers based on your latency and quality requirements:

STT Options:

  • parakeet-tdt – Fastest option with good quality and live transcription support (recommended default)
  • whisper – Use for maximum accuracy when latency is less critical
  • whisper-mlx – Optimized for Apple Silicon

LLM Options:

  • responses-api – OpenAI-compatible chat completions (recommended for production)
  • transformers – Local Hugging Face models
  • mlx-lm – Apple-optimized local inference

TTS Options:

  • qwen3 – High quality, runs on both CPU and GPU (recommended default)
  • pocket – Lowest latency for CPU-only deployments
  • kokoro – Experimental style-transfer capabilities

All back-ends expose handler-specific arguments via the --<handler>_handler_kwargs namespace. The rename_args helper (lines 20-26) automatically strips prefixes to build generation-specific keyword arguments.

Handling Live Transcription and Speculative Turns

Enable live transcription with --enable_live_transcription true. The VAD handler emits progressive audio chunks via _process_realtime (lines 69-102 in vad_handler.py), allowing your UI to display partial transcripts before the turn completes.

Speculative turns automatically activate when running in realtime mode with live transcription enabled. This feature, implemented via SpeculativeTurnTracker in pipeline/speculative_turns.py, allows the VAD to "re-open" a turn if the user speaks again before the assistant finishes responding, significantly reducing perceived latency.

Note that when running multiple pipelines on macOS (--num_pipelines > 1), live transcription is automatically disabled because the MLX lock would generate excessive warnings (see main() lines 1020-1028).

Implementation Examples

Local Development with CLI

Run a complete local pipeline with debug logging:

python -m speech_to_speech.s2s_pipeline \
  --mode local \
  --device cuda \
  --stt whisper \
  --llm_backend transformers \
  --tts qwen3 \
  --log_level debug

Embedding in Python Applications

Import the pipeline builder functions for programmatic control:

from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse arguments or build ParsedArguments manually

args = parse_arguments()

# Apply defaults and optimizations

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Initialize queues and events

queues_and_events = initialize_queues_and_events()

# Build the pipeline

pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues_and_events,
)

# Execute

pipeline_manager.start()
pipeline_manager.wait()  # Blocks until stop() is called

Production Realtime Server

Deploy a WebSocket server supporting multiple concurrent conversations:

python -m speech_to_speech.s2s_pipeline \
  --mode realtime \
  --ws_host 0.0.0.0 \
  --ws_port 8000 \
  --stt parakeet-tdt \
  --tts qwen3 \
  --llm_backend responses-api \
  --num_pipelines 4

Clients connect to ws://<host>:8000 streaming raw 16-kHz PCM audio. The server automatically manages speculative turns and live transcription for each connected client.

Graceful Shutdown and Debugging

The CLI registers SIGINT/SIGTERM handlers that invoke pipeline_manager.stop() (see main() lines 1056-1064). When embedding the pipeline, always call pipeline_manager.stop() before exiting to ensure all threads terminate cleanly.

Enable detailed logging with --log_level debug to monitor queue activity, turn management, and audio chunk processing. The VAD handler throttles logging to once per second to prevent console flooding (lines 48-56 in vad_handler.py).

Summary

  • Select mode based on scale: Use local for development, realtime for production multi-client deployments
  • Optimize for hardware: Enable --local_mac_optimal_settings on Apple Silicon; set --device cuda for NVIDIA GPUs
  • Balance speed and quality: Default to parakeet-tdt for STT and qwen3 for TTS; switch to whisper or transformers when accuracy matters more than latency
  • Enable progressive feedback: Turn on --enable_live_transcription for realtime UIs (disabled automatically with multiple pipelines on macOS)
  • Clean shutdown: Always call pipeline_manager.stop() or rely on the CLI's signal handlers to prevent zombie threads

Frequently Asked Questions

What is the difference between websocket and realtime modes?

The websocket mode runs a single pipeline served over WebSocket, suitable for one-to-one browser interactions. The realtime mode creates a pool of isolated pipelines (controlled by --num_pipelines) and provides OpenAI-compatible Realtime API support, speculative turn handling, and live transcription features required for production multi-client deployments.

How do I optimize performance on Apple Silicon Macs?

Enable --local_mac_optimal_settings to automatically configure parakeet-tdt for STT, mlx-lm for LLM inference, qwen3 for TTS, and mps as the device. This setting, defined in s2s_pipeline.py lines 31-46, ensures all back-ends use Apple-optimized code paths. Avoid running multiple pipelines (--num_pipelines > 1) with live transcription on macOS, as the MLX framework's global lock forces automatic disabling of progressive audio emission.

What are speculative turns and when should I use them?

Speculative turns allow the Voice Activity Detection system to re-open a conversation turn if the user speaks again before the assistant finishes responding. This feature, implemented via SpeculativeTurnTracker in pipeline/speculative_turns.py, reduces perceived latency in natural conversations. It activates automatically in realtime mode when live transcription is enabled, requiring no manual configuration.

How do I properly shut down the pipeline when embedding it in my application?

Always call pipeline_manager.stop() before exiting your application. The CLI handles this automatically via SIGINT/SIGTERM handlers registered in main() (lines 1056-1064), but programmatic users must invoke this method to ensure all handler threads terminate cleanly and system resources are released.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →