How to Use the Speech-to-Speech Command-Line Interface: A Complete Guide to huggingface/speech-to-speech

Run speech-to-speech with flags like --mode local, --stt whisper, and --llm_backend transformers to launch a modular pipeline that wires together VAD, STT, LLM, and TTS components via thread-safe queues.

The speech-to-speech command-line interface is the primary entry point for the huggingface/speech-to-speech repository, a modular framework for building real-time voice-to-voice applications. This guide explains how to use the CLI, what happens under the hood, and how to customize every stage of the pipeline.

CLI Entry Point and Architecture

The executable speech-to-speech is registered in pyproject.toml:

[project.scripts]
speech-to-speech = "speech_to_speech.s2s_pipeline:main"

When invoked, main() in src/speech_to_speech/s2s_pipeline.py orchestrates six distinct phases:

  1. Argument parsing — parse_arguments() (lines 29-50) builds dataclasses for each pipeline component
  2. Logging setup — setup_logger() (lines 6-14) adds pipeline-specific prefixes to log lines
  3. Argument normalization — prepare_all_args() (lines 80-104) applies defaults and macOS optimizations
  4. Queue creation — initialize_queues_and_events() (lines 34-48) instantiates thread-safe communication channels
  5. Pipeline construction — build_pipeline() wires handlers for audio I/O, VAD, STT, LLM, and TTS
  6. Thread management — ThreadManager starts components and handles graceful shutdown on SIGINT/SIGTERM (lines 840-870)

Each handler implements the BaseHandler interface, reading from input Queue objects and writing to output queues. This design lets you swap any component by changing a single flag.

Basic Usage: Local Microphone Mode

The simplest invocation runs the full pipeline on your local machine:

speech-to-speech \
  --mode local \
  --stt whisper \
  --tts qwen3 \
  --llm_backend transformers \
  --model_name Qwen/Qwen3-4B-Instruct-2507 \
  --device cuda

What each flag does:

Networking Modes: WebSocket and Realtime

WebSocket Server Mode

Deploy the pipeline as a networked service for remote clients:

speech-to-speech \
  --mode websocket \
  --ws_host 0.0.0.0 \
  --ws_port 8765 \
  --stt faster-whisper \
  --tts pocket \
  --llm_backend responses-api \
  --responses_api_api_key $OPENAI_API_KEY

Key options:

Parallel Realtime Mode

Scale to multiple concurrent pipelines with OpenAI-style realtime serving:

speech-to-speech \
  --mode realtime \
  --num_pipelines 3 \
  --stt parakeet-tdt \
  --tts qwen3 \
  --llm_backend chat-completions \
  --chat_completions_api_key $OPENAI_API_KEY \
  --ws_host 0.0.0.0 \
  --ws_port 8000
  • --num_pipelines 3 — Creates three independent pipeline instances sharing a single RealtimeServer from src/speech_to_speech/api/openai_realtime/server.py
  • --stt parakeet-tdt — Selects streaming STT with ParakeetTDTHandler
  • --llm_backend chat-completions — Uses OpenAI Chat Completion API

Configuration via JSON File

For complex setups, store parameters in a JSON file instead of command-line flags:

{
  "mode": "local",
  "stt": "whisper",
  "tts": "qwen3",
  "llm_backend": "transformers",
  "model_name": "Qwen/Qwen3-4B-Instruct-2507",
  "device": "cuda",
  "log_level": "info"
}

Launch with:

speech-to-speech config.json

The parser detects the .json suffix and merges values with dataclass defaults (lines 38-44 of s2s_pipeline.py). This approach simplifies version-controlled deployments and A/B testing.

macOS Optimization

Apple Silicon users can enable automatic performance tuning:

speech-to-speech \
  --mode local \
  --local_mac_optimal_settings \
  --device mps \
  --stt whisper \
  --tts qwen3 \
  --llm_backend mlx-lm

The optimal_mac_settings() helper (lines 31-48) rewrites arguments to use Metal Performance Shaders (mps device) and MLX-optimized backends where available.

Available STT, LLM, and TTS Backends

Speech-to-Text Options

Flag Handler File Notes
whisper whisper_stt_handler.py Default OpenAI Whisper
whisper-mlx whisper_mlx_handler.py MLX-optimized for Apple Silicon
mlx-audio-whisper mlx_audio_whisper_handler.py Alternative MLX implementation
faster-whisper faster_whisper_handler.py CTranslate2 backend, lower latency
paraformer paraformer_handler.py Alibaba Paraformer model
parakeet-tdt parakeet_tdt_handler.py Streaming-optimized NVIDIA model

STT handlers are instantiated in get_stt_handler() (lines 650-720).

LLM Backend Options

Flag Handler File Use Case
transformers language_model.py Local HuggingFace models
mlx-lm mlx_lm_handler.py Apple Silicon optimized
responses-api openai_responses_handler.py OpenAI Realtime API
chat-completions chat_completion_handler.py Standard OpenAI/compatible APIs

Defined in get_llm_handler() (lines 850-910) with arguments in language_model_arguments.py.

Text-to-Speech Options

Flag Handler File Characteristics
qwen3 qwen3_tts_handler.py High quality, multilingual
kokoro kokoro_tts_handler.py Fast, lightweight
chatTTS chattts_handler.py Conversational quality
pocket pocket_tts_handler.py Minimal dependencies
facebookMMS facebook_mms_handler.py Massively Multilingual Speech

Selected in get_tts_handler() (lines 950-1020) with arguments in files like qwen3_tts_arguments.py.

Key Source Files for CLI Customization

Component Path
CLI entry point & orchestration src/speech_to_speech/s2s_pipeline.py
Module-level argument dataclasses src/speech_to_speech/arguments_classes/module_arguments.py
STT argument definitions src/speech_to_speech/arguments_classes/whisper_stt_arguments.py (and parallel files for other STTs)
LLM argument definitions src/speech_to_speech/arguments_classes/language_model_arguments.py
TTS argument definitions src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py
VAD implementation src/speech_to_speech/VAD/vad_handler.py

Summary

  • Install the package to get the speech-to-speech executable, defined in pyproject.toml
  • Choose a mode: local for development, websocket for networked clients, realtime for scaled deployments
  • Select backends via --stt, --llm_backend, and --tts flags to trade latency, quality, and hardware compatibility
  • Use JSON configs for reproducible, version-controlled pipeline definitions
  • Enable macOS optimizations with --local_mac_optimal_settings for Apple Silicon performance

Frequently Asked Questions

How do I change the STT model without modifying code?

Pass a different --stt flag. The CLI supports whisper, faster-whisper, paraformer, parakeet-tdt, and MLX variants. Each maps to a distinct handler class instantiated in get_stt_handler().

Can I run multiple pipeline instances on one server?

Yes. Use --mode realtime with --num_pipelines N. The RealtimeServer distributes connections across N independent pipelines that share no state, enabling concurrent user sessions.

What happens if I don't specify --device?

The prepare_all_args() function applies platform-specific defaults. On CUDA systems it selects cuda; on macOS with --local_mac_optimal_settings, it selects mps. Otherwise CPU is used.

How do I add a custom TTS or LLM backend?

Implement a BaseHandler subclass in the appropriate TTS/ or LLM/ directory, create a corresponding argument dataclass in arguments_classes/, and register the new flag in module_arguments.py. The modular queue-based architecture requires no changes to core orchestration code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →