How to Use the Speech-to-Speech Command-Line Interface: A Complete Guide to huggingface/speech-to-speech
Run speech-to-speech with flags like --mode local, --stt whisper, and --llm_backend transformers to launch a modular pipeline that wires together VAD, STT, LLM, and TTS components via thread-safe queues.
The speech-to-speech command-line interface is the primary entry point for the huggingface/speech-to-speech repository, a modular framework for building real-time voice-to-voice applications. This guide explains how to use the CLI, what happens under the hood, and how to customize every stage of the pipeline.
CLI Entry Point and Architecture
The executable speech-to-speech is registered in pyproject.toml:
[project.scripts]
speech-to-speech = "speech_to_speech.s2s_pipeline:main"
When invoked, main() in src/speech_to_speech/s2s_pipeline.py orchestrates six distinct phases:
- Argument parsing —
parse_arguments()(lines 29-50) builds dataclasses for each pipeline component - Logging setup —
setup_logger()(lines 6-14) adds pipeline-specific prefixes to log lines - Argument normalization —
prepare_all_args()(lines 80-104) applies defaults and macOS optimizations - Queue creation —
initialize_queues_and_events()(lines 34-48) instantiates thread-safe communication channels - Pipeline construction —
build_pipeline()wires handlers for audio I/O, VAD, STT, LLM, and TTS - Thread management —
ThreadManagerstarts components and handles graceful shutdown on SIGINT/SIGTERM (lines 840-870)
Each handler implements the BaseHandler interface, reading from input Queue objects and writing to output queues. This design lets you swap any component by changing a single flag.
Basic Usage: Local Microphone Mode
The simplest invocation runs the full pipeline on your local machine:
speech-to-speech \
--mode local \
--stt whisper \
--tts qwen3 \
--llm_backend transformers \
--model_name Qwen/Qwen3-4B-Instruct-2507 \
--device cuda
What each flag does:
--mode local— ActivatesLocalAudioStreamerfor microphone input and speaker output--stt whisper— LoadsWhisperSTTHandlerfromsrc/speech_to_speech/STT/whisper_stt_handler.py--tts qwen3— UsesQwen3TTSHandlerfromsrc/speech_to_speech/TTS/qwen3_tts_handler.py--llm_backend transformers— Enables generic HuggingFace Transformers LLM support--model_nameand--device— Forwarded via autogeneratedgen_kwargsto underlying models
Networking Modes: WebSocket and Realtime
WebSocket Server Mode
Deploy the pipeline as a networked service for remote clients:
speech-to-speech \
--mode websocket \
--ws_host 0.0.0.0 \
--ws_port 8765 \
--stt faster-whisper \
--tts pocket \
--llm_backend responses-api \
--responses_api_api_key $OPENAI_API_KEY
Key options:
--mode websocket— Spins upWebSocketStreamerfromsrc/speech_to_speech/connections/websocket_streamer.py--stt faster-whisper— UsesFasterWhisperSTTHandlerfor lower latency--tts pocket— LoadsPocketTTSHandlerfromsrc/speech_to_speech/TTS/pocket_tts_handler.py--llm_backend responses-api— Connects to OpenAI's Realtime "responses" endpoint
Parallel Realtime Mode
Scale to multiple concurrent pipelines with OpenAI-style realtime serving:
speech-to-speech \
--mode realtime \
--num_pipelines 3 \
--stt parakeet-tdt \
--tts qwen3 \
--llm_backend chat-completions \
--chat_completions_api_key $OPENAI_API_KEY \
--ws_host 0.0.0.0 \
--ws_port 8000
--num_pipelines 3— Creates three independent pipeline instances sharing a singleRealtimeServerfromsrc/speech_to_speech/api/openai_realtime/server.py--stt parakeet-tdt— Selects streaming STT withParakeetTDTHandler--llm_backend chat-completions— Uses OpenAI Chat Completion API
Configuration via JSON File
For complex setups, store parameters in a JSON file instead of command-line flags:
{
"mode": "local",
"stt": "whisper",
"tts": "qwen3",
"llm_backend": "transformers",
"model_name": "Qwen/Qwen3-4B-Instruct-2507",
"device": "cuda",
"log_level": "info"
}
Launch with:
speech-to-speech config.json
The parser detects the .json suffix and merges values with dataclass defaults (lines 38-44 of s2s_pipeline.py). This approach simplifies version-controlled deployments and A/B testing.
macOS Optimization
Apple Silicon users can enable automatic performance tuning:
speech-to-speech \
--mode local \
--local_mac_optimal_settings \
--device mps \
--stt whisper \
--tts qwen3 \
--llm_backend mlx-lm
The optimal_mac_settings() helper (lines 31-48) rewrites arguments to use Metal Performance Shaders (mps device) and MLX-optimized backends where available.
Available STT, LLM, and TTS Backends
Speech-to-Text Options
| Flag | Handler File | Notes |
|---|---|---|
whisper |
whisper_stt_handler.py |
Default OpenAI Whisper |
whisper-mlx |
whisper_mlx_handler.py |
MLX-optimized for Apple Silicon |
mlx-audio-whisper |
mlx_audio_whisper_handler.py |
Alternative MLX implementation |
faster-whisper |
faster_whisper_handler.py |
CTranslate2 backend, lower latency |
paraformer |
paraformer_handler.py |
Alibaba Paraformer model |
parakeet-tdt |
parakeet_tdt_handler.py |
Streaming-optimized NVIDIA model |
STT handlers are instantiated in get_stt_handler() (lines 650-720).
LLM Backend Options
| Flag | Handler File | Use Case |
|---|---|---|
transformers |
language_model.py |
Local HuggingFace models |
mlx-lm |
mlx_lm_handler.py |
Apple Silicon optimized |
responses-api |
openai_responses_handler.py |
OpenAI Realtime API |
chat-completions |
chat_completion_handler.py |
Standard OpenAI/compatible APIs |
Defined in get_llm_handler() (lines 850-910) with arguments in language_model_arguments.py.
Text-to-Speech Options
| Flag | Handler File | Characteristics |
|---|---|---|
qwen3 |
qwen3_tts_handler.py |
High quality, multilingual |
kokoro |
kokoro_tts_handler.py |
Fast, lightweight |
chatTTS |
chattts_handler.py |
Conversational quality |
pocket |
pocket_tts_handler.py |
Minimal dependencies |
facebookMMS |
facebook_mms_handler.py |
Massively Multilingual Speech |
Selected in get_tts_handler() (lines 950-1020) with arguments in files like qwen3_tts_arguments.py.
Key Source Files for CLI Customization
| Component | Path |
|---|---|
| CLI entry point & orchestration | src/speech_to_speech/s2s_pipeline.py |
| Module-level argument dataclasses | src/speech_to_speech/arguments_classes/module_arguments.py |
| STT argument definitions | src/speech_to_speech/arguments_classes/whisper_stt_arguments.py (and parallel files for other STTs) |
| LLM argument definitions | src/speech_to_speech/arguments_classes/language_model_arguments.py |
| TTS argument definitions | src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py |
| VAD implementation | src/speech_to_speech/VAD/vad_handler.py |
Summary
- Install the package to get the
speech-to-speechexecutable, defined inpyproject.toml - Choose a mode:
localfor development,websocketfor networked clients,realtimefor scaled deployments - Select backends via
--stt,--llm_backend, and--ttsflags to trade latency, quality, and hardware compatibility - Use JSON configs for reproducible, version-controlled pipeline definitions
- Enable macOS optimizations with
--local_mac_optimal_settingsfor Apple Silicon performance
Frequently Asked Questions
How do I change the STT model without modifying code?
Pass a different --stt flag. The CLI supports whisper, faster-whisper, paraformer, parakeet-tdt, and MLX variants. Each maps to a distinct handler class instantiated in get_stt_handler().
Can I run multiple pipeline instances on one server?
Yes. Use --mode realtime with --num_pipelines N. The RealtimeServer distributes connections across N independent pipelines that share no state, enabling concurrent user sessions.
What happens if I don't specify --device?
The prepare_all_args() function applies platform-specific defaults. On CUDA systems it selects cuda; on macOS with --local_mac_optimal_settings, it selects mps. Otherwise CPU is used.
How do I add a custom TTS or LLM backend?
Implement a BaseHandler subclass in the appropriate TTS/ or LLM/ directory, create a corresponding argument dataclass in arguments_classes/, and register the new flag in module_arguments.py. The modular queue-based architecture requires no changes to core orchestration code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →