How to Use the Speech-to-Speech Command-Line Interface: A Complete Guide
Run the speech-to-speech command with flags like --mode, --stt, --tts, and --llm_backend to launch a modular pipeline that chains together voice activity detection, speech recognition, language models, and text-to-speech synthesis.
The huggingface/speech-to-speech repository provides a production-ready CLI for building real-time speech-to-speech applications. The entry point is defined in pyproject.toml as [project.scripts] speech-to-speech = "speech_to_speech.s2s_pipeline:main", which routes to the main() function in src/speech_to_speech/s2s_pipeline.py. This executable orchestrates a threaded pipeline where each component communicates via thread-safe queues, allowing you to swap STT, LLM, or TTS backends without changing the core architecture.
Installation and Entry Point
After installing the package, the speech-to-speech command becomes available in your shell. The CLI entry point is declared in pyproject.toml and points to the main() function inside src/speech_to_speech/s2s_pipeline.py. When executed, the script performs six distinct phases: argument parsing, logging setup, argument normalization, queue initialization, pipeline construction, and thread management.
Core Architecture of the CLI
The command-line interface follows a strict initialization sequence defined in s2s_pipeline.py:
-
Argument Parsing – The
parse_arguments()function (lines 29-50) pre-parses--llm_backendto avoid field name conflicts, then usesHfArgumentParserto process CLI flags or JSON configuration files. -
Logging Setup –
setup_logger()(lines 6-14) installs a custom filter that prefixes every log line with the pipeline name, essential when running multiple pipelines simultaneously. -
Argument Normalisation –
prepare_all_args()(lines 80-104) consolidates defaults, applies macOS-specific optimizations if requested, and rewrites prefixed arguments into a genericgen_kwargsdictionary consumed by model handlers. -
Queue Creation –
initialize_queues_and_events()(lines 34-48) instantiates the thread-safe queues that shuttle audio frames, VAD events, transcriptions, LLM prompts, and final audio packets between components. -
Pipeline Construction –
build_pipeline()wires together the VADHandler, selected STT handler, LLM handler, LM output processor, and TTS handler based on your CLI flags. -
Thread Management – A
ThreadManagerstarts each handler in its own thread and monitors health until SIGINT/SIGTERM (lines 840-870).
Configuration Methods
You can configure the pipeline via command-line flags or a JSON configuration file. The parser detects the .json suffix automatically and merges the file contents with the argument dataclasses (lines 38-44).
To use a JSON config:
{
"mode": "local",
"stt": "whisper",
"tts": "qwen3",
"llm_backend": "transformers",
"model_name": "Qwen/Qwen3-4B-Instruct-2507",
"device": "cuda",
"log_level": "info"
}
Then run:
speech-to-speech config.json
Running the Pipeline
Local Microphone Mode
Use local mode when processing audio from your machine's microphone and playing output through the speakers. This instantiates LocalAudioStreamer and routes audio through the full pipeline locally.
speech-to-speech \
--mode local \
--stt whisper \
--tts qwen3 \
--llm_backend transformers \
--model_name Qwen/Qwen3-4B-Instruct-2507 \
--device cuda
--stt whisperloads theWhisperSTTHandlerfromsrc/speech_to_speech/STT/whisper_stt_handler.py.--tts qwen3activates the Qwen-3 TTS handler found insrc/speech_to_speech/TTS/qwen3_tts_handler.py.--llm_backend transformersuses the generic HuggingFace transformers LLM handler.
WebSocket Server Mode
WebSocket mode exposes the pipeline as a network service, allowing remote clients to stream audio over WebSocket connections.
speech-to-speech \
--mode websocket \
--ws_host 0.0.0.0 \
--ws_port 8765 \
--stt faster-whisper \
--tts pocket \
--llm_backend responses-api \
--responses_api_api_key $OPENAI_API_KEY
--mode websocketspawns aWebSocketStreamer(implemented insrc/speech_to_speech/connections/websocket_streamer.py).--stt faster-whisperselects the Faster-Whisper implementation fromsrc/speech_to_speech/STT/faster_whisper_handler.py.--tts pocketenables the Pocket TTS handler.--llm_backend responses-apiconnects to the OpenAI Realtime responses endpoint.
Realtime Multi-Pipeline Mode
Realtime mode creates multiple independent pipeline instances sharing a single RealtimeServer, useful for high-concurrency scenarios.
speech-to-speech \
--mode realtime \
--num_pipelines 3 \
--stt parakeet-tdt \
--tts qwen3 \
--llm_backend chat-completions \
--chat_completions_api_key $OPENAI_API_KEY \
--ws_host 0.0.0.0 \
--ws_port 8000
--num_pipelines 3instantiates three parallel pipelines managed by theThreadManager.--stt parakeet-tdtselects the streaming Parakeet-TDT handler.--llm_backend chat-completionsuses the OpenAI Chat Completions API compatible with realtime services.
macOS Optimized Settings
For Apple Silicon devices, use the --local_mac_optimal_settings flag to automatically apply optimal configurations.
speech-to-speech \
--mode local \
--local_mac_optimal_settings \
--device mps \
--stt whisper \
--tts qwen3 \
--llm_backend mlx-lm
The optimal_mac_settings() helper (lines 31-48) rewrites arguments to use device="mps", tts="qwen3", and llm_backend="mlx-lm" for best performance on Apple hardware.
Key Components and Handler Selection
The CLI supports modular swapping of every pipeline component via specific flags:
Speech-to-Text (STT) options include whisper, whisper-mlx, mlx-audio-whisper, faster-whisper, paraformer, and parakeet-tdt. Handlers are instantiated in get_stt_handler() (lines 650-720).
Language Model (LLM) backends include responses-api, chat-completions, transformers, and mlx-lm. These are loaded in get_llm_handler() (lines 850-910) located in src/speech_to_speech/LLM/language_model.py.
Text-to-Speech (TTS) options include chatTTS, facebookMMS, pocket, kokoro, and qwen3. The factory function get_tts_handler() (lines 950-1020) routes to the appropriate handler file in src/speech_to_speech/TTS/.
Summary
- The
speech-to-speechCLI entry point is defined inpyproject.tomland implemented insrc/speech_to_speech/s2s_pipeline.py. - Argument parsing uses
HfArgumentParserwith dataclasses defined inarguments_classes/modules. - The pipeline supports three primary modes: local (microphone), websocket (network), and realtime (multi-pipeline).
- You can override any component (STT, LLM, TTS) via CLI flags without modifying source code.
- Configuration can be passed via command-line arguments or JSON files.
- macOS users should leverage
--local_mac_optimal_settingsfor automatic MPS optimization.
Frequently Asked Questions
How do I switch from Whisper to Faster-Whisper for STT?
Pass --stt faster-whisper when launching the CLI. This instantiates the FasterWhisperSTTHandler from src/speech_to_speech/STT/faster_whisper_handler.py instead of the default Whisper handler. The flag is processed in get_stt_handler() (lines 650-720) and requires the faster-whisper Python package to be installed.
Can I run multiple pipelines simultaneously for different users?
Yes. Use --mode realtime combined with --num_pipelines N to spawn N independent pipeline instances. The ThreadManager (lines 840-870) handles thread lifecycle for each pipeline, while the RealtimeServer manages shared WebSocket connections. Each pipeline maintains its own set of queues initialized by initialize_queues_and_events().
Where are the argument dataclasses defined?
Module-specific arguments reside in src/speech_to_speech/arguments_classes/. For example, STT arguments are in whisper_stt_arguments.py, LLM arguments in language_model_arguments.py, and TTS arguments in qwen3_tts_arguments.py. These dataclasses are registered with HfArgumentParser in parse_arguments() (lines 29-50).
How does the CLI handle configuration files?
If the final positional argument ends with .json, parse_arguments() (lines 38-44) loads the JSON and merges its keys with the dataclass instances. This allows you to store complex configurations (e.g., model paths, generation parameters) in version-controlled files rather than long shell commands.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →