# How to Use the Speech-to-Speech Command-Line Interface: A Complete Guide to huggingface/speech-to-speech

> Master the speech-to-speech command-line interface. Learn to configure VAD, STT, LLM, and TTS components for a custom audio pipeline. Get the complete guide.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**Run `speech-to-speech` with flags like `--mode local`, `--stt whisper`, and `--llm_backend transformers` to launch a modular pipeline that wires together VAD, STT, LLM, and TTS components via thread-safe queues.**

The **speech-to-speech command-line interface** is the primary entry point for the `huggingface/speech-to-speech` repository, a modular framework for building real-time voice-to-voice applications. This guide explains how to use the CLI, what happens under the hood, and how to customize every stage of the pipeline.

## CLI Entry Point and Architecture

The executable `speech-to-speech` is registered in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml):

```toml
[project.scripts]
speech-to-speech = "speech_to_speech.s2s_pipeline:main"

```

When invoked, `main()` in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) orchestrates six distinct phases:

1. **Argument parsing** — `parse_arguments()` (lines 29-50) builds dataclasses for each pipeline component
2. **Logging setup** — `setup_logger()` (lines 6-14) adds pipeline-specific prefixes to log lines
3. **Argument normalization** — `prepare_all_args()` (lines 80-104) applies defaults and macOS optimizations
4. **Queue creation** — `initialize_queues_and_events()` (lines 34-48) instantiates thread-safe communication channels
5. **Pipeline construction** — `build_pipeline()` wires handlers for audio I/O, VAD, STT, LLM, and TTS
6. **Thread management** — `ThreadManager` starts components and handles graceful shutdown on SIGINT/SIGTERM (lines 840-870)

Each handler implements the `BaseHandler` interface, reading from input `Queue` objects and writing to output queues. This design lets you swap any component by changing a single flag.

## Basic Usage: Local Microphone Mode

The simplest invocation runs the full pipeline on your local machine:

```bash
speech-to-speech \
  --mode local \
  --stt whisper \
  --tts qwen3 \
  --llm_backend transformers \
  --model_name Qwen/Qwen3-4B-Instruct-2507 \
  --device cuda

```

**What each flag does:**

- `--mode local` — Activates `LocalAudioStreamer` for microphone input and speaker output
- `--stt whisper` — Loads `WhisperSTTHandler` from [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py)
- `--tts qwen3` — Uses `Qwen3TTSHandler` from [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)
- `--llm_backend transformers` — Enables generic HuggingFace Transformers LLM support
- `--model_name` and `--device` — Forwarded via autogenerated `gen_kwargs` to underlying models

## Networking Modes: WebSocket and Realtime

### WebSocket Server Mode

Deploy the pipeline as a networked service for remote clients:

```bash
speech-to-speech \
  --mode websocket \
  --ws_host 0.0.0.0 \
  --ws_port 8765 \
  --stt faster-whisper \
  --tts pocket \
  --llm_backend responses-api \
  --responses_api_api_key $OPENAI_API_KEY

```

Key options:
- `--mode websocket` — Spins up `WebSocketStreamer` from [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py)
- `--stt faster-whisper` — Uses `FasterWhisperSTTHandler` for lower latency
- `--tts pocket` — Loads `PocketTTSHandler` from [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py)
- `--llm_backend responses-api` — Connects to OpenAI's Realtime "responses" endpoint

### Parallel Realtime Mode

Scale to multiple concurrent pipelines with OpenAI-style realtime serving:

```bash
speech-to-speech \
  --mode realtime \
  --num_pipelines 3 \
  --stt parakeet-tdt \
  --tts qwen3 \
  --llm_backend chat-completions \
  --chat_completions_api_key $OPENAI_API_KEY \
  --ws_host 0.0.0.0 \
  --ws_port 8000

```

- `--num_pipelines 3` — Creates three independent pipeline instances sharing a single `RealtimeServer` from [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py)
- `--stt parakeet-tdt` — Selects streaming STT with `ParakeetTDTHandler`
- `--llm_backend chat-completions` — Uses OpenAI Chat Completion API

## Configuration via JSON File

For complex setups, store parameters in a JSON file instead of command-line flags:

```json
{
  "mode": "local",
  "stt": "whisper",
  "tts": "qwen3",
  "llm_backend": "transformers",
  "model_name": "Qwen/Qwen3-4B-Instruct-2507",
  "device": "cuda",
  "log_level": "info"
}

```

Launch with:

```bash
speech-to-speech config.json

```

The parser detects the `.json` suffix and merges values with dataclass defaults (lines 38-44 of [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)). This approach simplifies version-controlled deployments and A/B testing.

## macOS Optimization

Apple Silicon users can enable automatic performance tuning:

```bash
speech-to-speech \
  --mode local \
  --local_mac_optimal_settings \
  --device mps \
  --stt whisper \
  --tts qwen3 \
  --llm_backend mlx-lm

```

The `optimal_mac_settings()` helper (lines 31-48) rewrites arguments to use Metal Performance Shaders (`mps` device) and MLX-optimized backends where available.

## Available STT, LLM, and TTS Backends

### Speech-to-Text Options

| Flag | Handler File | Notes |
|------|-----------|-------|
| `whisper` | [`whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_handler.py) | Default OpenAI Whisper |
| `whisper-mlx` | [`whisper_mlx_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_mlx_handler.py) | MLX-optimized for Apple Silicon |
| `mlx-audio-whisper` | [`mlx_audio_whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/mlx_audio_whisper_handler.py) | Alternative MLX implementation |
| `faster-whisper` | [`faster_whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/faster_whisper_handler.py) | CTranslate2 backend, lower latency |
| `paraformer` | [`paraformer_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/paraformer_handler.py) | Alibaba Paraformer model |
| `parakeet-tdt` | [`parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/parakeet_tdt_handler.py) | Streaming-optimized NVIDIA model |

STT handlers are instantiated in `get_stt_handler()` (lines 650-720).

### LLM Backend Options

| Flag | Handler File | Use Case |
|------|-----------|----------|
| `transformers` | [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py) | Local HuggingFace models |
| `mlx-lm` | [`mlx_lm_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/mlx_lm_handler.py) | Apple Silicon optimized |
| `responses-api` | [`openai_responses_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/openai_responses_handler.py) | OpenAI Realtime API |
| `chat-completions` | [`chat_completion_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat_completion_handler.py) | Standard OpenAI/compatible APIs |

Defined in `get_llm_handler()` (lines 850-910) with arguments in [`language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model_arguments.py).

### Text-to-Speech Options

| Flag | Handler File | Characteristics |
|------|-----------|---------------|
| `qwen3` | [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) | High quality, multilingual |
| `kokoro` | [`kokoro_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/kokoro_tts_handler.py) | Fast, lightweight |
| `chatTTS` | [`chattts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/chattts_handler.py) | Conversational quality |
| `pocket` | [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py) | Minimal dependencies |
| `facebookMMS` | [`facebook_mms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/facebook_mms_handler.py) | Massively Multilingual Speech |

Selected in `get_tts_handler()` (lines 950-1020) with arguments in files like [`qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_arguments.py).

## Key Source Files for CLI Customization

| Component | Path |
|-----------|------|
| CLI entry point & orchestration | [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) |
| Module-level argument dataclasses | [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) |
| STT argument definitions | [`src/speech_to_speech/arguments_classes/whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/whisper_stt_arguments.py) (and parallel files for other STTs) |
| LLM argument definitions | [`src/speech_to_speech/arguments_classes/language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/language_model_arguments.py) |
| TTS argument definitions | [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py) |
| VAD implementation | [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) |

## Summary

- **Install** the package to get the `speech-to-speech` executable, defined in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml)
- **Choose a mode**: `local` for development, `websocket` for networked clients, `realtime` for scaled deployments
- **Select backends** via `--stt`, `--llm_backend`, and `--tts` flags to trade latency, quality, and hardware compatibility
- **Use JSON configs** for reproducible, version-controlled pipeline definitions
- **Enable macOS optimizations** with `--local_mac_optimal_settings` for Apple Silicon performance

## Frequently Asked Questions

### How do I change the STT model without modifying code?

Pass a different `--stt` flag. The CLI supports `whisper`, `faster-whisper`, `paraformer`, `parakeet-tdt`, and MLX variants. Each maps to a distinct handler class instantiated in `get_stt_handler()`.

### Can I run multiple pipeline instances on one server?

Yes. Use `--mode realtime` with `--num_pipelines N`. The `RealtimeServer` distributes connections across N independent pipelines that share no state, enabling concurrent user sessions.

### What happens if I don't specify `--device`?

The `prepare_all_args()` function applies platform-specific defaults. On CUDA systems it selects `cuda`; on macOS with `--local_mac_optimal_settings`, it selects `mps`. Otherwise CPU is used.

### How do I add a custom TTS or LLM backend?

Implement a `BaseHandler` subclass in the appropriate `TTS/` or `LLM/` directory, create a corresponding argument dataclass in `arguments_classes/`, and register the new flag in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py). The modular queue-based architecture requires no changes to core orchestration code.