# How to Install Hugging Face Speech-to-Speech: Complete Setup Guide

> Install Hugging Face Speech-to-Speech easily with pip. Get the complete VAD STT LLM TTS pipeline for Linux CUDA, macOS MLX, or CPU. Follow our setup guide.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: getting-started
- Published: 2026-08-01

---

**The fastest way to install Hugging Face Speech-to-Speech is via `pip install speech-to-speech`, which bundles the full VAD → STT → LLM → TTS pipeline with platform-specific binary wheels for Linux CUDA, macOS MLX, or CPU-only inference.**

The `huggingface/speech-to-speech` repository provides a fully modular, local voice agent that processes audio through a complete pipeline without requiring cloud APIs. This Python package automatically handles platform-specific dependencies, making it possible to deploy a realtime speech-to-speech server on Linux, macOS, or CPU-only systems with minimal configuration.

## Basic Installation via pip

The simplest method installs the default pipeline with automatic platform detection. According to the quick-start section in [`README.md`](https://github.com/huggingface/speech-to-speech/blob/main/README.md), run:

```bash
pip install speech-to-speech

```

This command pulls the correct binary wheels for your system—whether Linux CUDA, macOS MLX, or CPU-only—and installs the core orchestration code in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). After installation, start the realtime server immediately:

```bash
export OPENAI_API_KEY=your-api-key-here  # Optional: only required for remote LLM providers

speech-to-speech

```

## CUDA-Specific Setup for Qwen3-TTS Backend

On Linux systems, the default TTS backend (`qwen3`) uses the `faster-qwen3-tts[ggml]` wheel compiled for CUDA 12.8. If your system runs a different CUDA runtime, you must install the matching `qwentts-cpp-python` wheel **before** installing the main package to avoid runtime errors.

### Matching CUDA Versions to Wheels

Install the specific wheel for your CUDA version using the Hugging Face wheels repository:

**CUDA 13.x:**

```bash
pip install "qwentts-cpp-python==0.3.1+cu130" -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu130
pip install speech-to-speech

```

**CUDA 12.4:**

```bash
pip install "qwentts-cpp-python==0.3.1+cu124" -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124
pip install speech-to-speech

```

**CPU-only systems:**

```bash
pip install "qwentts-cpp-python==0.3.1+cpu" -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cpu
pip install speech-to-speech

```

*Source: The CUDA compatibility note in the repository's [`README.md`](https://github.com/huggingface/speech-to-speech/blob/main/README.md).*

## Optional Backend Extras

The pipeline supports interchangeable STT, LLM, and TTS implementations defined in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml). These are provided as pip extras so you install only the components you need:

- **TTS backends:** `kokoro` (Kokoro-82M, Linux/Windows), `pocket` (Pocket TTS), `chattts` (ChatTTS), `facebook-mms` (MMS TTS)
- **STT backends:** `faster-whisper` (Faster Whisper), `whisper-mlx` (Lightning Whisper for macOS), `paraformer` (FunASR Paraformer)
- **macOS acceleration:** `mlx-lm` (vision model support via mlx-lm)

Install specific extras using bracket notation:

```bash
pip install "speech-to-speech[kokoro]"
pip install "speech-to-speech[faster-whisper]"
pip install "speech-to-speech[whisper-mlx]"  # macOS only

```

## Install from Source for Development

To modify the pipeline or contribute to development, clone the repository and use `uv` for an editable install. This approach creates the `speech-to-speech` CLI in your current virtual environment while referencing the source code directly:

```bash
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync

```

The `uv sync` command resolves all dependencies from [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) and installs the package in editable mode, allowing changes to [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) or argument classes in `src/speech_to_speech/arguments_classes/` to take effect immediately.

## Docker Deployment

For containerized environments, the repository includes a [`docker-compose.yml`](https://github.com/huggingface/speech-to-speech/blob/main/docker-compose.yml) that orchestrates a complete stack. First install the NVIDIA Container Toolkit, then run:

```bash
docker compose up

```

This deployment brings up a llama.cpp server, a TCP socket server, and exposes ports `8080`, `12345`, and `12346` for API access. The Docker configuration handles GPU passthrough automatically on compatible systems.

## Verify Installation and Run the Server

After installation, verify the CLI is available and start the default realtime server:

```bash

# Check installation

speech-to-speech --help

# Start the server (set API key only if using remote LLM providers)

export OPENAI_API_KEY=sk-...
speech-to-speech

```

To test the pipeline, use the provided client script that streams microphone audio to the server and plays back synthesized responses:

```bash
python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765

```

## Summary

- **Standard install:** Use `pip install speech-to-speech` for automatic platform detection and binary wheel installation.
- **CUDA compatibility:** Install matching `qwentts-cpp-python` wheels before the main package if running CUDA versions other than 12.8.
- **Modular backends:** Add specific STT/TTS engines via extras like `[faster-whisper]` or `[kokoro]` defined in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml).
- **Development workflow:** Clone the repo and run `uv sync` for editable installs that track source changes in `src/speech_to_speech/`.
- **Container deployment:** Use `docker compose up` after installing NVIDIA Container Toolkit for GPU-accelerated containerized instances.

## Frequently Asked Questions

### Do I need an OpenAI API key to run Speech-to-Speech locally?

No. The `OPENAI_API_KEY` environment variable is only required if you configure the pipeline to use remote LLM providers like OpenAI GPT-4. The default configuration runs entirely locally using llama.cpp or other local model backends, allowing fully offline operation after the initial model downloads.

### Which CUDA version does the default Qwen3-TTS backend require?

The default installation targets **CUDA 12.8**. If your system runs CUDA 12.4, 13.x, or CPU-only, you must manually install the corresponding `qwentts-cpp-python` wheel (e.g., `0.3.1+cu124` or `0.3.1+cpu`) from the Hugging Face wheels repository before installing `speech-to-speech`.

### Can I install Speech-to-Speech on macOS without CUDA?

Yes. macOS installations automatically use **MLX** acceleration where available. Install the `[whisper-mlx]` extra for Lightning Whisper STT support and `[mlx-lm]` for vision model capabilities. The Qwen3-TTS backend falls back to CPU or Metal performance shaders on Apple Silicon without requiring CUDA.

### How do I add additional TTS or STT backends after the initial installation?

Run the install command again with the specific extra in brackets. For example, to add ChatTTS support to an existing installation, execute `pip install "speech-to-speech[chattts]"`. The [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) defines all optional dependencies, allowing modular expansion of the pipeline's capabilities without reinstalling the core package.