# How to Set Up the Speech-to-Speech Project Locally: A Complete Guide

> Set up speech-to-speech locally with ease. Install the PyPI package, configure your API key, and run the command to launch the WebSocket server on localhost. Follow our complete guide.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-07

---

**To set up the speech-to-speech project locally, install the PyPI package with `pip install speech-to-speech`, configure your `OPENAI_API_KEY` environment variable, and run the `speech-to-speech` command to launch the Realtime-compatible WebSocket server on `localhost:8765`.**

The **huggingface/speech-to-speech** repository implements a low-latency voice-agent pipeline that chains four interchangeable components: Voice Activity Detection (VAD), Speech-to-Text (STT), a Language Model (LLM), and Text-to-Speech (TTS). Each stage runs in separate threads and communicates through thread-safe queues, making it easy to swap backends without modifying the core orchestration logic in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

## Prerequisites

Before you begin, ensure your environment meets the following requirements:

- **Python 3.10 or higher** is required for all installation methods.
- An API key for your chosen LLM provider (OpenAI by default, or Hugging Face Inference Providers).
- A microphone and speakers for local mode testing.

## Installation Methods

You can install the speech-to-speech pipeline either from PyPI for immediate use or from source for development.

### Install from PyPI (Recommended)

The simplest way to set up the project locally is using the published wheel:

```bash
pip install speech-to-speech

```

This installs the console script `speech-to-speech` and all core dependencies.

### Build from Source (Development)

If you need to modify handlers or contribute to the repository, clone and install in editable mode:

```bash
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync

```

The `uv sync` command creates an editable install and resolves all dependencies, pointing the `speech-to-speech` command to your local checkout.

## Configure API Keys

The default LLM backend (`responses-api`) requires an OpenAI-compatible API key. Set this in your environment before launching:

```bash
export OPENAI_API_KEY=sk-...

```

If you prefer Hugging Face Inference Providers, export your token instead:

```bash
export HF_TOKEN=hf_...

```

You will pass the corresponding base URL and key via CLI flags when starting the server.

## Launch the Pipeline

### Start the Realtime Server

Running the installed CLI starts an OpenAI Realtime-compatible WebSocket server on port `8765` with the default stack (Parakeet TDT for STT, OpenAI LLM, Qwen3-TTS for synthesis):

```bash
speech-to-speech

```

This command initializes the four-stage pipeline defined in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), where the `VADHandlerArguments` class (located in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py)) configures speech detection thresholds and turn-taking behavior.

### Connect with a Client

From a second terminal, run the example client to record audio, stream it to the server, and play back the synthesized response:

```bash
python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765

```

This script handles PCM audio encoding and WebSocket framing compatible with the `/v1/realtime` endpoint.

## Customization: Swap Backends and Run Modes

The pipeline supports swapping any component via command-line flags. Below are common configurations for different hardware and privacy requirements.

### Change STT Backend

To use Faster Whisper instead of the default Parakeet TDT:

```bash
speech-to-speech --stt faster-whisper --stt_model_name large-v2

```

STT implementations reside in `src/speech_to_speech/STT/`, including Whisper, Faster Whisper, and Parakeet TDT handlers.

### Change TTS Backend

To synthesize speech locally with Pocket TTS:

```bash
speech-to-speech --tts pocket --pocket_tts_voice jean --pocket_tts_device cpu

```

Alternative TTS handlers in `src/speech_to_speech/TTS/` include Qwen3-TTS, Pocket TTS, and Kokoro.

### Use Local Mode

To bypass the WebSocket protocol and run microphone-to-speaker directly on your machine:

```bash
speech-to-speech --mode local

```

Local mode uses the same thread-queue architecture but connects audio I/O directly to the pipeline without network transport.

### macOS Optimized Settings

For Apple Silicon machines, a convenience preset applies CPU-optimized defaults:

```bash
speech-to-speech --local_mac_optimal_settings

```

## Verify Your Installation

Confirm functionality by running the test suite:

```bash
pytest
ruff check

```

All tests should pass, indicating that the VAD, STT, LLM, and TTS handlers are correctly installed and the pipeline orchestration in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) is functional.

## Summary

- **Install** the package with `pip install speech-to-speech` (Python 3.10+ required).
- **Configure** your `OPENAI_API_KEY` or `HF_TOKEN` environment variable.
- **Launch** the server with the `speech-to-speech` command, which exposes an OpenAI Realtime-compatible WebSocket API at `ws://localhost:8765/v1/realtime`.
- **Customize** any pipeline stage (STT, LLM, TTS) using CLI flags or run in `local` mode for offline microphone-to-speaker operation.
- **Develop** using an editable install via `git clone` and `uv sync` to modify handlers in `src/speech_to_speech/`.

## Frequently Asked Questions

### What Python version is required to set up the speech-to-speech project locally?

The speech-to-speech package requires **Python 3.10 or higher**. This is enforced by the package metadata and confirmed in the repository README installation instructions.

### Can I run the speech-to-speech pipeline without an internet connection?

Yes, by using local backends for all four stages. Specify `--mode local`, use `--stt faster-whisper` with a local model path, select `--llm_backend mlx-lm` or `--llm_backend transformers` with a local model (e.g., `mlx-community/Qwen3-4B-Instruct-2507-bf16`), and use `--tts pocket` or `--tts kokoro` for on-device synthesis. No API keys are required for fully offline operation.

### How do I switch from OpenAI to a Hugging Face Inference Provider for the LLM?

Set your `HF_TOKEN` environment variable, then launch with the responses-api backend pointed to the Hugging Face router:

```bash
speech-to-speech \
    --llm_backend responses-api \
    --responses_api_base_url https://router.huggingface.co/v1 \
    --responses_api_api_key "$HF_TOKEN" \
    --model_name "Qwen/Qwen3.5-9B:together"

```

### Where is the pipeline orchestration logic defined in the source code?

The core coordination of VAD → STT → LLM → TTS threads is implemented in **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)**. This file manages the thread-safe queues and component handoff, while configuration-specific arguments (such as VAD thresholds) are defined in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py).