How to Set Up the Speech-to-Speech Project Locally: A Complete Guide

To set up the speech-to-speech project locally, install the PyPI package with pip install speech-to-speech, configure your OPENAI_API_KEY environment variable, and run the speech-to-speech command to launch the Realtime-compatible WebSocket server on localhost:8765.

The huggingface/speech-to-speech repository implements a low-latency voice-agent pipeline that chains four interchangeable components: Voice Activity Detection (VAD), Speech-to-Text (STT), a Language Model (LLM), and Text-to-Speech (TTS). Each stage runs in separate threads and communicates through thread-safe queues, making it easy to swap backends without modifying the core orchestration logic in src/speech_to_speech/s2s_pipeline.py.

Prerequisites

Before you begin, ensure your environment meets the following requirements:

  • Python 3.10 or higher is required for all installation methods.
  • An API key for your chosen LLM provider (OpenAI by default, or Hugging Face Inference Providers).
  • A microphone and speakers for local mode testing.

Installation Methods

You can install the speech-to-speech pipeline either from PyPI for immediate use or from source for development.

The simplest way to set up the project locally is using the published wheel:

pip install speech-to-speech

This installs the console script speech-to-speech and all core dependencies.

Build from Source (Development)

If you need to modify handlers or contribute to the repository, clone and install in editable mode:

git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync

The uv sync command creates an editable install and resolves all dependencies, pointing the speech-to-speech command to your local checkout.

Configure API Keys

The default LLM backend (responses-api) requires an OpenAI-compatible API key. Set this in your environment before launching:

export OPENAI_API_KEY=sk-...

If you prefer Hugging Face Inference Providers, export your token instead:

export HF_TOKEN=hf_...

You will pass the corresponding base URL and key via CLI flags when starting the server.

Launch the Pipeline

Start the Realtime Server

Running the installed CLI starts an OpenAI Realtime-compatible WebSocket server on port 8765 with the default stack (Parakeet TDT for STT, OpenAI LLM, Qwen3-TTS for synthesis):

speech-to-speech

This command initializes the four-stage pipeline defined in src/speech_to_speech/s2s_pipeline.py, where the VADHandlerArguments class (located in src/speech_to_speech/arguments_classes/vad_arguments.py) configures speech detection thresholds and turn-taking behavior.

Connect with a Client

From a second terminal, run the example client to record audio, stream it to the server, and play back the synthesized response:

python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765

This script handles PCM audio encoding and WebSocket framing compatible with the /v1/realtime endpoint.

Customization: Swap Backends and Run Modes

The pipeline supports swapping any component via command-line flags. Below are common configurations for different hardware and privacy requirements.

Change STT Backend

To use Faster Whisper instead of the default Parakeet TDT:

speech-to-speech --stt faster-whisper --stt_model_name large-v2

STT implementations reside in src/speech_to_speech/STT/, including Whisper, Faster Whisper, and Parakeet TDT handlers.

Change TTS Backend

To synthesize speech locally with Pocket TTS:

speech-to-speech --tts pocket --pocket_tts_voice jean --pocket_tts_device cpu

Alternative TTS handlers in src/speech_to_speech/TTS/ include Qwen3-TTS, Pocket TTS, and Kokoro.

Use Local Mode

To bypass the WebSocket protocol and run microphone-to-speaker directly on your machine:

speech-to-speech --mode local

Local mode uses the same thread-queue architecture but connects audio I/O directly to the pipeline without network transport.

macOS Optimized Settings

For Apple Silicon machines, a convenience preset applies CPU-optimized defaults:

speech-to-speech --local_mac_optimal_settings

Verify Your Installation

Confirm functionality by running the test suite:

pytest
ruff check

All tests should pass, indicating that the VAD, STT, LLM, and TTS handlers are correctly installed and the pipeline orchestration in s2s_pipeline.py is functional.

Summary

  • Install the package with pip install speech-to-speech (Python 3.10+ required).
  • Configure your OPENAI_API_KEY or HF_TOKEN environment variable.
  • Launch the server with the speech-to-speech command, which exposes an OpenAI Realtime-compatible WebSocket API at ws://localhost:8765/v1/realtime.
  • Customize any pipeline stage (STT, LLM, TTS) using CLI flags or run in local mode for offline microphone-to-speaker operation.
  • Develop using an editable install via git clone and uv sync to modify handlers in src/speech_to_speech/.

Frequently Asked Questions

What Python version is required to set up the speech-to-speech project locally?

The speech-to-speech package requires Python 3.10 or higher. This is enforced by the package metadata and confirmed in the repository README installation instructions.

Can I run the speech-to-speech pipeline without an internet connection?

Yes, by using local backends for all four stages. Specify --mode local, use --stt faster-whisper with a local model path, select --llm_backend mlx-lm or --llm_backend transformers with a local model (e.g., mlx-community/Qwen3-4B-Instruct-2507-bf16), and use --tts pocket or --tts kokoro for on-device synthesis. No API keys are required for fully offline operation.

How do I switch from OpenAI to a Hugging Face Inference Provider for the LLM?

Set your HF_TOKEN environment variable, then launch with the responses-api backend pointed to the Hugging Face router:

speech-to-speech \
    --llm_backend responses-api \
    --responses_api_base_url https://router.huggingface.co/v1 \
    --responses_api_api_key "$HF_TOKEN" \
    --model_name "Qwen/Qwen3.5-9B:together"

Where is the pipeline orchestration logic defined in the source code?

The core coordination of VAD → STT → LLM → TTS threads is implemented in src/speech_to_speech/s2s_pipeline.py. This file manages the thread-safe queues and component handoff, while configuration-specific arguments (such as VAD thresholds) are defined in src/speech_to_speech/arguments_classes/vad_arguments.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →