How to Set Up the Speech-to-Speech Project Locally: A Complete Guide
To set up the speech-to-speech project locally, install the PyPI package with pip install speech-to-speech, configure your OPENAI_API_KEY environment variable, and run the speech-to-speech command to launch the Realtime-compatible WebSocket server on localhost:8765.
The huggingface/speech-to-speech repository implements a low-latency voice-agent pipeline that chains four interchangeable components: Voice Activity Detection (VAD), Speech-to-Text (STT), a Language Model (LLM), and Text-to-Speech (TTS). Each stage runs in separate threads and communicates through thread-safe queues, making it easy to swap backends without modifying the core orchestration logic in src/speech_to_speech/s2s_pipeline.py.
Prerequisites
Before you begin, ensure your environment meets the following requirements:
- Python 3.10 or higher is required for all installation methods.
- An API key for your chosen LLM provider (OpenAI by default, or Hugging Face Inference Providers).
- A microphone and speakers for local mode testing.
Installation Methods
You can install the speech-to-speech pipeline either from PyPI for immediate use or from source for development.
Install from PyPI (Recommended)
The simplest way to set up the project locally is using the published wheel:
pip install speech-to-speech
This installs the console script speech-to-speech and all core dependencies.
Build from Source (Development)
If you need to modify handlers or contribute to the repository, clone and install in editable mode:
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync
The uv sync command creates an editable install and resolves all dependencies, pointing the speech-to-speech command to your local checkout.
Configure API Keys
The default LLM backend (responses-api) requires an OpenAI-compatible API key. Set this in your environment before launching:
export OPENAI_API_KEY=sk-...
If you prefer Hugging Face Inference Providers, export your token instead:
export HF_TOKEN=hf_...
You will pass the corresponding base URL and key via CLI flags when starting the server.
Launch the Pipeline
Start the Realtime Server
Running the installed CLI starts an OpenAI Realtime-compatible WebSocket server on port 8765 with the default stack (Parakeet TDT for STT, OpenAI LLM, Qwen3-TTS for synthesis):
speech-to-speech
This command initializes the four-stage pipeline defined in src/speech_to_speech/s2s_pipeline.py, where the VADHandlerArguments class (located in src/speech_to_speech/arguments_classes/vad_arguments.py) configures speech detection thresholds and turn-taking behavior.
Connect with a Client
From a second terminal, run the example client to record audio, stream it to the server, and play back the synthesized response:
python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765
This script handles PCM audio encoding and WebSocket framing compatible with the /v1/realtime endpoint.
Customization: Swap Backends and Run Modes
The pipeline supports swapping any component via command-line flags. Below are common configurations for different hardware and privacy requirements.
Change STT Backend
To use Faster Whisper instead of the default Parakeet TDT:
speech-to-speech --stt faster-whisper --stt_model_name large-v2
STT implementations reside in src/speech_to_speech/STT/, including Whisper, Faster Whisper, and Parakeet TDT handlers.
Change TTS Backend
To synthesize speech locally with Pocket TTS:
speech-to-speech --tts pocket --pocket_tts_voice jean --pocket_tts_device cpu
Alternative TTS handlers in src/speech_to_speech/TTS/ include Qwen3-TTS, Pocket TTS, and Kokoro.
Use Local Mode
To bypass the WebSocket protocol and run microphone-to-speaker directly on your machine:
speech-to-speech --mode local
Local mode uses the same thread-queue architecture but connects audio I/O directly to the pipeline without network transport.
macOS Optimized Settings
For Apple Silicon machines, a convenience preset applies CPU-optimized defaults:
speech-to-speech --local_mac_optimal_settings
Verify Your Installation
Confirm functionality by running the test suite:
pytest
ruff check
All tests should pass, indicating that the VAD, STT, LLM, and TTS handlers are correctly installed and the pipeline orchestration in s2s_pipeline.py is functional.
Summary
- Install the package with
pip install speech-to-speech(Python 3.10+ required). - Configure your
OPENAI_API_KEYorHF_TOKENenvironment variable. - Launch the server with the
speech-to-speechcommand, which exposes an OpenAI Realtime-compatible WebSocket API atws://localhost:8765/v1/realtime. - Customize any pipeline stage (STT, LLM, TTS) using CLI flags or run in
localmode for offline microphone-to-speaker operation. - Develop using an editable install via
git cloneanduv syncto modify handlers insrc/speech_to_speech/.
Frequently Asked Questions
What Python version is required to set up the speech-to-speech project locally?
The speech-to-speech package requires Python 3.10 or higher. This is enforced by the package metadata and confirmed in the repository README installation instructions.
Can I run the speech-to-speech pipeline without an internet connection?
Yes, by using local backends for all four stages. Specify --mode local, use --stt faster-whisper with a local model path, select --llm_backend mlx-lm or --llm_backend transformers with a local model (e.g., mlx-community/Qwen3-4B-Instruct-2507-bf16), and use --tts pocket or --tts kokoro for on-device synthesis. No API keys are required for fully offline operation.
How do I switch from OpenAI to a Hugging Face Inference Provider for the LLM?
Set your HF_TOKEN environment variable, then launch with the responses-api backend pointed to the Hugging Face router:
speech-to-speech \
--llm_backend responses-api \
--responses_api_base_url https://router.huggingface.co/v1 \
--responses_api_api_key "$HF_TOKEN" \
--model_name "Qwen/Qwen3.5-9B:together"
Where is the pipeline orchestration logic defined in the source code?
The core coordination of VAD → STT → LLM → TTS threads is implemented in src/speech_to_speech/s2s_pipeline.py. This file manages the thread-safe queues and component handoff, while configuration-specific arguments (such as VAD thresholds) are defined in src/speech_to_speech/arguments_classes/vad_arguments.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →