Speech-to-Speech System Requirements: Complete Hardware & Software Guide
The Hugging Face speech-to-speech repository requires Python 3.10+, runs on CPU or GPU, and automatically selects CUDA or Apple Silicon backends based on your platform.
The speech-to-speech repository is a pure-Python voice agent pipeline that adapts its dependencies to your hardware. This guide covers every system requirement—from Python version to GPU acceleration—based on the official source code and documentation.
Python Version Requirement
The project requires Python 3.10 or newer. This is enforced across all installation methods.
# Verify your Python version
python --version # Must output 3.10.x or higher
No additional system packages are required. All core components install through standard pip commands.
Core Dependencies & Installation
The base pipeline installs with a single command. Platform-specific wheels (including CUDA-enabled libraries) are selected automatically via pyproject.toml markers.
# Install the core speech-to-speech pipeline
pip install speech-to-speech
# Set required API key for default LLM backend
export OPENAI_API_KEY=your_key_here
# Start the realtime WebSocket server
speech-to-speech # Runs on ws://localhost:8765/v1/realtime
Core components include:
- VAD (Voice Activity Detection): Silero VAD via
vad_handler.py - STT (Speech-to-Text): Default Whisper-based backend
- LLM: OpenAI API client (local backends available via extras)
- TTS (Text-to-Speech): Qwen3-TTS GGML backend
CUDA & GPU Requirements
Default GPU Backend (Qwen3-TTS GGML)
The default TTS backend uses CUDA 12.8 wheels automatically. For optimal performance, match this CUDA version or install a compatible wheel manually.
| CUDA Version | Installation Command |
|---|---|
| CUDA 12.8 (default) | pip install speech-to-speech |
| CUDA 13.x | Manual wheel install (see below) |
| CUDA 12.4 | Manual wheel install |
| CPU-only | Works without GPU |
Installing Custom CUDA Versions
For CUDA versions other than 12.8, install the matching qwentts-cpp-python wheel before the main package:
# Example: CUDA 13.x
pip install "qwentts-cpp-python==0.3.1+cu130" \
-f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu130
# Then install the main package
pip install speech-to-speech
Replace cu130 with your CUDA version (cu124, cu128, etc.).
GPU Recommendations
- CUDA-compatible NVIDIA GPU: Required for accelerated Qwen3-TTS and Faster-Whisper STT
- VRAM: Varies by model size; 8GB+ recommended for larger TTS models
- CPU fallback: Functional but significantly slower for real-time inference
Apple Silicon & macOS Requirements
macOS systems use MLX-based backends automatically, optimized for Apple Silicon (M1/M2/M3).
# Run with macOS-optimized settings
speech-to-speech --local_mac_optimal_settings
The pyproject.toml platform markers automatically select:
mlx-audiofor audio processingmlx-lmfor local LLM inference (install via extras)whisper-mlxfor STT (install via extras)
Key macOS files:
s2s_pipeline.pyorchestrates the pipeline- Argument classes in
arguments_classes/expose--local_mac_optimal_settings
Optional Backend Extras
Install additional STT, LLM, and TTS engines via pip extras. These are defined in README.md, lines 124–136.
| Extra | Backend | Use Case |
|---|---|---|
kokoro |
Kokoro TTS | Lightweight neural TTS |
pocket |
Pocket TTS | CPU/CUDA neural TTS |
chattts |
ChatTTS | Conversational TTS |
facebook-mms |
Meta MMS | Multilingual STT/TTS |
faster-whisper |
Faster Whisper | Optimized Whisper STT |
whisper-mlx |
Whisper MLX | Apple Silicon STT |
paraformer |
Paraformer | Streaming STT |
mlx-lm |
MLX-LM | Local LLM on Apple Silicon |
# Install multiple extras
pip install "speech-to-speech[pocket,faster-whisper]"
pip install "speech-to-speech[mlx-lm,whisper-mlx]" # macOS
DeepFilterNet Audio Enhancement Conflict
The optional DeepFilterNet noise suppression filter has a dependency conflict:
- Requires
numpy < 2 - Incompatible with Pocket TTS (requires
numpy >= 2)
# Only install DeepFilterNet if NOT using Pocket TTS
pip install "speech-to-speech[deepfilternet]"
Enable in the VAD handler via configuration flags in VAD/vad_handler.py.
Hardware Configuration Summary
| Configuration | Requirements | Performance |
|---|---|---|
| Minimal (CPU) | Python 3.10+, 4GB RAM | Functional, higher latency |
| Recommended (CUDA) | Python 3.10+, CUDA 12.8 GPU, 8GB+ VRAM | Real-time voice conversation |
| Apple Silicon | Python 3.10+, M1/M2/M3 Mac | Optimized MLX acceleration |
| Development | All above + optional extras | Full backend flexibility |
Key Configuration Files
Understanding these source files helps troubleshoot system requirements:
README.md(lines 87–142): Installation and platform-specific notespyproject.toml: Platform markers for automatic wheel selectionsrc/speech_to_speech/s2s_pipeline.py: Pipeline orchestration showing component wiringsrc/speech_to_speech/arguments_classes/: CLI argument definitions for each backendsrc/speech_to_speech/VAD/vad_handler.py: VAD and DeepFilterNet integration
Summary
- Python 3.10+ is mandatory for all installations
- CUDA 12.8 is the default GPU target; other CUDA versions require manual wheel installation
- Apple Silicon uses MLX backends automatically—no CUDA needed
- CPU-only works but sacrifices real-time performance
- Optional extras extend STT/LLM/TTS capabilities without breaking core functionality
- DeepFilterNet conflicts with Pocket TTS due to NumPy version requirements
Frequently Asked Questions
Can I run speech-to-speech without a GPU?
Yes. The pipeline runs on CPU-only systems out of the box, though latency increases significantly. Install with pip install speech-to-speech and the GGML backends will use CPU inference. For better CPU performance, consider the faster-whisper or paraformer extras for STT.
What if my CUDA version doesn't match the default?
Install a matching qwentts-cpp-python wheel from the Hugging Face wheelhouse before installing the main package. The README.md provides examples for CUDA 13.x and 12.4. Mismatched CUDA versions will cause TTS backend initialization failures.
Is Apple Silicon fully supported?
Yes. macOS systems automatically receive MLX-optimized audio processing. Use --local_mac_optimal_settings for best performance. Install whisper-mlx and mlx-lm extras to replace cloud LLM and default STT with local Apple Silicon alternatives.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →