System Requirements for Running Speech-to-Speech Models: Hardware, Software, and Setup Guide
The Hugging Face speech-to-speech pipeline requires Python 3.10 or newer, runs on macOS and Linux/Windows, and needs either CUDA 12 for GPU-accelerated TTS on Linux or Apple Silicon with MLX for optimized macOS inference.
The huggingface/speech-to-speech repository delivers a low-latency, modular voice agent that chains Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) components. Before installing, verify that your system meets the hardware and software dependencies declared in pyproject.toml and documented in the README.
Python Version and Core Dependencies
The package requires Python 3.10, 3.11, or 3.12, as specified by the requires-python field in pyproject.toml (line 10). Installation on older Python versions will fail.
Mandatory runtime dependencies declared in pyproject.toml (lines 27-49) include:
torchandtorchaudiofor neural network operationstransformersfor LLM and model managementfastapi,uvicorn, andwebsocketsfor the OpenAI Realtime-compatible API serversounddeviceandsoundfilefor cross-platform audio I/Onano-parakeetfor the default STT backendfaster-qwen3-tts[ggml]for the default TTS backendlingua-language-detectorfor automatic language detection
Operating System Support
The pipeline supports macOS (Darwin), Linux, and Windows. Platform-specific wheels are selected automatically via environment markers in pyproject.toml (lines 31-46).
- macOS: Uses the MLX stack (
mlx,mlx-audio,mlx-lm,mlx-metal) andminiaudiofor optimized Apple Silicon inference, as defined inpyproject.toml(lines 52-57). - Linux/Windows: Uses standard PyTorch CUDA wheels and
sounddevicefor audio capture.
Hardware Requirements by Component
Each pipeline component in src/speech_to_speech/s2s_pipeline.py has distinct hardware needs:
Voice Activity Detection (VAD)
- Runs on CPU only using Silero VAD. No GPU required.
Speech-to-Text (STT)
- Parakeet-TDT (default): Benefits from CUDA GPU on Linux; uses Apple Silicon via MLX on macOS.
- Faster Whisper: Requires GPU for real-time performance.
- Paraformer: GPU recommended for low latency.
Large Language Model (LLM)
- Most compute-intensive component. Requires CUDA GPU for large transformer models on Linux, or Apple Silicon + MLX on macOS.
- CPU fallback available for smaller models only.
Text-to-Speech (TTS)
- Qwen3-TTS GGML: Requires CUDA 12 runtime on Linux (specifically CUDA 12.8 as per the wheel requirements).
- MLX Audio: Uses Apple Silicon Metal Performance Shaders on macOS.
- Pocket TTS: CPU-only fallback option.
CUDA Requirements for Linux GPU Acceleration
For the default Qwen3-TTS GGML backend on Linux, you must have CUDA 12 installed. The wheel targets CUDA 12.8 specifically. If your system runs a different CUDA version (e.g., 12.4), install the matching wheel from the Hugging Face wheelhouse as documented in the README:
pip install "qwentts-cpp-python==0.3.0+cu124" \
-f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124
Mismatched CUDA versions will cause runtime errors when initializing the TTS backend.
macOS and Apple Silicon Optimization
On macOS, the pipeline automatically utilizes the MLX stack when you install the package. The pyproject.toml (lines 52-57) declares platform-specific dependencies including mlx, mlx-audio, and mlx-lm.
For optimal performance, use the dedicated flag:
pip install "speech-to-speech[whisper-mlx]"
speech-to-speech --local_mac_optimal_settings
This configures Metal GPU usage (--device mps) and selects MLX-based LLM and TTS backends automatically via the argument classes in src/speech_to_speech/arguments_classes/.
Optional Backend Dependencies
The optional-dependencies section in pyproject.toml (lines 63-97) defines extras for alternative backends:
speech-to-speech[kokoro]– Kokoro-82M TTS (CUDA/CPU on non-macOS, built-in on macOS)speech-to-speech[pocket]– Pocket TTS CPU-only backendspeech-to-speech[chattts]– ChatTTS backendspeech-to-speech[facebook-mms]– MMS TTS backendspeech-to-speech[faster-whisper]– Faster Whisper STTspeech-to-speech[whisper-mlx]– Lightning Whisper MLX for macOSspeech-to-speech[paraformer]– Paraformer STT backend
Install only the extras you need to minimize dependency overhead.
Installation and Verification Examples
Basic Installation (CPU/GPU Auto-detect)
# Requires Python 3.10+
pip install speech-to-speech
# Set API key for OpenAI LLM backend
export OPENAI_API_KEY=your-key-here
# Launch WebSocket server (ws://localhost:8765/v1/realtime)
speech-to-speech
macOS Apple Silicon Setup
# Install with MLX Whisper support
pip install "speech-to-speech[whisper-mlx]"
# Run with Metal optimization
speech-to-speech --local_mac_optimal_settings
Custom Linux GPU Configuration
# Install specific CUDA wheel if needed (example for CUDA 12.4)
pip install "qwentts-cpp-python==0.3.0+cu124" \
-f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124
# Run with specific backends
speech-to-speech \
--stt faster-whisper \
--tts qwen3 \
--qwen3_tts_backend ggml \
--model_name "gpt-4o-mini"
Summary
- Python: Version 3.10, 3.11, or 3.12 is mandatory according to
pyproject.toml. - OS: macOS (Darwin), Linux, and Windows supported with platform-specific wheels.
- GPU: CUDA 12 required for Qwen3-TTS GGML on Linux; Apple Silicon recommended for macOS MLX backends.
- Audio:
sounddeviceandsoundfilehandle I/O across platforms. - Modularity: Install optional extras (
[kokoro],[faster-whisper], etc.) only for backends you intend to use.
Frequently Asked Questions
Do I need a GPU to run the speech-to-speech pipeline?
No, but it is strongly recommended for real-time performance. The VAD component runs on CPU, but STT, LLM, and TTS components benefit significantly from GPU acceleration. You can run smaller models on CPU using the Pocket TTS backend or smaller LLM variants, though latency will increase substantially.
Can I run this on Windows?
Yes, the repository supports Windows, Linux, and macOS. Windows installations use the same pip install speech-to-speech command, but ensure you have the appropriate audio drivers for sounddevice to function correctly. Platform-specific logic in pyproject.toml handles dependency resolution automatically.
How do I resolve CUDA version mismatches with Qwen3-TTS?
If you encounter errors loading the Qwen3-TTS GGML backend, verify your CUDA version with nvcc --version. The default wheel requires CUDA 12.8. For other versions (e.g., 12.4), install the matching wheel from the Hugging Face wheelhouse as shown in the installation examples above, or use the CPU fallback wheels documented in the README.
What is the minimum RAM requirement?
While not explicitly stated in the configuration files, running the full pipeline with default models (Parakeet-TDT, GPT-4o-mini, Qwen3-TTS) requires approximately 8GB of system RAM for CPU-only operation, and 4GB+ VRAM for GPU-accelerated inference. Larger LLM models require proportionally more memory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →