# Speech-to-Speech System Requirements: Complete Hardware & Software Guide

> Discover Hugging Face speech-to-speech system requirements. Learn about Python 3.10+, CPU/GPU compatibility, and automatic CUDA/Apple Silicon backend selection for your platform.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: getting-started
- Published: 2026-08-01

---

**The Hugging Face speech-to-speech repository requires Python 3.10+, runs on CPU or GPU, and automatically selects CUDA or Apple Silicon backends based on your platform.**

The `speech-to-speech` repository is a pure-Python voice agent pipeline that adapts its dependencies to your hardware. This guide covers every system requirement—from Python version to GPU acceleration—based on the official source code and documentation.

## Python Version Requirement

The project requires **Python 3.10 or newer**. This is enforced across all installation methods.

```bash

# Verify your Python version

python --version  # Must output 3.10.x or higher

```

No additional system packages are required. All core components install through standard `pip` commands.

## Core Dependencies & Installation

The base pipeline installs with a single command. Platform-specific wheels (including CUDA-enabled libraries) are selected automatically via [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) markers.

```bash

# Install the core speech-to-speech pipeline

pip install speech-to-speech

# Set required API key for default LLM backend

export OPENAI_API_KEY=your_key_here

# Start the realtime WebSocket server

speech-to-speech  # Runs on ws://localhost:8765/v1/realtime

```

Core components include:
- **VAD (Voice Activity Detection)**: Silero VAD via [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)
- **STT (Speech-to-Text)**: Default Whisper-based backend
- **LLM**: OpenAI API client (local backends available via extras)
- **TTS (Text-to-Speech)**: Qwen3-TTS GGML backend

## CUDA & GPU Requirements

### Default GPU Backend (Qwen3-TTS GGML)

The default TTS backend uses **CUDA 12.8** wheels automatically. For optimal performance, match this CUDA version or install a compatible wheel manually.

| CUDA Version | Installation Command |
|-------------|----------------------|
| CUDA 12.8 (default) | `pip install speech-to-speech` |
| CUDA 13.x | Manual wheel install (see below) |
| CUDA 12.4 | Manual wheel install |
| CPU-only | Works without GPU |

### Installing Custom CUDA Versions

For CUDA versions other than 12.8, install the matching `qwentts-cpp-python` wheel before the main package:

```bash

# Example: CUDA 13.x

pip install "qwentts-cpp-python==0.3.1+cu130" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu130

# Then install the main package

pip install speech-to-speech

```

Replace `cu130` with your CUDA version (`cu124`, `cu128`, etc.).

### GPU Recommendations

- **CUDA-compatible NVIDIA GPU**: Required for accelerated Qwen3-TTS and Faster-Whisper STT
- **VRAM**: Varies by model size; 8GB+ recommended for larger TTS models
- **CPU fallback**: Functional but significantly slower for real-time inference

## Apple Silicon & macOS Requirements

macOS systems use **MLX-based backends** automatically, optimized for Apple Silicon (M1/M2/M3).

```bash

# Run with macOS-optimized settings

speech-to-speech --local_mac_optimal_settings

```

The [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) platform markers automatically select:
- `mlx-audio` for audio processing
- `mlx-lm` for local LLM inference (install via extras)
- `whisper-mlx` for STT (install via extras)

Key macOS files:
- [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) orchestrates the pipeline
- Argument classes in `arguments_classes/` expose `--local_mac_optimal_settings`

## Optional Backend Extras

Install additional STT, LLM, and TTS engines via pip extras. These are defined in [`README.md`](https://github.com/huggingface/speech-to-speech/blob/main/README.md), lines 124–136.

| Extra | Backend | Use Case |
|-------|---------|----------|
| `kokoro` | Kokoro TTS | Lightweight neural TTS |
| `pocket` | Pocket TTS | CPU/CUDA neural TTS |
| `chattts` | ChatTTS | Conversational TTS |
| `facebook-mms` | Meta MMS | Multilingual STT/TTS |
| `faster-whisper` | Faster Whisper | Optimized Whisper STT |
| `whisper-mlx` | Whisper MLX | Apple Silicon STT |
| `paraformer` | Paraformer | Streaming STT |
| `mlx-lm` | MLX-LM | Local LLM on Apple Silicon |

```bash

# Install multiple extras

pip install "speech-to-speech[pocket,faster-whisper]"
pip install "speech-to-speech[mlx-lm,whisper-mlx]"  # macOS

```

## DeepFilterNet Audio Enhancement Conflict

The optional **DeepFilterNet** noise suppression filter has a dependency conflict:

- Requires `numpy < 2`
- **Incompatible with Pocket TTS** (requires `numpy >= 2`)

```bash

# Only install DeepFilterNet if NOT using Pocket TTS

pip install "speech-to-speech[deepfilternet]"

```

Enable in the VAD handler via configuration flags in [`VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/VAD/vad_handler.py).

## Hardware Configuration Summary

| Configuration | Requirements | Performance |
|-------------|--------------|-------------|
| **Minimal (CPU)** | Python 3.10+, 4GB RAM | Functional, higher latency |
| **Recommended (CUDA)** | Python 3.10+, CUDA 12.8 GPU, 8GB+ VRAM | Real-time voice conversation |
| **Apple Silicon** | Python 3.10+, M1/M2/M3 Mac | Optimized MLX acceleration |
| **Development** | All above + optional extras | Full backend flexibility |

## Key Configuration Files

Understanding these source files helps troubleshoot system requirements:

- **[`README.md`](https://github.com/huggingface/speech-to-speech/blob/main/README.md)** (lines 87–142): Installation and platform-specific notes
- **[`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml)**: Platform markers for automatic wheel selection
- **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)**: Pipeline orchestration showing component wiring
- **`src/speech_to_speech/arguments_classes/`**: CLI argument definitions for each backend
- **[`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py)**: VAD and DeepFilterNet integration

## Summary

- **Python 3.10+** is mandatory for all installations
- **CUDA 12.8** is the default GPU target; other CUDA versions require manual wheel installation
- **Apple Silicon** uses MLX backends automatically—no CUDA needed
- **CPU-only** works but sacrifices real-time performance
- **Optional extras** extend STT/LLM/TTS capabilities without breaking core functionality
- **DeepFilterNet conflicts** with Pocket TTS due to NumPy version requirements

## Frequently Asked Questions

### Can I run speech-to-speech without a GPU?

Yes. The pipeline runs on CPU-only systems out of the box, though latency increases significantly. Install with `pip install speech-to-speech` and the GGML backends will use CPU inference. For better CPU performance, consider the `faster-whisper` or `paraformer` extras for STT.

### What if my CUDA version doesn't match the default?

Install a matching `qwentts-cpp-python` wheel from the Hugging Face wheelhouse before installing the main package. The [`README.md`](https://github.com/huggingface/speech-to-speech/blob/main/README.md) provides examples for CUDA 13.x and 12.4. Mismatched CUDA versions will cause TTS backend initialization failures.

### Is Apple Silicon fully supported?

Yes. macOS systems automatically receive MLX-optimized audio processing. Use `--local_mac_optimal_settings` for best performance. Install `whisper-mlx` and `mlx-lm` extras to replace cloud LLM and default STT with local Apple Silicon alternatives.