# Dependencies for Hugging Face Speech-to-Speech: Complete Installation Guide

> Install Hugging Face speech-to-speech effortlessly. Discover core dependencies like FastAPI and PyTorch, plus platform-specific ML libraries for macOS and optional extras for advanced TTS/STT.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: getting-started
- Published: 2026-08-02

---

**The speech-to-speech package declares all runtime requirements in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml), with core dependencies including FastAPI, PyTorch, and platform-specific ML libraries like MLX on macOS, plus optional extras for specialized TTS and STT backends.**

The Hugging Face **speech-to-speech** repository provides a modular toolkit for building real-time voice-to-voice AI pipelines. Understanding the dependencies for Hugging Face Speech-to-Speech is essential before deployment, as the package uses platform-specific constraints and optional extras to support diverse hardware configurations ranging from Apple Silicon to CUDA-enabled GPUs.

## Core Runtime Dependencies

The mandatory dependencies are defined in the `dependencies` key of **[`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml)** (lines 27-62). These packages are required for all installations regardless of platform.

### Cross-Platform Requirements

Every installation requires these base packages:

- **FastAPI** (`>=0.115.0`) and **Uvicorn** (`>=0.30.0`) for the web server layer
- **HTTPX** (`>=0.28.0`) for async HTTP client functionality
- **OpenAI** (`==2.28.0`) for API-compatible LLM interactions
- **Pydantic** (`>=2.0`) for data validation
- **Pillow** (`>=10.0.0`) for image processing utilities
- **Rich** (`>=13.0`) for terminal formatting
- **NLTK** (`==3.10.0`) for natural language processing
- **SciPy** (`>=1.10.0`) for audio signal processing

### Platform-Specific Constraints

The package declares different versions for **Darwin** (macOS) versus Linux and Windows systems:

| Package | macOS (Darwin) | Linux/Windows |
|---------|---------------|---------------|
| **NumPy** | `>=1.26.0,<2.4.4` | `>=1.26.0` |
| **SoundDevice** | `==0.5.3` | `>=0.5.0` |
| **SoundFile** | `>=0.13.0` | Not constrained |
| **Torch** | `==2.11.0` | `>=2.4.0` |
| **TorchAudio** | `==2.11.0` | `>=2.4.0` |
| **Transformers** | `==5.6.2` | `>=4.57.0` |

### macOS-Specific ML Stack

Darwin installations include additional Apple Silicon optimizations:

- **MLX** (`==0.31.1`), **MLX-Audio** (`==0.4.2`), **MLX-LM** (`==0.31.1`), and **MLX-Metal** (`==0.31.1`)
- **Miniaudio** (`==1.61`)
- **Misaki** (`>=0.9.4`)
- **EspeakNG-Loader** (`>=0.2.4`)
- **SpaCy** (`>=3.8.4`)
- **Phonemizer-Fork** (`>=3.3.2`)

### Linux-Specific Additions

Non-Darwin platforms receive:

- **Nano-Parakeet** (`>=0.2.0`)
- **Faster-Qwen3-TTS** with GGML support (`>=0.3.2`) on non-Windows systems, or the base package on Windows

## Optional Dependency Groups

The repository defines extras under `optional-dependencies` for specific backends. Install these using `pip install speech-to-speech[extra-name]`.

### Speech-to-Text Backends

- **Faster Whisper** (`faster-whisper>=1.0.3`): Optimized Whisper implementation for CPU/GPU
- **Whisper-MLX** (`lightning-whisper-mlx>=0.0.10`): Apple Silicon optimized Whisper, available via `whisper-mlx` extra

### Text-to-Speech Engines

- **ChatTTS** (`ChatTTS>=0.1.1`): Conversational TTS model via `chattts` extra
- **Kokoro** (`kokoro>=0.9.2`): Lightweight TTS for non-Darwin systems via `kokoro` extra
- **Pocket TTS** (`pocket-tts>=0.1.0`): Minimal TTS implementation via `pocket` extra

### Model Support and Utilities

- **Facebook MMS** (`transformers>=4.57.0`): Additional transformer support via `facebook-mms` extra
- **MLX LLM** (`mlx-lm==0.31.1`, `mlx-vlm==0.4.1`): Vision-language model support on Darwin via `mlx-lm` extra
- **Paraformer** (`funasr>=1.1.6`, `modelscope>=1.17.1`, `onnxruntime<1.24` for Python < 3.11): Alibaba's Paraformer model via `paraformer` extra
- **WebRTC** (`aiortc>=1.9.0`): Real-time communication support via `webrtc` extra

### Language Detection

- **Lingua** (`lingua-language-detector>=2.0.2`): Automatic language identification available via `language-detection` extra

## Installing Speech-to-Speech

Use these commands to install the package with your required configuration:

```bash

# Core installation only

pip install speech-to-speech

# With ChatTTS support

pip install "speech-to-speech[chattts]"

# With Faster Whisper STT

pip install "speech-to-speech[faster-whisper]"

# Multiple extras (e.g., Kokoro TTS and Pocket TTS)

pip install "speech-to-speech[kokoro,pocket]"

```

## Where Dependencies Are Used in the Source

The declared dependencies power specific handlers throughout the codebase:

- **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)**: The main `SpeechToSpeechPipeline` class wires together STT, LLM, and TTS components using **Torch**, **Transformers**, and **Pydantic** for configuration validation.

- **`src/speech_to_speech/TTS/*_handler.py`**: Individual TTS implementations like [`chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/chatTTS_handler.py) import **ChatTTS**, while [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py) uses the **pocket-tts** package.

- **`src/speech_to_speech/STT/*_handler.py`**: [`faster_whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/faster_whisper_handler.py) requires the **faster-whisper** package, whereas [`whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_handler.py) uses the base **Transformers** and **Torch** stack.

- **`src/speech_to_speech/LLM/*_language_model.py`**: Language model handlers utilize **OpenAI** for API-compatible endpoints and **MLX-LM** on Darwin platforms for local inference.

## Summary

- **Core dependencies** are defined in **[`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml)** and include **FastAPI**, **PyTorch**, **Transformers**, and **NLTK** with strict platform constraints for macOS.
- **macOS** installations include **MLX** optimizations and audio processing libraries like **Miniaudio** and **SpaCy**.
- **Nine optional extras** provide specialized backends including **Faster Whisper**, **ChatTTS**, **Kokoro**, and **WebRTC**.
- Install extras using bracket syntax: `pip install "speech-to-speech[extra1,extra2]"`.
- The dependency architecture maps directly to handler classes in the **`TTS/`**, **`STT/`**, and **`LLM/`** directories.

## Frequently Asked Questions

### What is the minimum PyTorch version required for Linux?

According to the **[`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml)** configuration, Linux and Windows systems require **Torch** `>=2.4.0` and **TorchAudio** `>=2.4.0`, while macOS (Darwin) uses pinned versions `==2.11.0` for both packages.

### Can I run Speech-to-Speech without installing the optional extras?

Yes. The core package provides full pipeline functionality using the base **Transformers** library and **OpenAI**-compatible APIs. Optional extras like **Faster Whisper** or **ChatTTS** provide alternative backends but are not required for basic operation.

### Why does macOS require different dependencies than Linux?

Darwin (macOS) installations include **MLX**-specific packages (`mlx`, `mlx-audio`, `mlx-lm`) optimized for Apple Silicon GPUs, along with audio libraries like **Miniaudio** and phonemization tools. Linux installations instead receive **Nano-Parakeet** and **Faster-Qwen3-TTS** with GGML acceleration support.

### Do I need ONNX Runtime installed?

**ONNX Runtime** is only required when using the **Paraformer** extra on Python versions below 3.11, constrained to `<1.24`. If you do not install the `paraformer` extra or use Python 3.11+, this dependency is not required.