# Dependencies for Speech-to-Speech: Complete Guide to Installation Requirements

> Discover the essential dependencies for speech-to-speech installation. Learn about core runtime requirements and optional extras for specific model backends.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**The `speech-to-speech` package declares its dependencies in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml), organized into core runtime requirements (22 packages), optional extras for specific model backends (11 groups), and development tools.**

The Hugging Face `speech-to-speech` repository is a real-time voice-to-voice conversation system that requires specific Python dependencies for audio processing, model inference, and web server functionality. Understanding these dependencies is essential for successful installation and deployment, whether you are running the minimal pipeline or integrating specialized backends like ChatTTS or Faster-Whisper.

## Core Runtime Dependencies

The essential packages required to run the base pipeline are defined in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) starting at line 27. These include web framework components, audio libraries, and machine learning frameworks:

- **Web server and API**: `fastapi`, `uvicorn`, and `websockets` provide the asynchronous HTTP/WebSocket server infrastructure
- **Networking**: `httpx` for async HTTP client operations
- **Audio I/O**: `sounddevice` and `soundfile` handle microphone capture and audio playback
- **Scientific computing**: `numpy`, `scipy`, `torch`, and `torchaudio` enable tensor manipulation and neural network inference
- **ML/NLP**: `transformers` for model loading, `nltk==3.10.0` for text tokenization, and `openai==2.28.0` for API compatibility
- **Data validation**: `pydantic>=2.0` for configuration management
- **Utilities**: `rich` for formatted console output and `pillow` for image processing

```bash

# Install only core dependencies

pip install speech-to-speech

```

## Optional Extras for Specific Backends

The repository supports modular installation via optional extras defined at line 64 in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml). These allow you to install only the dependencies needed for your specific speech-to-text (STT) or text-to-speech (TTS) backend:

**`chattts`** – `ChatTTS>=0.1.1` for the ChatTTS neural TTS model

**`facebook-mms`** – `transformers>=4.57.0` for Facebook's Massively Multilingual Speech models

**`faster-whisper`** – `faster-whisper>=1.0.3` for optimized Whisper inference

**`kokoro`** – `kokoro>=0.9.2` (skipped on macOS) for the Kokoro TTS engine

**`language-detection`** – `lingua-language-detector>=2.0.2` for automatic language identification

**`mlx-lm`** – `mlx-lm`, `mlx-vlm`, and `mlx-metal` for Apple Silicon optimized inference (macOS only)

**`paraformer`** – `funasr`, `modelscope`, and `onnxruntime` (Python < 3.11 only) for Alibaba's Paraformer models

**`pocket`** – `pocket-tts>=0.1.0` for Pocket TTS integration

**`webrtc`** – `aiortc>=1.9.0` for WebRTC peer-to-peer communication support

**`websocket`** – `websockets>=12.0` (redundant with core but pinned version)

**`whisper-mlx`** – `lightning-whisper-mlx>=0.0.10` for optimized Whisper on Apple Silicon (macOS only)

```bash

# Example: Install with ChatTTS and Faster-Whisper support

pip install "speech-to-speech[chattts,faster-whisper]"

```

## Development Dependencies

For contributors and those running the test suite, the `dev` dependency group at line 111 includes linting and testing tools:

- **Linting**: `ruff` for code formatting and import sorting
- **Type checking**: `mypy` for static type analysis
- **Testing**: `pytest` and `pytest-asyncio` for unit and integration tests
- **Network testing**: `websockets` and `aiortc` for testing real-time communication components

```bash

# Install with development dependencies

pip install "speech-to-speech[dev]"

```

## Platform-Specific Requirements

Several dependencies have platform-specific constraints, particularly for macOS (Darwin). According to the source code in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml), certain packages require specific binary builds on macOS for compatibility:

- **`numpy`**: Pinned to specific builds on macOS for binary compatibility
- **`torch`**: Uses macOS-specific wheels to ensure Metal Performance Shaders (MPS) support
- **`kokoro`**: Excluded from macOS installations due to compatibility issues
- **MLX packages**: The `mlx-lm`, `mlx-vlm`, `mlx-metal`, and `whisper-mlx` extras are exclusively for macOS with Apple Silicon

Linux and Windows users should install the standard PyTorch wheels with CUDA support where applicable, while macOS users should leverage the MLX extras for optimal performance on M1/M2/M3 chips.

## How Dependencies Map to Architecture

The dependency structure directly supports the pipeline architecture implemented in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) and related modules:

**Web Interface Layer** – `fastapi` and `uvicorn` instantiate the HTTP server that hosts the speech-to-speech endpoint, while `websockets` and optional `aiortc` enable real-time bidirectional streaming.

**Audio Processing Layer** – `sounddevice` handles microphone input streams, `soundfile` manages audio file I/O, and `torch`/`torchaudio` provide the tensor operations necessary for audio preprocessing and neural vocoding.

**Model Inference Layer** – `transformers` loads the underlying speech recognition and language models, while specific extras like `faster-whisper` or `ChatTTS` provide optimized implementations of particular model classes found in `src/speech_to_speech/STT/` and `src/speech_to_speech/TTS/`.

**Orchestration Layer** – `pydantic` validates the argument classes defined in `src/speech_to_speech/arguments_classes/`, ensuring type safety when configuring handlers like `WhisperSTTHandlerArguments` or `ChatTTSArguments`.

## Installation Examples

Install the core package for basic functionality:

```python

# Basic installation (includes torch, transformers, fastapi, etc.)

pip install "speech-to-speech==0.2.11"

```

Install with specific model support:

```python

# For ChatTTS TTS backend

pip install "speech-to-speech[chattts]==0.2.11"

# For Paraformer STT (requires Python < 3.11)

pip install "speech-to-speech[paraformer]==0.2.11"

# For macOS with MLX acceleration

pip install "speech-to-speech[mlx-lm,whisper-mlx]==0.2.11"

```

Launch the pipeline after installation:

```python
python -m speech_to_speech.s2s_pipeline

```

## Summary

- **Core dependencies** (22 packages) are defined in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) lines 27-62 and include `fastapi`, `torch`, `transformers`, and audio processing libraries.
- **Optional extras** (11 groups) at lines 64-99 allow modular installation of specific backends like ChatTTS, Faster-Whisper, and MLX-optimized models without bloating the base install.
- **Development tools** at lines 111-118 include `ruff`, `mypy`, and `pytest` for code quality and testing.
- **Platform constraints** apply particularly to macOS users, who should use MLX-specific extras for optimal performance while avoiding Linux-only packages like `kokoro`.
- The dependency structure supports the modular architecture in `src/speech_to_speech/`, separating concerns between web serving, audio I/O, and model inference.

## Frequently Asked Questions

### What are the minimum dependencies required to run speech-to-speech?

The minimum installation requires the 22 core runtime packages defined in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) including `fastapi`, `uvicorn`, `torch`, `transformers`, `sounddevice`, and `websockets`. These provide sufficient functionality to run the base pipeline with default STT and TTS handlers, though specific model backends may require additional optional extras.

### How do I install support for ChatTTS or Faster-Whisper specifically?

Use the optional extras syntax: `pip install "speech-to-speech[chattts]"` for ChatTTS support or `pip install "speech-to-speech[faster-whisper]"` for optimized Whisper inference. You can combine multiple extras using commas, such as `"speech-to-speech[chattts,faster-whisper]"`, to install dependencies for multiple backends simultaneously.

### Are there any macOS-specific dependency requirements?

Yes, macOS users with Apple Silicon should install the `mlx-lm` and `whisper-mlx` extras for optimal performance, while avoiding the `kokoro` extra which is skipped on Darwin platforms. The [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) specifies platform-specific markers that automatically handle version pinning for `numpy` and `torch` on macOS to ensure Metal Performance Shaders compatibility.

### What Python version is required for Paraformer support?

The Paraformer STT backend requires Python versions less than 3.11 due to dependencies on `funasr`, `modelscope`, and `onnxruntime` that are not compatible with Python 3.11+. This constraint is explicitly defined in the `paraformer` optional extra group in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml), whereas the core package supports standard Python versions.