# System Requirements for Running Speech-to-Speech Models: Hardware, Software, and Setup Guide

> Discover system requirements for Hugging Face speech-to-speech models. Learn about Python 3.10+, GPU/MLX needs, and OS compatibility for NLP projects. Get started now.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: getting-started
- Published: 2026-07-07

---

**The Hugging Face speech-to-speech pipeline requires Python 3.10 or newer, runs on macOS and Linux/Windows, and needs either CUDA 12 for GPU-accelerated TTS on Linux or Apple Silicon with MLX for optimized macOS inference.**

The huggingface/speech-to-speech repository delivers a low-latency, modular voice agent that chains Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) components. Before installing, verify that your system meets the hardware and software dependencies declared in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) and documented in the README.

## Python Version and Core Dependencies

The package requires **Python 3.10, 3.11, or 3.12**, as specified by the `requires-python` field in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) (line 10). Installation on older Python versions will fail.

Mandatory runtime dependencies declared in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) (lines 27-49) include:
- `torch` and `torchaudio` for neural network operations
- `transformers` for LLM and model management
- `fastapi`, `uvicorn`, and `websockets` for the OpenAI Realtime-compatible API server
- `sounddevice` and `soundfile` for cross-platform audio I/O
- `nano-parakeet` for the default STT backend
- `faster-qwen3-tts[ggml]` for the default TTS backend
- `lingua-language-detector` for automatic language detection

## Operating System Support

The pipeline supports **macOS (Darwin)**, **Linux**, and **Windows**. Platform-specific wheels are selected automatically via environment markers in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) (lines 31-46).

- **macOS**: Uses the MLX stack (`mlx`, `mlx-audio`, `mlx-lm`, `mlx-metal`) and `miniaudio` for optimized Apple Silicon inference, as defined in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) (lines 52-57).
- **Linux/Windows**: Uses standard PyTorch CUDA wheels and `sounddevice` for audio capture.

## Hardware Requirements by Component

Each pipeline component in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) has distinct hardware needs:

**Voice Activity Detection (VAD)**
- Runs on CPU only using Silero VAD. No GPU required.

**Speech-to-Text (STT)**
- **Parakeet-TDT** (default): Benefits from CUDA GPU on Linux; uses Apple Silicon via MLX on macOS.
- **Faster Whisper**: Requires GPU for real-time performance.
- **Paraformer**: GPU recommended for low latency.

**Large Language Model (LLM)**
- Most compute-intensive component. Requires **CUDA GPU** for large transformer models on Linux, or **Apple Silicon + MLX** on macOS.
- CPU fallback available for smaller models only.

**Text-to-Speech (TTS)**
- **Qwen3-TTS GGML**: Requires **CUDA 12 runtime** on Linux (specifically CUDA 12.8 as per the wheel requirements).
- **MLX Audio**: Uses Apple Silicon Metal Performance Shaders on macOS.
- **Pocket TTS**: CPU-only fallback option.

## CUDA Requirements for Linux GPU Acceleration

For the default Qwen3-TTS GGML backend on Linux, you must have **CUDA 12** installed. The wheel targets CUDA 12.8 specifically. If your system runs a different CUDA version (e.g., 12.4), install the matching wheel from the Hugging Face wheelhouse as documented in the README:

```bash
pip install "qwentts-cpp-python==0.3.0+cu124" \
    -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124

```

Mismatched CUDA versions will cause runtime errors when initializing the TTS backend.

## macOS and Apple Silicon Optimization

On macOS, the pipeline automatically utilizes the **MLX stack** when you install the package. The [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) (lines 52-57) declares platform-specific dependencies including `mlx`, `mlx-audio`, and `mlx-lm`.

For optimal performance, use the dedicated flag:

```bash
pip install "speech-to-speech[whisper-mlx]"
speech-to-speech --local_mac_optimal_settings

```

This configures Metal GPU usage (`--device mps`) and selects MLX-based LLM and TTS backends automatically via the argument classes in `src/speech_to_speech/arguments_classes/`.

## Optional Backend Dependencies

The `optional-dependencies` section in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) (lines 63-97) defines extras for alternative backends:

- `speech-to-speech[kokoro]` – Kokoro-82M TTS (CUDA/CPU on non-macOS, built-in on macOS)
- `speech-to-speech[pocket]` – Pocket TTS CPU-only backend
- `speech-to-speech[chattts]` – ChatTTS backend
- `speech-to-speech[facebook-mms]` – MMS TTS backend
- `speech-to-speech[faster-whisper]` – Faster Whisper STT
- `speech-to-speech[whisper-mlx]` – Lightning Whisper MLX for macOS
- `speech-to-speech[paraformer]` – Paraformer STT backend

Install only the extras you need to minimize dependency overhead.

## Installation and Verification Examples

### Basic Installation (CPU/GPU Auto-detect)

```bash

# Requires Python 3.10+

pip install speech-to-speech

# Set API key for OpenAI LLM backend

export OPENAI_API_KEY=your-key-here

# Launch WebSocket server (ws://localhost:8765/v1/realtime)

speech-to-speech

```

### macOS Apple Silicon Setup

```bash

# Install with MLX Whisper support

pip install "speech-to-speech[whisper-mlx]"

# Run with Metal optimization

speech-to-speech --local_mac_optimal_settings

```

### Custom Linux GPU Configuration

```bash

# Install specific CUDA wheel if needed (example for CUDA 12.4)

pip install "qwentts-cpp-python==0.3.0+cu124" \
    -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124

# Run with specific backends

speech-to-speech \
    --stt faster-whisper \
    --tts qwen3 \
    --qwen3_tts_backend ggml \
    --model_name "gpt-4o-mini"

```

## Summary

- **Python**: Version 3.10, 3.11, or 3.12 is mandatory according to [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml).
- **OS**: macOS (Darwin), Linux, and Windows supported with platform-specific wheels.
- **GPU**: CUDA 12 required for Qwen3-TTS GGML on Linux; Apple Silicon recommended for macOS MLX backends.
- **Audio**: `sounddevice` and `soundfile` handle I/O across platforms.
- **Modularity**: Install optional extras (`[kokoro]`, `[faster-whisper]`, etc.) only for backends you intend to use.

## Frequently Asked Questions

### Do I need a GPU to run the speech-to-speech pipeline?

No, but it is strongly recommended for real-time performance. The VAD component runs on CPU, but STT, LLM, and TTS components benefit significantly from GPU acceleration. You can run smaller models on CPU using the Pocket TTS backend or smaller LLM variants, though latency will increase substantially.

### Can I run this on Windows?

Yes, the repository supports Windows, Linux, and macOS. Windows installations use the same `pip install speech-to-speech` command, but ensure you have the appropriate audio drivers for `sounddevice` to function correctly. Platform-specific logic in [`pyproject.toml`](https://github.com/huggingface/speech-to-speech/blob/main/pyproject.toml) handles dependency resolution automatically.

### How do I resolve CUDA version mismatches with Qwen3-TTS?

If you encounter errors loading the Qwen3-TTS GGML backend, verify your CUDA version with `nvcc --version`. The default wheel requires CUDA 12.8. For other versions (e.g., 12.4), install the matching wheel from the Hugging Face wheelhouse as shown in the installation examples above, or use the CPU fallback wheels documented in the README.

### What is the minimum RAM requirement?

While not explicitly stated in the configuration files, running the full pipeline with default models (Parakeet-TDT, GPT-4o-mini, Qwen3-TTS) requires approximately **8GB of system RAM** for CPU-only operation, and **4GB+ VRAM** for GPU-accelerated inference. Larger LLM models require proportionally more memory.