# How to Run Speech-to-Speech on Apple Silicon (MPS/MLX), CUDA, or CPU: A Complete Hardware Guide

> Discover how to run speech-to-speech on Apple Silicon MPS MLX CUDA or CPU by setting the device flag. Optimize your hardware for STT LLM and TTS components. Get the complete guide.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-08

---

**You can run the huggingface/speech-to-speech pipeline on Apple Silicon (MPS/MLX), CUDA, or CPU by setting the `--device` flag and selecting hardware-specific backends for each modular component (STT, LLM, TTS).**

The huggingface/speech-to-speech repository provides a modular voice-to-voice pipeline that lets you mix and match hardware accelerators independently across its three core stages. Whether you are deploying on a MacBook with Apple Silicon, an NVIDIA GPU workstation, or a CPU-only server, the CLI arguments defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) give you granular control over device placement.

## Understanding Device Selection in the Pipeline

The pipeline consists of three independently configurable components: **STT** (speech-to-text), **LLM** (language model), and **TTS** (text-to-speech). Each component exposes its own device flag, while the global `--device` parameter acts as a fallback when specific flags are omitted.

According to the source code in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py), the device selection logic works as follows:

- **STT**: Controlled via `--stt` (backend selection) and `--stt_device` (hardware placement)
- **LLM**: Controlled via `--llm_backend` and `--model_device` (for local models) or the generic `--device`
- **TTS**: Controlled via `--tts` and `--tts_device` (e.g., `--qwen3_tts_device`)

## Running on Apple Silicon (MPS/MLX)

For macOS users with Apple Silicon, the fastest path is the `--local_mac_optimal_settings` flag. As implemented in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) (lines 17–22), this single flag automatically configures the entire pipeline for MLX acceleration:

```bash
speech-to-speech --local_mac_optimal_settings

```

This command performs three optimizations:
1. Sets `--device mps` globally
2. Selects **Parakeet TDT** for STT (which automatically uses the `mlx-audio` implementation)
3. Chooses **MLX-LM** for the LLM and **Qwen3-TTS** (MLX-audio backend) for speech synthesis

### Manual Configuration for Custom Setups

If you need specific backends while retaining MPS acceleration, pass the device flags explicitly:

```bash
speech-to-speech \
    --device mps \
    --stt whisper-mlx \
    --tts kokoro \
    --llm_backend mlx-lm \
    --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16

```

**Note**: The `whisper-mlx` backend requires the MLX-audio stack, which you install via the `[whisper-mlx]` extra target. The handler logic in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) and [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) automatically routes to MLX-audio when `---device mps` is detected on macOS.

## Running on CUDA GPUs

When an NVIDIA GPU is available, the pipeline defaults to CUDA unless overridden. Set `--device cuda` explicitly and choose CUDA-compatible backends:

```bash
speech-to-speech \
    --device cuda \
    --stt parakeet-tdt \
    --tts qwen3 \
    --llm_backend responses-api \
    --model_name gpt-4o-mini \
    --responses_api_api_key $OPENAI_API_KEY

```

For the **Qwen3-TTS** component, you can select the Torch CUDA backend instead of the default GGML by adding:

```bash
    --qwen3_tts_backend torch

```

The Torch backend requires specific CUDA wheels detailed in the README (lines 100–115). The device selection logic in [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py) and [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) handles the CUDA runtime initialization based on these flags.

## Running on CPU-Only Systems

For CPU-only deployment, specify `cpu` as the device and select compatible backends:

```bash
speech-to-speech \
    --device cpu \
    --stt whisper \
    --tts qwen3 \
    --llm_backend transformers \
    --model_name google/gemma-2b-it

```

When using the **Qwen3-TTS** GGML backend on Linux, you may need to install the CPU-specific wheel if the default GGML build does not load. The handler implementation in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) falls back to CPU inference when CUDA is unavailable and the GGML backend is selected.

## Summary

- **Apple Silicon**: Use `--local_mac_optimal_settings` for automatic MLX configuration, or manually set `--device mps` with `whisper-mlx`, `mlx-lm`, and `kokoro` backends.
- **CUDA**: Set `--device cuda` and use `parakeet-tdt` or `qwen3` with the Torch backend for maximum GPU utilization.
- **CPU**: Specify `--device cpu` with `whisper` STT and `transformers` LLM backend for broad compatibility.
- **Source files**: Device routing logic lives in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py), while handler-specific implementations are in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py), [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py), and [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py).

## Frequently Asked Questions

### Can I mix different hardware devices for different pipeline components?

**Yes.** The modular architecture allows you to run STT on CPU, LLM on CUDA, and TTS on MPS by specifying the device flag for each component. For example, use `--stt_device cpu` alongside `--device cuda` to offload the transcription to CPU while keeping the language model on GPU.

### What is the difference between MPS and MLX backends on Apple Silicon?

**MPS** (Metal Performance Shaders) is Apple's PyTorch backend for GPU acceleration, while **MLX** is Apple's native machine learning framework. The `whisper-mlx` and `mlx-lm` backends use the MLX framework, which often provides better performance on Apple Silicon than MPS-based PyTorch, particularly for the STT and LLM components.

### How do I install the correct dependencies for my hardware?

**Install extras correspond to specific backends.** For Apple Silicon with Whisper, use `pip install speech-to-speech[whisper-mlx]`. For CUDA-specific Qwen3-TTS wheels, install the appropriate CUDA version listed in the README (lines 113–115). The base `pip install speech-to-speech` provides CPU-compatible defaults.

### Why does Qwen3-TTS use different backends on macOS versus Linux?

**Platform-specific optimizations.** According to [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), the handler selects `mlx-audio` on macOS when `--device mps` is set, while defaulting to the **GGML** backend on Linux for CPU inference or Torch for CUDA. This ensures the most efficient inference library is used for each operating system.