How to Run Speech-to-Speech on Apple Silicon (MPS/MLX), CUDA, or CPU: A Complete Hardware Guide
You can run the huggingface/speech-to-speech pipeline on Apple Silicon (MPS/MLX), CUDA, or CPU by setting the --device flag and selecting hardware-specific backends for each modular component (STT, LLM, TTS).
The huggingface/speech-to-speech repository provides a modular voice-to-voice pipeline that lets you mix and match hardware accelerators independently across its three core stages. Whether you are deploying on a MacBook with Apple Silicon, an NVIDIA GPU workstation, or a CPU-only server, the CLI arguments defined in src/speech_to_speech/arguments_classes/module_arguments.py give you granular control over device placement.
Understanding Device Selection in the Pipeline
The pipeline consists of three independently configurable components: STT (speech-to-text), LLM (language model), and TTS (text-to-speech). Each component exposes its own device flag, while the global --device parameter acts as a fallback when specific flags are omitted.
According to the source code in src/speech_to_speech/arguments_classes/module_arguments.py, the device selection logic works as follows:
- STT: Controlled via
--stt(backend selection) and--stt_device(hardware placement) - LLM: Controlled via
--llm_backendand--model_device(for local models) or the generic--device - TTS: Controlled via
--ttsand--tts_device(e.g.,--qwen3_tts_device)
Running on Apple Silicon (MPS/MLX)
For macOS users with Apple Silicon, the fastest path is the --local_mac_optimal_settings flag. As implemented in src/speech_to_speech/arguments_classes/module_arguments.py (lines 17–22), this single flag automatically configures the entire pipeline for MLX acceleration:
speech-to-speech --local_mac_optimal_settings
This command performs three optimizations:
- Sets
--device mpsglobally - Selects Parakeet TDT for STT (which automatically uses the
mlx-audioimplementation) - Chooses MLX-LM for the LLM and Qwen3-TTS (MLX-audio backend) for speech synthesis
Manual Configuration for Custom Setups
If you need specific backends while retaining MPS acceleration, pass the device flags explicitly:
speech-to-speech \
--device mps \
--stt whisper-mlx \
--tts kokoro \
--llm_backend mlx-lm \
--model_name mlx-community/Qwen3-4B-Instruct-2507-bf16
Note: The whisper-mlx backend requires the MLX-audio stack, which you install via the [whisper-mlx] extra target. The handler logic in src/speech_to_speech/STT/parakeet_tdt_handler.py and src/speech_to_speech/TTS/qwen3_tts_handler.py automatically routes to MLX-audio when ---device mps is detected on macOS.
Running on CUDA GPUs
When an NVIDIA GPU is available, the pipeline defaults to CUDA unless overridden. Set --device cuda explicitly and choose CUDA-compatible backends:
speech-to-speech \
--device cuda \
--stt parakeet-tdt \
--tts qwen3 \
--llm_backend responses-api \
--model_name gpt-4o-mini \
--responses_api_api_key $OPENAI_API_KEY
For the Qwen3-TTS component, you can select the Torch CUDA backend instead of the default GGML by adding:
--qwen3_tts_backend torch
The Torch backend requires specific CUDA wheels detailed in the README (lines 100–115). The device selection logic in src/speech_to_speech/TTS/kokoro_handler.py and src/speech_to_speech/TTS/qwen3_tts_handler.py handles the CUDA runtime initialization based on these flags.
Running on CPU-Only Systems
For CPU-only deployment, specify cpu as the device and select compatible backends:
speech-to-speech \
--device cpu \
--stt whisper \
--tts qwen3 \
--llm_backend transformers \
--model_name google/gemma-2b-it
When using the Qwen3-TTS GGML backend on Linux, you may need to install the CPU-specific wheel if the default GGML build does not load. The handler implementation in src/speech_to_speech/TTS/qwen3_tts_handler.py falls back to CPU inference when CUDA is unavailable and the GGML backend is selected.
Summary
- Apple Silicon: Use
--local_mac_optimal_settingsfor automatic MLX configuration, or manually set--device mpswithwhisper-mlx,mlx-lm, andkokorobackends. - CUDA: Set
--device cudaand useparakeet-tdtorqwen3with the Torch backend for maximum GPU utilization. - CPU: Specify
--device cpuwithwhisperSTT andtransformersLLM backend for broad compatibility. - Source files: Device routing logic lives in
src/speech_to_speech/arguments_classes/module_arguments.py, while handler-specific implementations are insrc/speech_to_speech/STT/parakeet_tdt_handler.py,src/speech_to_speech/TTS/kokoro_handler.py, andsrc/speech_to_speech/TTS/qwen3_tts_handler.py.
Frequently Asked Questions
Can I mix different hardware devices for different pipeline components?
Yes. The modular architecture allows you to run STT on CPU, LLM on CUDA, and TTS on MPS by specifying the device flag for each component. For example, use --stt_device cpu alongside --device cuda to offload the transcription to CPU while keeping the language model on GPU.
What is the difference between MPS and MLX backends on Apple Silicon?
MPS (Metal Performance Shaders) is Apple's PyTorch backend for GPU acceleration, while MLX is Apple's native machine learning framework. The whisper-mlx and mlx-lm backends use the MLX framework, which often provides better performance on Apple Silicon than MPS-based PyTorch, particularly for the STT and LLM components.
How do I install the correct dependencies for my hardware?
Install extras correspond to specific backends. For Apple Silicon with Whisper, use pip install speech-to-speech[whisper-mlx]. For CUDA-specific Qwen3-TTS wheels, install the appropriate CUDA version listed in the README (lines 113–115). The base pip install speech-to-speech provides CPU-compatible defaults.
Why does Qwen3-TTS use different backends on macOS versus Linux?
Platform-specific optimizations. According to src/speech_to_speech/TTS/qwen3_tts_handler.py, the handler selects mlx-audio on macOS when --device mps is set, while defaulting to the GGML backend on Linux for CPU inference or Torch for CUDA. This ensures the most efficient inference library is used for each operating system.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →