Configuring MPS Device and MLX Backends with --local_mac_optimal_settings in Speech-to-Speech
The --local_mac_optimal_settings flag automatically configures Apple Silicon Macs to use the MPS device and MLX-optimized backends for the STT, LLM, and TTS pipeline stages.
The speech-to-speech repository from Hugging Face provides a low-latency, fully-modular voice-agent framework. When deploying on Apple Silicon, you can simplify hardware configuration using a single CLI argument that automatically selects the appropriate compute device and accelerated backends for local inference.
What --local_mac_optimal_settings Configures
Enabling this flag triggers three automatic optimizations designed specifically for macOS running on Apple Silicon (M1/M2/M3/M4).
Automatic MPS Device Assignment
The flag sets --device mps for every component in the four-stage pipeline. According to src/speech_to_speech/s2s_pipeline.py, the optimal_mac_settings function (lines 231-242) ensures that Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) handlers all target the Metal Performance Shaders (MPS) backend.
MLX Backend Selection
Beyond device configuration, the flag automatically chooses backends optimized for Apple's MLX framework:
- STT: Parakeet TDT for local transcription
- LLM: mlx-lm for local language model inference
- TTS: Qwen3-TTS (MLX implementation) for speech synthesis
These selections are enforced by the check_mac_settings validation routine (lines 251-267) within s2s_pipeline.py.
Source Code Implementation
The configuration logic resides in the pipeline orchestration file and integrates with the handler factory. In src/speech_to_speech/s2s_pipeline.py, the system implements two key functions:
-
optimal_mac_settings(lines 231-242): Applies default values for Apple Silicon, setting the device to MPS and selecting MLX-compatible handlers. -
check_mac_settings(lines 251-267): Validates that the chosen STT, LLM, and TTS backends support MLX acceleration when the optimal settings flag is active.
The CLI flag itself is defined in src/speech_to_speech/arguments_classes/module_arguments.py. Handler instantiation is managed by src/speech_to_speech/handler_factory.py, which maps these configuration flags to concrete MLX-compatible implementations.
Running the Pipeline on macOS
To start the speech-to-speech pipeline with automatic MPS and MLX configuration, execute:
speech-to-speech --local_mac_optimal_settings
This single flag is equivalent to manually specifying:
speech-to-speech \
--device mps \
--stt parakeet-tdt \
--llm_backend mlx-lm \
--tts qwen3 \
--mode local
The --mode local switch disables the WebSocket server, running the entire pipeline locally without requiring an external API key.
Pipeline Architecture Context
The system consists of four interchangeable stages communicating through thread-safe queues defined in initialize_queues_and_events:
- VAD: Silero V5 for voice activity detection
- STT: Transcription via the selected handler (Parakeet TDT when using optimal settings)
- LLM: Response generation via mlx-lm on MPS
- TTS: Audio synthesis via Qwen3-TTS utilizing MLX optimizations
Each handler runs in separate threads, connecting through queues such as stt_output_queue and lm_response_queue to enable real-time streaming without blocking the main execution thread.
Summary
--local_mac_optimal_settingsautomatically configures MPS device usage across VAD, STT, LLM, and TTS components.- The flag selects Parakeet TDT, mlx-lm, and Qwen3-TTS backends optimized for Apple Silicon.
- Implementation resides in
src/speech_to_speech/s2s_pipeline.pywithinoptimal_mac_settings(lines 231-242) andcheck_mac_settings(lines 251-267). - Running with this flag equals manually setting
--device mpswith MLX-compatible handlers and--mode local. - The architecture uses queue-based communication between pipeline stages to maintain low latency.
Frequently Asked Questions
What hardware supports --local_mac_optimal_settings?
This flag is designed exclusively for Apple Silicon Macs (M1, M2, M3, and M4 series). The implementation validates the macOS platform and ARM architecture before applying MPS and MLX configurations in src/speech_to_speech/s2s_pipeline.py.
Can I override specific backends when using --local_mac_optimal_settings?
No, the validation routine check_mac_settings enforces specific backend choices to ensure MLX compatibility. To use custom backends, omit the optimal settings flag and manually configure --device, --stt, --llm_backend, and --tts arguments as defined in module_arguments.py.
Does this flag work with the Realtime WebSocket API?
No, enabling --local_mac_optimal_settings automatically switches --mode to local, disabling the WebSocket server endpoint. For Realtime API compatibility with MPS acceleration, manually specify --device mps while maintaining --mode realtime.
Where is the device validation logic implemented?
The device and backend validation occurs in src/speech_to_speech/s2s_pipeline.py within the check_mac_settings function (lines 251-267). This routine verifies compatibility between the MPS device requirement and the selected MLX backends before pipeline initialization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →