# Configuring MPS Device and MLX Backends with --local_mac_optimal_settings in Speech-to-Speech

> Optimize your Speech-to-Speech pipeline on Apple Silicon Macs. Discover how `--local_mac_optimal_settings` configures MPS device and MLX backends for faster STT, LLM, and TTS.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-09

---

**The `--local_mac_optimal_settings` flag automatically configures Apple Silicon Macs to use the MPS device and MLX-optimized backends for the STT, LLM, and TTS pipeline stages.**

The `speech-to-speech` repository from Hugging Face provides a low-latency, fully-modular voice-agent framework. When deploying on Apple Silicon, you can simplify hardware configuration using a single CLI argument that automatically selects the appropriate compute device and accelerated backends for local inference.

## What --local_mac_optimal_settings Configures

Enabling this flag triggers three automatic optimizations designed specifically for macOS running on Apple Silicon (M1/M2/M3/M4).

### Automatic MPS Device Assignment

The flag sets `--device mps` for every component in the four-stage pipeline. According to [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the `optimal_mac_settings` function (lines 231-242) ensures that **Voice Activity Detection (VAD)**, **Speech-to-Text (STT)**, **Large Language Model (LLM)**, and **Text-to-Speech (TTS)** handlers all target the **Metal Performance Shaders (MPS)** backend.

### MLX Backend Selection

Beyond device configuration, the flag automatically chooses backends optimized for Apple's MLX framework:

- **STT**: Parakeet TDT for local transcription
- **LLM**: mlx-lm for local language model inference
- **TTS**: Qwen3-TTS (MLX implementation) for speech synthesis

These selections are enforced by the `check_mac_settings` validation routine (lines 251-267) within [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).

## Source Code Implementation

The configuration logic resides in the pipeline orchestration file and integrates with the handler factory. In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the system implements two key functions:

1. **`optimal_mac_settings`** (lines 231-242): Applies default values for Apple Silicon, setting the device to MPS and selecting MLX-compatible handlers.

2. **`check_mac_settings`** (lines 251-267): Validates that the chosen STT, LLM, and TTS backends support MLX acceleration when the optimal settings flag is active.

The CLI flag itself is defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py). Handler instantiation is managed by [`src/speech_to_speech/handler_factory.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/handler_factory.py), which maps these configuration flags to concrete MLX-compatible implementations.

## Running the Pipeline on macOS

To start the speech-to-speech pipeline with automatic MPS and MLX configuration, execute:

```bash
speech-to-speech --local_mac_optimal_settings

```

This single flag is equivalent to manually specifying:

```bash
speech-to-speech \
    --device mps \
    --stt parakeet-tdt \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --mode local

```

The `--mode local` switch disables the WebSocket server, running the entire pipeline locally without requiring an external API key.

## Pipeline Architecture Context

The system consists of four interchangeable stages communicating through thread-safe queues defined in `initialize_queues_and_events`:

- **VAD**: Silero V5 for voice activity detection
- **STT**: Transcription via the selected handler (Parakeet TDT when using optimal settings)
- **LLM**: Response generation via mlx-lm on MPS
- **TTS**: Audio synthesis via Qwen3-TTS utilizing MLX optimizations

Each handler runs in separate threads, connecting through queues such as `stt_output_queue` and `lm_response_queue` to enable real-time streaming without blocking the main execution thread.

## Summary

- **`--local_mac_optimal_settings`** automatically configures MPS device usage across VAD, STT, LLM, and TTS components.
- The flag selects **Parakeet TDT**, **mlx-lm**, and **Qwen3-TTS** backends optimized for Apple Silicon.
- Implementation resides in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) within `optimal_mac_settings` (lines 231-242) and `check_mac_settings` (lines 251-267).
- Running with this flag equals manually setting `--device mps` with MLX-compatible handlers and `--mode local`.
- The architecture uses queue-based communication between pipeline stages to maintain low latency.

## Frequently Asked Questions

### What hardware supports --local_mac_optimal_settings?

This flag is designed exclusively for Apple Silicon Macs (M1, M2, M3, and M4 series). The implementation validates the macOS platform and ARM architecture before applying MPS and MLX configurations in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

### Can I override specific backends when using --local_mac_optimal_settings?

No, the validation routine `check_mac_settings` enforces specific backend choices to ensure MLX compatibility. To use custom backends, omit the optimal settings flag and manually configure `--device`, `--stt`, `--llm_backend`, and `--tts` arguments as defined in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py).

### Does this flag work with the Realtime WebSocket API?

No, enabling `--local_mac_optimal_settings` automatically switches `--mode` to `local`, disabling the WebSocket server endpoint. For Realtime API compatibility with MPS acceleration, manually specify `--device mps` while maintaining `--mode realtime`.

### Where is the device validation logic implemented?

The device and backend validation occurs in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) within the `check_mac_settings` function (lines 251-267). This routine verifies compatibility between the MPS device requirement and the selected MLX backends before pipeline initialization.