# How the Speech-to-Speech Pipeline Handles Device Allocation Across VAD, STT, LLM, and TTS: CUDA, MPS, and CPU Support

> Discover how the speech-to-speech pipeline manages device allocation for VAD, STT, LLM, and TTS on CUDA, MPS, and CPU. Learn about macOS optimizations and automatic CPU fallback.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-10

---

**The speech-to-speech pipeline uses a global device override system that propagates a single `--device` argument to all component-specific handlers (VAD, STT, LLM, TTS) through dataclass arguments, with special macOS optimizations for MPS and automatic CPU fallback.**

The Hugging Face `speech-to-speech` repository provides a modular pipeline for real-time voice conversion. Understanding how the pipeline manages **device allocation across VAD/STT/LLM/TTS** components is essential for optimizing performance on CUDA GPUs, Apple Silicon (MPS), or CPU-only systems.

## Global Device Override Strategy

The central orchestration logic resides in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). When you specify `--device <dev>` via command line or set `module_kwargs.device` programmatically, the **overwrite_device_argument()** function (lines 263-277) copies that value to every component-specific device field.

This function automatically populates `llm_device`, `tts_device`, `stt_device`, and other handler-specific device arguments with your chosen backend. By default, individual handlers assume `"cuda"`, but the global override ensures consistency across the entire pipeline.

## macOS MPS Optimization and Platform Validation

### Automatic MPS Configuration

When running on Apple Silicon, setting `module_kwargs.local_mac_optimal_settings` to `True` triggers **optimal_mac_settings()** (lines 31-45). This function forces every component to use the **Metal Performance Shaders (MPS)** backend by setting `device = "mps"` and applies macOS-compatible defaults such as `tts = "qwen3"`.

### CUDA Validation on macOS

The pipeline explicitly validates device compatibility for macOS. If you attempt to specify `cuda` on macOS, the code raises an error (lines 50-53) because CUDA is unavailable on that platform.

## Component-Level Device Configuration

Each handler receives device configuration through dedicated **dataclasses** in `src/speech_to_speech/arguments_classes/`:

- **STT Handler**: **WhisperSTTHandlerArguments** exposes `stt_device` (default `"cuda"`) in [`whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_arguments.py)
- **LLM Handler**: **LanguageModelHandlerArguments** exposes `llm_device` (default `"cuda"`) in [`language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model_arguments.py)  
- **TTS Handler**: **Qwen3TTSHandlerArguments** exposes `qwen3_tts_device` (default `"cuda"`) in [`qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_arguments.py)

These dataclasses populate the `setup_kwargs` dictionaries that handlers receive during instantiation.

## Handler Instantiation and Model Placement

### Factory Function Dispatch

The pipeline uses **factory functions** to instantiate handlers with the correct device settings:

1. **get_stt_handler()** passes `vars(whisper_stt_handler_kwargs)` to `WhisperSTTHandler` (line ~669) in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py)
2. **get_llm_handler()** passes `lm_kwargs` to `LanguageModelHandler` (line ~858) in [`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py)
3. **get_tts_handler()** passes `vars(qwen3_tts_handler_kwargs)` to `Qwen3TTSHandler` (line ~971) in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)

Each handler internally moves its model to the specified device using `self.model.to(self.device)`.

### MPS Memory Synchronization

When running on MPS, handlers requiring explicit synchronization call **torch.mps.synchronize()** and **torch.mps.empty_cache()** after generation. For example, `ChatTTSHandler` in [`src/speech_to_speech/TTS/chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/chatTTS_handler.py) (lines 81-86) implements this pattern to prevent memory accumulation on Apple Silicon.

## Deployment Scenarios and Command Examples

### Standard Linux with CUDA

By default, all components use `"cuda"` as specified in their respective argument dataclasses. The models run on NVIDIA GPUs via standard CUDA streams without additional configuration.

```bash
python -m speech_to_speech \
    --device cuda \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3

```

### Apple Silicon with MPS

Specify `--local_mac_optimal_settings` or `--device mps` to enable Metal Performance Shaders. The `optimal_mac_settings()` function ensures all handlers receive `"mps"` as their device.

```bash
python -m speech_to_speech \
    --local_mac_optimal_settings \
    --stt whisper \
    --llm_backend mlx-lm \
    --tts qwen3

```

### CPU-Only Execution

When `"cpu"` is specified (either explicitly or as a fallback), handlers like `PocketTTSHandler` load models onto the CPU and skip GPU-specific optimizations. No extra synchronization is required for CPU execution.

## Summary

- The **global device override** in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) ensures consistency across all pipeline components through `overwrite_device_argument()`
- **macOS optimization** automatically configures MPS via `optimal_mac_settings()` and rejects invalid CUDA requests on Apple Silicon
- **Per-component dataclasses** (`WhisperSTTHandlerArguments`, `LanguageModelHandlerArguments`, etc.) receive device values through their `*_device` fields
- **Factory functions** instantiate handlers with pre-populated device arguments, and handlers move models via `self.model.to(self.device)`
- **MPS-specific synchronization** occurs in handlers like `ChatTTSHandler` to manage memory on Apple Silicon

## Frequently Asked Questions

### How do I force all pipeline components to use CPU instead of GPU?

Pass `--device cpu` when launching the pipeline. The `overwrite_device_argument()` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) will propagate this value to `stt_device`, `llm_device`, and `tts_device`, causing each handler to load its model onto the CPU.

### Can I run the STT on CUDA while keeping the LLM on CPU?

The current implementation uses a global device override that applies the same device to all components. While the underlying dataclasses support individual device arguments, the `overwrite_device_argument()` function overrides all component-specific settings with the global value. To achieve mixed devices, you would need to modify the pipeline initialization to bypass the global override.

### What happens if I specify `--device cuda` on a Mac?

The pipeline raises an explicit error during initialization. In [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (lines 50-53), the code checks for macOS and rejects `"cuda"` because CUDA is not available on Apple Silicon platforms. Use `--device mps` or `--local_mac_optimal_settings` instead.

### Does the VAD component respect the same device allocation rules?

Yes, the Voice Activity Detection (VAD) handler follows the same pattern as STT, LLM, and TTS components. It receives its device configuration through the global override mechanism and loads its model accordingly, though the implementation focuses primarily on the three main processing stages.