How the Speech-to-Speech Pipeline Handles Device Allocation Across VAD, STT, LLM, and TTS: CUDA, MPS, and CPU Support

The speech-to-speech pipeline uses a global device override system that propagates a single --device argument to all component-specific handlers (VAD, STT, LLM, TTS) through dataclass arguments, with special macOS optimizations for MPS and automatic CPU fallback.

The Hugging Face speech-to-speech repository provides a modular pipeline for real-time voice conversion. Understanding how the pipeline manages device allocation across VAD/STT/LLM/TTS components is essential for optimizing performance on CUDA GPUs, Apple Silicon (MPS), or CPU-only systems.

Global Device Override Strategy

The central orchestration logic resides in src/speech_to_speech/s2s_pipeline.py. When you specify --device <dev> via command line or set module_kwargs.device programmatically, the overwrite_device_argument() function (lines 263-277) copies that value to every component-specific device field.

This function automatically populates llm_device, tts_device, stt_device, and other handler-specific device arguments with your chosen backend. By default, individual handlers assume "cuda", but the global override ensures consistency across the entire pipeline.

macOS MPS Optimization and Platform Validation

Automatic MPS Configuration

When running on Apple Silicon, setting module_kwargs.local_mac_optimal_settings to True triggers optimal_mac_settings() (lines 31-45). This function forces every component to use the Metal Performance Shaders (MPS) backend by setting device = "mps" and applies macOS-compatible defaults such as tts = "qwen3".

CUDA Validation on macOS

The pipeline explicitly validates device compatibility for macOS. If you attempt to specify cuda on macOS, the code raises an error (lines 50-53) because CUDA is unavailable on that platform.

Component-Level Device Configuration

Each handler receives device configuration through dedicated dataclasses in src/speech_to_speech/arguments_classes/:

These dataclasses populate the setup_kwargs dictionaries that handlers receive during instantiation.

Handler Instantiation and Model Placement

Factory Function Dispatch

The pipeline uses factory functions to instantiate handlers with the correct device settings:

  1. get_stt_handler() passes vars(whisper_stt_handler_kwargs) to WhisperSTTHandler (line ~669) in src/speech_to_speech/STT/whisper_stt_handler.py
  2. get_llm_handler() passes lm_kwargs to LanguageModelHandler (line ~858) in src/speech_to_speech/LLM/language_model.py
  3. get_tts_handler() passes vars(qwen3_tts_handler_kwargs) to Qwen3TTSHandler (line ~971) in src/speech_to_speech/TTS/qwen3_tts_handler.py

Each handler internally moves its model to the specified device using self.model.to(self.device).

MPS Memory Synchronization

When running on MPS, handlers requiring explicit synchronization call torch.mps.synchronize() and torch.mps.empty_cache() after generation. For example, ChatTTSHandler in src/speech_to_speech/TTS/chatTTS_handler.py (lines 81-86) implements this pattern to prevent memory accumulation on Apple Silicon.

Deployment Scenarios and Command Examples

Standard Linux with CUDA

By default, all components use "cuda" as specified in their respective argument dataclasses. The models run on NVIDIA GPUs via standard CUDA streams without additional configuration.

python -m speech_to_speech \
    --device cuda \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3

Apple Silicon with MPS

Specify --local_mac_optimal_settings or --device mps to enable Metal Performance Shaders. The optimal_mac_settings() function ensures all handlers receive "mps" as their device.

python -m speech_to_speech \
    --local_mac_optimal_settings \
    --stt whisper \
    --llm_backend mlx-lm \
    --tts qwen3

CPU-Only Execution

When "cpu" is specified (either explicitly or as a fallback), handlers like PocketTTSHandler load models onto the CPU and skip GPU-specific optimizations. No extra synchronization is required for CPU execution.

Summary

  • The global device override in s2s_pipeline.py ensures consistency across all pipeline components through overwrite_device_argument()
  • macOS optimization automatically configures MPS via optimal_mac_settings() and rejects invalid CUDA requests on Apple Silicon
  • Per-component dataclasses (WhisperSTTHandlerArguments, LanguageModelHandlerArguments, etc.) receive device values through their *_device fields
  • Factory functions instantiate handlers with pre-populated device arguments, and handlers move models via self.model.to(self.device)
  • MPS-specific synchronization occurs in handlers like ChatTTSHandler to manage memory on Apple Silicon

Frequently Asked Questions

How do I force all pipeline components to use CPU instead of GPU?

Pass --device cpu when launching the pipeline. The overwrite_device_argument() function in s2s_pipeline.py will propagate this value to stt_device, llm_device, and tts_device, causing each handler to load its model onto the CPU.

Can I run the STT on CUDA while keeping the LLM on CPU?

The current implementation uses a global device override that applies the same device to all components. While the underlying dataclasses support individual device arguments, the overwrite_device_argument() function overrides all component-specific settings with the global value. To achieve mixed devices, you would need to modify the pipeline initialization to bypass the global override.

What happens if I specify --device cuda on a Mac?

The pipeline raises an explicit error during initialization. In s2s_pipeline.py (lines 50-53), the code checks for macOS and rejects "cuda" because CUDA is not available on Apple Silicon platforms. Use --device mps or --local_mac_optimal_settings instead.

Does the VAD component respect the same device allocation rules?

Yes, the Voice Activity Detection (VAD) handler follows the same pattern as STT, LLM, and TTS components. It receives its device configuration through the global override mechanism and loads its model accordingly, though the implementation focuses primarily on the three main processing stages.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →