How the Speech-to-Speech Pipeline Handles Device Allocation Across VAD, STT, LLM, and TTS: CUDA, MPS, and CPU Support
The speech-to-speech pipeline uses a global device override system that propagates a single --device argument to all component-specific handlers (VAD, STT, LLM, TTS) through dataclass arguments, with special macOS optimizations for MPS and automatic CPU fallback.
The Hugging Face speech-to-speech repository provides a modular pipeline for real-time voice conversion. Understanding how the pipeline manages device allocation across VAD/STT/LLM/TTS components is essential for optimizing performance on CUDA GPUs, Apple Silicon (MPS), or CPU-only systems.
Global Device Override Strategy
The central orchestration logic resides in src/speech_to_speech/s2s_pipeline.py. When you specify --device <dev> via command line or set module_kwargs.device programmatically, the overwrite_device_argument() function (lines 263-277) copies that value to every component-specific device field.
This function automatically populates llm_device, tts_device, stt_device, and other handler-specific device arguments with your chosen backend. By default, individual handlers assume "cuda", but the global override ensures consistency across the entire pipeline.
macOS MPS Optimization and Platform Validation
Automatic MPS Configuration
When running on Apple Silicon, setting module_kwargs.local_mac_optimal_settings to True triggers optimal_mac_settings() (lines 31-45). This function forces every component to use the Metal Performance Shaders (MPS) backend by setting device = "mps" and applies macOS-compatible defaults such as tts = "qwen3".
CUDA Validation on macOS
The pipeline explicitly validates device compatibility for macOS. If you attempt to specify cuda on macOS, the code raises an error (lines 50-53) because CUDA is unavailable on that platform.
Component-Level Device Configuration
Each handler receives device configuration through dedicated dataclasses in src/speech_to_speech/arguments_classes/:
- STT Handler: WhisperSTTHandlerArguments exposes
stt_device(default"cuda") inwhisper_stt_arguments.py - LLM Handler: LanguageModelHandlerArguments exposes
llm_device(default"cuda") inlanguage_model_arguments.py - TTS Handler: Qwen3TTSHandlerArguments exposes
qwen3_tts_device(default"cuda") inqwen3_tts_arguments.py
These dataclasses populate the setup_kwargs dictionaries that handlers receive during instantiation.
Handler Instantiation and Model Placement
Factory Function Dispatch
The pipeline uses factory functions to instantiate handlers with the correct device settings:
- get_stt_handler() passes
vars(whisper_stt_handler_kwargs)toWhisperSTTHandler(line ~669) insrc/speech_to_speech/STT/whisper_stt_handler.py - get_llm_handler() passes
lm_kwargstoLanguageModelHandler(line ~858) insrc/speech_to_speech/LLM/language_model.py - get_tts_handler() passes
vars(qwen3_tts_handler_kwargs)toQwen3TTSHandler(line ~971) insrc/speech_to_speech/TTS/qwen3_tts_handler.py
Each handler internally moves its model to the specified device using self.model.to(self.device).
MPS Memory Synchronization
When running on MPS, handlers requiring explicit synchronization call torch.mps.synchronize() and torch.mps.empty_cache() after generation. For example, ChatTTSHandler in src/speech_to_speech/TTS/chatTTS_handler.py (lines 81-86) implements this pattern to prevent memory accumulation on Apple Silicon.
Deployment Scenarios and Command Examples
Standard Linux with CUDA
By default, all components use "cuda" as specified in their respective argument dataclasses. The models run on NVIDIA GPUs via standard CUDA streams without additional configuration.
python -m speech_to_speech \
--device cuda \
--stt whisper \
--llm_backend transformers \
--tts qwen3
Apple Silicon with MPS
Specify --local_mac_optimal_settings or --device mps to enable Metal Performance Shaders. The optimal_mac_settings() function ensures all handlers receive "mps" as their device.
python -m speech_to_speech \
--local_mac_optimal_settings \
--stt whisper \
--llm_backend mlx-lm \
--tts qwen3
CPU-Only Execution
When "cpu" is specified (either explicitly or as a fallback), handlers like PocketTTSHandler load models onto the CPU and skip GPU-specific optimizations. No extra synchronization is required for CPU execution.
Summary
- The global device override in
s2s_pipeline.pyensures consistency across all pipeline components throughoverwrite_device_argument() - macOS optimization automatically configures MPS via
optimal_mac_settings()and rejects invalid CUDA requests on Apple Silicon - Per-component dataclasses (
WhisperSTTHandlerArguments,LanguageModelHandlerArguments, etc.) receive device values through their*_devicefields - Factory functions instantiate handlers with pre-populated device arguments, and handlers move models via
self.model.to(self.device) - MPS-specific synchronization occurs in handlers like
ChatTTSHandlerto manage memory on Apple Silicon
Frequently Asked Questions
How do I force all pipeline components to use CPU instead of GPU?
Pass --device cpu when launching the pipeline. The overwrite_device_argument() function in s2s_pipeline.py will propagate this value to stt_device, llm_device, and tts_device, causing each handler to load its model onto the CPU.
Can I run the STT on CUDA while keeping the LLM on CPU?
The current implementation uses a global device override that applies the same device to all components. While the underlying dataclasses support individual device arguments, the overwrite_device_argument() function overrides all component-specific settings with the global value. To achieve mixed devices, you would need to modify the pipeline initialization to bypass the global override.
What happens if I specify --device cuda on a Mac?
The pipeline raises an explicit error during initialization. In s2s_pipeline.py (lines 50-53), the code checks for macOS and rejects "cuda" because CUDA is not available on Apple Silicon platforms. Use --device mps or --local_mac_optimal_settings instead.
Does the VAD component respect the same device allocation rules?
Yes, the Voice Activity Detection (VAD) handler follows the same pattern as STT, LLM, and TTS components. It receives its device configuration through the global override mechanism and loads its model accordingly, though the implementation focuses primarily on the three main processing stages.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →