# Performance Implications of Running STT, LLM, and TTS on Same vs Separate Devices

> Compare performance of running STT LLM TTS on same vs separate devices. Discover how device allocation impacts GPU memory, PCIe transfers, and overall latency for your speech applications.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-07-09

---

**Running STT, LLM, and TTS on separate devices eliminates GPU memory contention and MLX lock serialization, but introduces PCIe transfer overhead that may dominate latency for short utterances.**

The `huggingface/speech-to-speech` repository implements a fully modular speech-to-speech pipeline where Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) stages run as independent handlers communicating via thread-safe queues. Understanding the performance implications of running STT LLM TTS on same vs separate devices requires analyzing resource contention, memory pressure, and data transfer overhead in the pipeline architecture defined in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

## Pipeline Architecture and Device Allocation

The central orchestration in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) instantiates each pipeline stage as a handler running in its own thread, with the `overwrite_device_argument` function (lines 63-72) managing device assignment. By default, all handlers inherit the global `--device` argument, but the system supports per-handler device overrides through arguments defined in [`arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/arguments_classes/module_arguments.py), [`arguments_classes/whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/arguments_classes/whisper_stt_arguments.py), [`arguments_classes/language_model_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/arguments_classes/language_model_arguments.py), and [`arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/arguments_classes/qwen3_tts_arguments.py).

## Same-Device Execution: Low Latency, High Contention

### Resource Contention and the MLX Lock

When running STT, LLM, and TTS on the same device, all handlers share a single GPU or Apple Silicon chip. On macOS, the global MLX lock implemented in [`utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/utils/mlx_lock.py) serializes inference operations, causing a flood of warnings and significant slowdowns for progressive STT→LLM→TTS flows. This serialization forces the pipeline to process stages sequentially rather than concurrently, eliminating the benefits of the multi-threaded architecture.

### Memory Pressure Constraints

A single GPU must simultaneously hold weights for Whisper (STT), the LLM (e.g., Qwen-3-4B consuming >10GB), and the TTS model. This concentration often exceeds VRAM capacity, triggering out-of-memory crashes or forcing the use of quantized models that sacrifice quality. The [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) guard blocks `--num_pipelines > 1` unless `mode==realtime` (lines 10,846-850) to prevent overload on shared devices.

### Zero-Copy Data Transfer

Same-device execution benefits from in-memory queue communication through `Queue[AudioInItem]` instances, eliminating inter-device copy latency. Audio tensors flow directly between handlers without PCIe or NVLink transfer costs, minimizing per-utterance latency for short interactions.

## Separate-Device Execution: Scalability with Overhead

### Eliminating Resource Contention

Assigning distinct devices to each handler—such as STT on CPU, LLM on GPU 0, and TTS on GPU 1—removes the global lock contention present in same-device configurations. On non-macOS platforms, this allows true concurrent execution where the LLM can process speculative turns via [`pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline/speculative_turns.py) while STT handles new audio input.

### Memory Distribution Benefits

Distributing models across devices spreads memory usage, enabling simultaneous deployment of larger unquantized models. Each device only stores its assigned model weights, preventing the out-of-memory conditions common in single-GPU setups and allowing high-throughput deployments with multiple parallel pipelines.

### Inter-Device Transfer Costs

Moving audio features or token tensors between devices incurs PCIe or NVLink copy costs. For short utterances, this transfer latency can dominate total response time, potentially negating the benefits of parallel processing. The overhead scales with tensor size and bus bandwidth, making this architecture optimal for longer-form content where compute time exceeds transfer time.

## Platform-Specific Considerations

On macOS, the code forces `--llm_backend mlx-lm` and applies `check_mac_settings` and `optimal_mac_settings` that automatically disable live transcription when using multiple pipelines to avoid MLX lock contention (lines 48-61). Linux and Windows platforms do not suffer from this global lock limitation, making multi-device setups more effective for exploiting multiple GPUs.

## Configuration Examples

### Single-GPU Deployment

Run all stages on one CUDA device using the default inheritance mechanism:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --device cuda \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3

```

The `overwrite_device_argument` function (lines 63-72) ensures all handlers inherit the `cuda` device assignment.

### Multi-Device Distribution

Assign specific devices to each handler to eliminate contention:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --device cpu \
    --stt_device cpu \
    --llm_device cuda \
    --tts_device cuda:1 \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3

```

Per-handler arguments (`--stt_device`, `--llm_device`, `--tts_device`) override the global device setting via the merging logic in `overwrite_device_argument`.

### macOS Multi-Pipeline Limitations

Disable live transcription on Apple Silicon to avoid lock contention:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --num_pipelines 2 \
    --local_mac_optimal_settings true

```

This configuration detects macOS and disables features that trigger the global MLX lock serialization.

## Summary

- **Same-device execution** minimizes data transfer latency through in-memory queues but suffers from GPU memory pressure and, on Apple Silicon, MLX lock serialization that forces sequential processing.
- **Separate-device execution** eliminates resource contention and distributes memory load, enabling larger models and higher throughput, but introduces PCIe/NVLink transfer overhead that impacts short-utterance latency.
- **Platform-specific locks** on macOS make multi-device setups more beneficial than on Linux/Windows, where the global MLX lock does not constrain same-device multi-pipeline deployments.
- **Configuration flexibility** via [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) allows granular device assignment through per-handler arguments, supporting both simple single-GPU deployments and complex multi-GPU orchestrations.

## Frequently Asked Questions

### What causes the MLX lock warnings when running all pipeline stages on a Mac?

The global MLX lock in [`utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/utils/mlx_lock.py) serializes inference operations on Apple Silicon, forcing STT, LLM, and TTS handlers to execute sequentially rather than concurrently. This lock generates warnings when multiple handlers attempt simultaneous access to the MLX device, causing performance degradation in progressive speech-to-speech flows.

### How does splitting STT, LLM, and TTS across devices affect memory requirements?

Distributing models across separate devices reduces per-device memory requirements significantly. Each device only needs to load its assigned model weights—for example, the LLM on GPU 0 and TTS on GPU 1—rather than requiring a single GPU to hold all three models simultaneously. This prevents out-of-memory errors and allows deployment of larger unquantized models.

### Is the data transfer overhead between devices significant enough to avoid multi-device setups?

For short utterances, PCIe or NVLink transfer overhead can dominate total latency, potentially making same-device execution faster despite contention. However, for production workloads where the LLM dominates compute time or when using large models that exceed single-GPU memory, the benefits of eliminating lock contention and memory pressure typically outweigh transfer costs.

### Why does the pipeline block multiple pipelines on the same device unless in realtime mode?

The [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) guard (lines 10,846-850) prevents `--num_pipelines > 1` unless `mode==realtime` to avoid overwhelming shared GPU resources. Multiple pipelines on the same device create resource contention where the slowest stage bottlenecks all others, potentially causing system instability or out-of-memory conditions under load.