Performance Implications of Running STT, LLM, and TTS on Same vs Separate Devices

Running STT, LLM, and TTS on separate devices eliminates GPU memory contention and MLX lock serialization, but introduces PCIe transfer overhead that may dominate latency for short utterances.

The huggingface/speech-to-speech repository implements a fully modular speech-to-speech pipeline where Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) stages run as independent handlers communicating via thread-safe queues. Understanding the performance implications of running STT LLM TTS on same vs separate devices requires analyzing resource contention, memory pressure, and data transfer overhead in the pipeline architecture defined in src/speech_to_speech/s2s_pipeline.py.

Pipeline Architecture and Device Allocation

The central orchestration in s2s_pipeline.py instantiates each pipeline stage as a handler running in its own thread, with the overwrite_device_argument function (lines 63-72) managing device assignment. By default, all handlers inherit the global --device argument, but the system supports per-handler device overrides through arguments defined in arguments_classes/module_arguments.py, arguments_classes/whisper_stt_arguments.py, arguments_classes/language_model_arguments.py, and arguments_classes/qwen3_tts_arguments.py.

Same-Device Execution: Low Latency, High Contention

Resource Contention and the MLX Lock

When running STT, LLM, and TTS on the same device, all handlers share a single GPU or Apple Silicon chip. On macOS, the global MLX lock implemented in utils/mlx_lock.py serializes inference operations, causing a flood of warnings and significant slowdowns for progressive STT→LLM→TTS flows. This serialization forces the pipeline to process stages sequentially rather than concurrently, eliminating the benefits of the multi-threaded architecture.

Memory Pressure Constraints

A single GPU must simultaneously hold weights for Whisper (STT), the LLM (e.g., Qwen-3-4B consuming >10GB), and the TTS model. This concentration often exceeds VRAM capacity, triggering out-of-memory crashes or forcing the use of quantized models that sacrifice quality. The s2s_pipeline.py guard blocks --num_pipelines > 1 unless mode==realtime (lines 10,846-850) to prevent overload on shared devices.

Zero-Copy Data Transfer

Same-device execution benefits from in-memory queue communication through Queue[AudioInItem] instances, eliminating inter-device copy latency. Audio tensors flow directly between handlers without PCIe or NVLink transfer costs, minimizing per-utterance latency for short interactions.

Separate-Device Execution: Scalability with Overhead

Eliminating Resource Contention

Assigning distinct devices to each handler—such as STT on CPU, LLM on GPU 0, and TTS on GPU 1—removes the global lock contention present in same-device configurations. On non-macOS platforms, this allows true concurrent execution where the LLM can process speculative turns via pipeline/speculative_turns.py while STT handles new audio input.

Memory Distribution Benefits

Distributing models across devices spreads memory usage, enabling simultaneous deployment of larger unquantized models. Each device only stores its assigned model weights, preventing the out-of-memory conditions common in single-GPU setups and allowing high-throughput deployments with multiple parallel pipelines.

Inter-Device Transfer Costs

Moving audio features or token tensors between devices incurs PCIe or NVLink copy costs. For short utterances, this transfer latency can dominate total response time, potentially negating the benefits of parallel processing. The overhead scales with tensor size and bus bandwidth, making this architecture optimal for longer-form content where compute time exceeds transfer time.

Platform-Specific Considerations

On macOS, the code forces --llm_backend mlx-lm and applies check_mac_settings and optimal_mac_settings that automatically disable live transcription when using multiple pipelines to avoid MLX lock contention (lines 48-61). Linux and Windows platforms do not suffer from this global lock limitation, making multi-device setups more effective for exploiting multiple GPUs.

Configuration Examples

Single-GPU Deployment

Run all stages on one CUDA device using the default inheritance mechanism:

python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --device cuda \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3

The overwrite_device_argument function (lines 63-72) ensures all handlers inherit the cuda device assignment.

Multi-Device Distribution

Assign specific devices to each handler to eliminate contention:

python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --device cpu \
    --stt_device cpu \
    --llm_device cuda \
    --tts_device cuda:1 \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3

Per-handler arguments (--stt_device, --llm_device, --tts_device) override the global device setting via the merging logic in overwrite_device_argument.

macOS Multi-Pipeline Limitations

Disable live transcription on Apple Silicon to avoid lock contention:

python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --num_pipelines 2 \
    --local_mac_optimal_settings true

This configuration detects macOS and disables features that trigger the global MLX lock serialization.

Summary

  • Same-device execution minimizes data transfer latency through in-memory queues but suffers from GPU memory pressure and, on Apple Silicon, MLX lock serialization that forces sequential processing.
  • Separate-device execution eliminates resource contention and distributes memory load, enabling larger models and higher throughput, but introduces PCIe/NVLink transfer overhead that impacts short-utterance latency.
  • Platform-specific locks on macOS make multi-device setups more beneficial than on Linux/Windows, where the global MLX lock does not constrain same-device multi-pipeline deployments.
  • Configuration flexibility via s2s_pipeline.py allows granular device assignment through per-handler arguments, supporting both simple single-GPU deployments and complex multi-GPU orchestrations.

Frequently Asked Questions

What causes the MLX lock warnings when running all pipeline stages on a Mac?

The global MLX lock in utils/mlx_lock.py serializes inference operations on Apple Silicon, forcing STT, LLM, and TTS handlers to execute sequentially rather than concurrently. This lock generates warnings when multiple handlers attempt simultaneous access to the MLX device, causing performance degradation in progressive speech-to-speech flows.

How does splitting STT, LLM, and TTS across devices affect memory requirements?

Distributing models across separate devices reduces per-device memory requirements significantly. Each device only needs to load its assigned model weights—for example, the LLM on GPU 0 and TTS on GPU 1—rather than requiring a single GPU to hold all three models simultaneously. This prevents out-of-memory errors and allows deployment of larger unquantized models.

Is the data transfer overhead between devices significant enough to avoid multi-device setups?

For short utterances, PCIe or NVLink transfer overhead can dominate total latency, potentially making same-device execution faster despite contention. However, for production workloads where the LLM dominates compute time or when using large models that exceed single-GPU memory, the benefits of eliminating lock contention and memory pressure typically outweigh transfer costs.

Why does the pipeline block multiple pipelines on the same device unless in realtime mode?

The s2s_pipeline.py guard (lines 10,846-850) prevents --num_pipelines > 1 unless mode==realtime to avoid overwhelming shared GPU resources. Multiple pipelines on the same device create resource contention where the slowest stage bottlenecks all others, potentially causing system instability or out-of-memory conditions under load.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →