MLX Contention on Apple Silicon in Speech-to-Speech: How --num_pipelines Interacts with It

MLX contention arises when multiple inference pipelines compete for Apple’s single global Metal command queue on macOS, causing crashes or severe latency, and the --num_pipelines flag triggers an automatic guard in s2s_pipeline.py that forces single-pipeline mode and disables live transcription on Apple Silicon to prevent this conflict.

The speech-to-speech repository by Hugging Face enables real-time voice conversion on Apple Silicon using MLX, Apple’s Metal-accelerated machine learning framework. Because MLX relies on a single global Metal command queue that is not thread-safe, attempting to run multiple concurrent pipelines creates resource contention that can destabilize the application. This article examines the root cause of MLX contention on Apple Silicon and explains how the --num_pipelines argument interacts with the live transcription feature to maintain system stability.

What Is MLX Contention on Apple Silicon?

The Metal Queue Bottleneck

MLX utilizes Apple’s Metal Performance Shaders for GPU acceleration through a single global command queue. Unlike CUDA or other backends that support concurrent streams, MLX’s architecture serializes access to Metal resources. When the speech-to-speech pipeline spawns multiple workers, each attempts to acquire the same global queue lock simultaneously. This competition for the lock is MLX contention, which manifests as thread-blocking, memory pressure, or runtime crashes.

Symptoms of Resource Contention

When contention occurs on Apple Silicon devices, you may observe:

  • Dramatic inference slowdowns (latency spikes)
  • Out-of-memory errors in the Metal driver
  • Application crashes within the MLX runtime
  • Audio dropouts during live transcription

How --num_pipelines Interacts with MLX Contention

The Guard Clause in s2s_pipeline.py

The repository implements a protective guard in src/speech_to_speech/s2s_pipeline.py (around line 1022) that detects problematic configurations:

if args.module_kwargs.num_pipelines > 1 and platform == "darwin" \
        and args.module_kwargs.enable_live_transcription:
    # MLX contention: --num_pipelines=%d > 1 on Apple Silicon → disabling live transcription

    args.module_kwargs.num_pipelines = 1

This condition checks three criteria:

  1. --num_pipelines is set greater than 1
  2. The platform is Darwin (macOS/Apple Silicon)
  3. Live transcription is enabled

When all three are true, the code automatically reduces the pipeline count to 1 and disables live transcription to avoid the contention bottleneck.

Automatic Mitigation Behavior

The interaction between these flags follows a strict hierarchy to preserve stability:

  • Single pipeline (--num_pipelines 1): Uses the MLX lock in src/speech_to_speech/utils/mlx_lock.py to safely access Metal resources; live transcription operates normally.
  • Multiple pipelines (--num_pipelines > 1): Allowed only when live transcription is disabled; the system serializes MLX access through the process-wide _mlx_lock or runs pipelines sequentially.

Configuring Pipelines for Stable Performance

Live Transcription Mode (Single Pipeline)

For real-time speech-to-speech conversion on Apple Silicon, always use a single pipeline:

python -m speech_to_speech \
  --mode realtime \
  --enable_live_transcription \
  --num_pipelines 1

This configuration ensures exclusive access to the Metal queue, preventing MLX contention while maintaining low-latency transcription.

Batch Processing Mode (Multiple Pipelines)

If you need throughput for offline batch processing, disable live transcription to permit multiple pipelines:

python -m speech_to_speech \
  --mode file \
  --disable_live_transcription \
  --num_pipelines 4 \
  --input_dir ./audio_files/

In this mode, the s2s_pipeline.py guard does not trigger, and the system manages resource contention through the MLX lock mechanism or sequential processing.

Key Source Files

The following files implement the MLX contention detection and mitigation logic:

Summary

  • MLX contention occurs when multiple pipelines compete for Apple’s single global Metal command queue on macOS, causing instability.
  • The --num_pipelines flag directly triggers this issue when set greater than 1 on Apple Silicon with live transcription enabled.
  • The repository automatically mitigates contention by forcing num_pipelines = 1 and disabling live transcription when the problematic combination is detected in s2s_pipeline.py.
  • For stable real-time performance on Apple Silicon, use --num_pipelines 1; for batch processing, disable live transcription to safely use higher pipeline counts.
  • The mlx_lock.py utility provides coarse-grained locking as a secondary defense against Metal queue conflicts.

Frequently Asked Questions

What exactly is MLX contention?

MLX contention is a resource conflict that occurs when multiple threads or processes attempt to access Apple’s Metal command queue simultaneously through the MLX framework. Because MLX uses a single global queue that is not thread-safe, concurrent access causes blocking, memory errors, or crashes on Apple Silicon.

Why does --num_pipelines cause issues specifically on Apple Silicon?

Apple Silicon relies on the Metal graphics API for GPU compute, and MLX serializes all operations through one global queue. When --num_pipelines is greater than 1, each pipeline attempts to submit commands concurrently, creating a bottleneck that does not occur on CUDA or CPU backends where parallel streams are supported.

How can I run multiple pipelines on macOS without triggering MLX contention?

You can safely use --num_pipelines 2 or higher on macOS only when live transcription is disabled. The guard in s2s_pipeline.py permits multiple pipelines for batch processing tasks, as the absence of live transcription avoids the real-time Metal queue conflicts that cause crashes.

Does the MLX lock eliminate contention entirely?

The _mlx_lock in mlx_lock.py mitigates contention by serializing access to the Metal queue, but it does not eliminate the underlying limitation. It prevents crashes by ensuring only one pipeline uses MLX at a time, but this effectively reduces concurrent pipelines to serialized execution, which is why the repository defaults to forcing single-pipeline mode for live transcription.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →