# MLX Lock Contention on Apple Silicon: How `--num_pipelines` Affects Realtime Performance in Speech-to-Speech

> Understand MLX lock contention on Apple Silicon impacting speech-to-speech realtime performance. Learn how --num_pipelines affects GPU access and prevents timeouts.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-07-30

---

**When running multiple concurrent pipelines on Apple Silicon, the global MLX lock serializes GPU access, forcing the huggingface/speech-to-speech library to disable live transcription to prevent timeout floods and UI freezing.**

The **MLX lock contention issue on Apple Silicon** arises from a fundamental hardware limitation in Metal’s GPU architecture. In the `huggingface/speech-to-speech` repository, all MLX-based models—including STT, LLM, and TTS backends—share a single global lock that ensures only one Metal command buffer executes per process. Understanding how the `--num_pipelines` argument interacts with this lock is critical for optimizing realtime performance on macOS.

## Understanding the MLX Lock Contention Issue on Apple Silicon

Apple Silicon GPUs can execute only one Metal command buffer per process at any given time. To enforce this serialization, the library implements a **global re-entrant lock** in [`src/speech_to_speech/utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py). All MLX-based inference operations must acquire this lock before accessing the GPU.

The lock itself is an `RLock` (re-entrant lock) wrapped with bookkeeping to track acquisition depth and hold time:

```python

# src/speech_to_speech/utils/mlx_lock.py lines 17-27

_mlx_lock = RLock()                     # global re‑entrant lock

_mlx_lock_state = Lock()                # protects bookkeeping variables

_lock_owner_ident: int | None = None
...
def acquire_mlx_lock(...):
    ...

```

Because the lock is re-entrant, a single thread can acquire it multiple times without deadlocking. However, when multiple threads or processes compete for the same lock, contention occurs.

## How `--num_pipelines` Triggers the Contention Problem

The `--num_pipelines` argument controls how many realtime pipelines run concurrently. When this value exceeds 1 on macOS, the global MLX lock becomes a bottleneck because every pipeline attempts to acquire it for each inference step.

The **progressive live-transcription path**—which provides low-latency STT updates while the user is still speaking—uses a short timeout when trying to acquire the lock. If the lock is unavailable, the work is dropped to maintain UI responsiveness. With multiple pipelines, this creates a flood of warnings and renders live transcription essentially useless.

To mitigate this, the pipeline startup code in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) automatically detects the problematic configuration:

```python

# src/speech_to_speech/s2s_pipeline.py lines 48-60

if args.module_kwargs.num_pipelines > 1 and platform == "darwin" and args.module_kwargs.enable_live_transcription:
    logger.info(
        "MLX contention: --num_pipelines=%d > 1 on Apple Silicon → disabling live transcription "
        "(progressive STT contends on the global MLX lock)",
        args.module_kwargs.num_pipelines,
    )
    args.module_kwargs.enable_live_transcription = False

```

This automatic fallback ensures the pipeline remains stable by disabling live transcription while preserving the final transcript output.

## Performance Implications of Pipeline Count

**Single Pipeline (`--num_pipelines 1`)**: The sole pipeline holds the MLX lock almost continuously during inference. Because no other pipelines compete for the lock, the progressive STT feature operates normally without timeouts or dropped frames.

**Multiple Pipelines (`--num_pipelines > 1`)**: Each pipeline must wait for the global lock to become available. The progressive STT’s aggressive timeout strategy fails repeatedly, triggering the automatic disablement mechanism described above. While the final transcript remains accurate, the realtime word-by-word display is suppressed to avoid console warning spam.

## Practical Examples and Workarounds

### Running a Single Realtime Pipeline (Live Transcription Enabled)

When running only one pipeline, live transcription works as expected because the global lock has no contenders:

```bash
speech-to-speech \
  --mode realtime \
  --num-pipelines 1 \
  --stt parakeet-tdt \
  --tts qwen3

```

The console displays live transcription updates while you speak, as the single pipeline maintains uninterrupted lock ownership.

### Running Multiple Pipelines on Apple Silicon (Auto-Disabled)

When requesting multiple pipelines on macOS, the library automatically disables live transcription to prevent lock contention:

```bash
speech-to-speech \
  --mode realtime \
  --num-pipelines 4 \
  --stt parakeet-tdt \
  --tts qwen3

```

The startup log displays:  
`MLX contention: --num_pipelines=4 > 1 on Apple Silicon → disabling live transcription …`  
No live transcription updates appear during the session, but the final transcript is still produced after the utterance ends.

### Forcing Live Transcription (Not Recommended)

You can override the automatic disablement, but this results in persistent timeout warnings:

```bash
speech-to-speech \
  --mode realtime \
  --num-pipelines 2 \
  --enable-live-transcription true

```

This configuration generates repeated warnings like `MLX lock acquisition timeout` from [`utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/utils/mlx_lock.py), confirming the unsustainable contention between pipelines.

## Summary

- **Apple Silicon Metal GPUs** process only one command buffer per process, necessitating a global MLX lock in [`utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/utils/mlx_lock.py).
- **`--num_pipelines > 1`** on macOS causes multiple realtime pipelines to compete for the same global lock.
- **Live transcription** uses short timeouts that fail under contention, prompting the library to automatically disable this feature when multiple pipelines are detected.
- **The global lock** is a re-entrant `RLock` that tracks ownership and duration, serialized across all MLX-based STT, LLM, and TTS operations.
- **Best practice** on Apple Silicon is to use `--num-pipelines 1` if live transcription is required, or accept the automatic fallback when scaling to multiple pipelines.

## Frequently Asked Questions

### What causes MLX lock contention on Apple Silicon?

The contention stems from Metal’s architectural limitation allowing only one command buffer per process. All MLX models in the speech-to-speech pipeline share a single global `RLock` defined in [`src/speech_to_speech/utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py). When multiple pipelines run concurrently, they serialize their GPU access through this lock, creating bottlenecks.

### Why does live transcription stop working with multiple pipelines?

The progressive live-transcription feature requires frequent, low-latency access to the MLX lock to update the UI while audio streams in. With multiple pipelines competing for the lock, the short acquisition timeouts expire, causing the feature to drop updates. The library prevents a flood of timeout warnings by automatically disabling live transcription when `--num_pipelines` exceeds 1 on macOS.

### Can I force live transcription with multiple pipelines on macOS?

Yes, by explicitly passing `--enable-live-transcription true`, but this is not recommended. Forcing the feature on results in continuous `MLX lock acquisition timeout` warnings and degraded user experience, as the UI updates fail silently or lag significantly while pipelines wait for the global lock.

### Is there a way to avoid the MLX lock limitation?

To avoid contention entirely, use `--num_pipelines 1` on Apple Silicon devices. Alternatively, use non-MLX backends (such as PyTorch or ONNX-based models) for the STT, LLM, or TTS components, as these do not rely on the Metal-specific global lock mechanism.