# How to Debug Latency Issues in the Speech-to-Speech Pipeline: A Complete Guide

> Debug speech-to-speech pipeline latency by monitoring queues optimizing the VAD STT LLM TTS handler chain and identifying MLX lock contention on macOS Learn how to speed up your ASR.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-09

---

**Debugging latency in the Hugging Face speech-to-speech pipeline requires monitoring thread-safe queue sizes, identifying MLX lock contention on macOS, and optimizing the VAD→STT→LLM→TTS handler chain using specific configuration flags and bounded queue implementations.**

The huggingface/speech-to-speech repository provides a real-time audio processing system that chains Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) through thread-safe queues. When latency spikes occur, they typically manifest as queue back-pressure, GPU contention, or blocking operations within the handler stages defined in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).

## Understanding the Pipeline Architecture

The speech-to-speech system initializes its data flow through `initialize_queues_and_events()` in `src/speech_to_speech/s2s_pipeline.py#L34-L48`. This function creates unbounded queues and threading events that coordinate the six-stage pipeline:

1. **VADHandler** ([`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py)) detects voice activity
2. **STT handlers** (various implementations in `src/speech_to_speech/STT/`) transcribe audio
3. **TranscriptionNotifier** ([`src/speech_to_speech/STT/transcription_notifier.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/transcription_notifier.py)) forwards partial transcripts
4. **LLM handlers** ([`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py)) generate responses
5. **LMOutputProcessor** ([`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py)) formats output
6. **TTS handlers** (various implementations in `src/speech_to_speech/TTS/`) synthesize audio

Each stage connects via queues defined in the initialization dictionary:

```python
def initialize_queues_and_events() -> dict[str, Any]:
    return {
        "stop_event": Event(),
        "should_listen": Event(),
        "response_playing": Event(),
        "recv_audio_chunks_queue": Queue[AudioInItem](),      # Unbounded

        "send_audio_chunks_queue": Queue[AudioOutItem](),     # Unbounded

        "spoken_prompt_queue": Queue[VADOutItem](),
        "stt_output_queue": Queue[STTOutItem](),
        "text_prompt_queue": Queue[TextPromptItem](),
        "lm_response_queue": Queue[LMOutItem](),
        "lm_processed_queue": Queue[TTSInItem](),
        "text_output_queue": Queue[TextEventItem](),
    }

```

The **unbounded queue design** masks back-pressure but risks memory bloat when downstream stages stall. The `threading.Event` objects (`should_listen`, `response_playing`) gate audio chunk flow and are critical synchronization points for latency analysis.

## Common Latency Sources

Latency typically emerges from four specific architectural constraints:

**Queue Back-Pressure**
When TTS processing is slower than LLM generation, the `lm_processed_queue` fills indefinitely. Because queues default to unbounded sizes in `initialize_queues_and_events()`, memory usage grows linearly while latency remains hidden until system resources exhaust.

**MLX Global Lock Contention**
On Apple Silicon, the global MLX lock ([`src/speech_to_speech/utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py)) serializes all GPU inference. When running multiple pipelines (`--num_pipelines > 1`) with live transcription enabled, STT calls block waiting for the lock, creating cascading delays. The repository automatically disables live transcription in this scenario (lines 22-29 of [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)), but manual verification is essential.

**Speculative Turn Overhead**
The `SpeculativeTurnTracker` ([`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py)) pre-emptively starts LLM inference based on partial STT results. Misaligned speculation triggers cancellation overhead, adding jitter to the response pipeline.

**Model Warm-Up Delays**
Heavy STT models like Whisper or Faster-Whisper exhibit significant first-inference latency. The `VADHandler` emits `spoken_prompt_queue` items only after detecting voice activity, but subsequent STT initialization delays create perceptible gaps before transcription begins.

## Step-by-Step Debugging Procedure

### 1. Enable Queue Size Monitoring

Add this instrumentation to track real-time queue depths:

```python
import threading
import time
from queue import Queue

def monitor_queues(queues: dict, stop_event: Event):
    """Monitor all queue sizes every 500ms to identify bottlenecks."""
    while not stop_event.is_set():
        for name, q in queues.items():
            if isinstance(q, Queue):
                logger.debug(f"Queue {name} size={q.qsize()}")
        time.sleep(0.5)

# Start monitoring thread after pipeline initialization

threading.Thread(
    target=monitor_queues, 
    args=(queues_and_events, queues_and_events["stop_event"]), 
    daemon=True
).start()

```

**Expected thresholds:**
- `recv_audio_chunks_queue`: ≤ 5 items
- `stt_output_queue`: ≤ 1 items  
- `lm_response_queue`: ≤ 2 items
- `lm_processed_queue`: ≤ 1 items

### 2. Isolate the Pipeline Configuration

Run a minimal single-pipeline configuration to eliminate cross-pipeline contention:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode local \
    --stt whisper \
    --tts pocket \
    --llm_backend mlx-lm \
    --log_level debug \
    --num_pipelines 1 \
    --enable_live_transcription False

```

The `--num_pipelines 1` flag prevents MLX lock contention, while `--enable_live_transcription False` eliminates speculative STT calls that compete with the LLM for GPU resources.

### 3. Measure Handler Execution Time

Wrap handler calls in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) to identify slow stages:

```python
import time

# Inside the pipeline processing loop

start = time.time()
handler.process(item)
elapsed = time.time() - start
logger.debug(f"{handler.__class__.__name__} elapsed={elapsed:.3f}s")

```

**Target benchmarks:** Each handler should complete within **200ms** for lightweight models. Values exceeding this indicate the bottleneck stage.

### 4. Check Platform-Specific Resources

**macOS (MLX):** Monitor the global lock using Activity Monitor. If CPU usage spikes while GPU utilization drops, the process is waiting on the MLX lock in [`src/speech_to_speech/utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py).

**Linux (CUDA):** Use `nvidia-smi` to verify GPU memory isn't saturated. TTS models like `Qwen3TTSHandler` ([`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)) can exhaust VRAM, causing fallback to CPU processing.

## Code Solutions for Latency Reduction

### Converting Unbounded Queues to Bounded Queues

Modify `initialize_queues_and_events()` to surface back-pressure immediately:

```python
def initialize_queues_and_events() -> dict[str, Any]:
    return {
        "recv_audio_chunks_queue": Queue(maxsize=10),
        "send_audio_chunks_queue": Queue(maxsize=10),
        "spoken_prompt_queue": Queue(maxsize=5),
        "stt_output_queue": Queue(maxsize=2),
        "lm_response_queue": Queue(maxsize=3),
        "lm_processed_queue": Queue(maxsize=2),
        # ... remaining queues

    }

```

When a bounded queue fills, the upstream producer blocks, making the latency source explicit in logs rather than consuming unlimited memory.

### Disabling Speculative Turns

If the `SpeculativeTurnTracker` causes jitter, disable it by passing `None` to the pipeline builder:

```python

# In your pipeline initialization

speculative_turns = None  # Instead of SpeculativeTurnTracker()

# Or via CLI flags for supported configurations

--speculative_turns false

```

This eliminates pre-emptive LLM inference and cancellation overhead at the cost of increased perceived response time.

### Implementing End-to-End Latency Measurement

```python
import time
from speech_to_speech.s2s_pipeline import (
    parse_arguments, 
    initialize_queues_and_events, 
    build_pipeline
)

def benchmark_pipeline():
    args = parse_arguments()
    args.module_kwargs.num_pipelines = 1
    args.module_kwargs.enable_live_transcription = False
    
    queues = initialize_queues_and_events()
    pipeline = build_pipeline(
        args.module_kwargs,
        args.socket_receiver_kwargs,
        args.socket_sender_kwargs,
        args.websocket_streamer_kwargs,
        args.vad_handler_kwargs,
        args.whisper_stt_handler_kwargs,
        args.faster_whisper_stt_handler_kwargs,
        args.paraformer_stt_handler_kwargs,
        args.mlx_audio_whisper_stt_handler_kwargs,
        args.parakeet_tdt_stt_handler_kwargs,
        args.language_model_handler_kwargs,
        args.responses_api_language_model_handler_kwargs,
        args.chat_tts_handler_kwargs,
        args.facebook_mms_tts_handler_kwargs,
        args.pocket_tts_handler_kwargs,
        args.kokoro_tts_handler_kwargs,
        args.qwen3_tts_handler_kwargs,
        queues,
    )
    
    start = time.perf_counter()
    pipeline.start()
    
    # Run test audio through recv_audio_chunks_queue

    # ... test logic ...

    
    pipeline.stop()
    elapsed = time.perf_counter() - start
    print(f"End-to-end latency: {elapsed:.3f}s")

```

## Summary

- **Queue monitoring** is essential: Unbounded queues in `initialize_queues_and_events()` hide latency until memory exhaustion occurs; use bounded queues with `maxsize` parameters to surface back-pressure immediately.
- **MLX lock contention** on Apple Silicon requires single-pipeline operation (`--num_pipelines 1`) and disabling live transcription to prevent GPU serialization bottlenecks.
- **Speculative turns** in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py) can add jitter; disable them by setting `speculative_turns=None` if cancellation overhead exceeds benefits.
- **Handler benchmarks** should target <200ms per stage; wrap `handler.process()` calls with timers in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) to identify slow components.
- **Lightweight model selection** (e.g., `pocket` TTS instead of `qwen3`, `parakeet-tdt` STT via `--local_mac_optimal_settings`) reduces queue buildup and processing delays.

## Frequently Asked Questions

### How do I identify which pipeline stage is causing latency?

Add timing instrumentation around each handler call in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) and monitor queue sizes using the `monitor_queues` function provided above. The stage with the longest execution time or the queue that consistently grows indicates your bottleneck. For example, if `lm_processed_queue.qsize()` consistently exceeds 2 items while `stt_output_queue` remains empty, the TTS handler is the limiting factor.

### Why does macOS show higher latency when running multiple pipelines?

Apple Silicon uses a **global MLX lock** implemented in [`src/speech_to_speech/utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py) to prevent GPU conflicts. When `--num_pipelines > 1`, multiple STT and LLM instances compete for this single lock, causing serial execution rather than parallel processing. The repository automatically disables live transcription in this configuration (lines 22-29 of [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)), but you should explicitly use `--num_pipelines 1` for optimal latency.

### How do I reduce memory usage from queue buildup?

Convert the default unbounded queues in `initialize_queues_and_events()` to bounded queues by adding `maxsize` parameters (e.g., `Queue(maxsize=10)`). When a bounded queue fills, the upstream producer blocks, which prevents unlimited memory consumption and makes the bottleneck visible in your logs. Combine this with lighter model selection (e.g., `--tts pocket` instead of `--tts qwen3`) to increase throughput.

### What is the difference between live transcription and speculative turns?

**Live transcription** (controlled by `--enable_live_transcription`) forwards partial STT results to the LLM before the user finishes speaking, but requires additional GPU resources that may conflict with the MLX lock on macOS. **Speculative turns** (managed by `SpeculativeTurnTracker` in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py)) pre-emptively start LLM inference based on partial results, then cancel if the final transcription differs. Live transcription affects STT→LLM latency, while speculative turns affect LLM processing efficiency and cancellation overhead.