How to Debug Latency Issues in the Speech-to-Speech Pipeline: A Complete Guide

Debugging latency in the Hugging Face speech-to-speech pipeline requires monitoring thread-safe queue sizes, identifying MLX lock contention on macOS, and optimizing the VAD→STT→LLM→TTS handler chain using specific configuration flags and bounded queue implementations.

The huggingface/speech-to-speech repository provides a real-time audio processing system that chains Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) through thread-safe queues. When latency spikes occur, they typically manifest as queue back-pressure, GPU contention, or blocking operations within the handler stages defined in src/speech_to_speech/s2s_pipeline.py.

Understanding the Pipeline Architecture

The speech-to-speech system initializes its data flow through initialize_queues_and_events() in src/speech_to_speech/s2s_pipeline.py#L34-L48. This function creates unbounded queues and threading events that coordinate the six-stage pipeline:

  1. VADHandler (src/speech_to_speech/VAD/vad_handler.py) detects voice activity
  2. STT handlers (various implementations in src/speech_to_speech/STT/) transcribe audio
  3. TranscriptionNotifier (src/speech_to_speech/STT/transcription_notifier.py) forwards partial transcripts
  4. LLM handlers (src/speech_to_speech/LLM/language_model.py) generate responses
  5. LMOutputProcessor (src/speech_to_speech/LLM/lm_output_processor.py) formats output
  6. TTS handlers (various implementations in src/speech_to_speech/TTS/) synthesize audio

Each stage connects via queues defined in the initialization dictionary:

def initialize_queues_and_events() -> dict[str, Any]:
    return {
        "stop_event": Event(),
        "should_listen": Event(),
        "response_playing": Event(),
        "recv_audio_chunks_queue": Queue[AudioInItem](),      # Unbounded

        "send_audio_chunks_queue": Queue[AudioOutItem](),     # Unbounded

        "spoken_prompt_queue": Queue[VADOutItem](),
        "stt_output_queue": Queue[STTOutItem](),
        "text_prompt_queue": Queue[TextPromptItem](),
        "lm_response_queue": Queue[LMOutItem](),
        "lm_processed_queue": Queue[TTSInItem](),
        "text_output_queue": Queue[TextEventItem](),
    }

The unbounded queue design masks back-pressure but risks memory bloat when downstream stages stall. The threading.Event objects (should_listen, response_playing) gate audio chunk flow and are critical synchronization points for latency analysis.

Common Latency Sources

Latency typically emerges from four specific architectural constraints:

Queue Back-Pressure When TTS processing is slower than LLM generation, the lm_processed_queue fills indefinitely. Because queues default to unbounded sizes in initialize_queues_and_events(), memory usage grows linearly while latency remains hidden until system resources exhaust.

MLX Global Lock Contention On Apple Silicon, the global MLX lock (src/speech_to_speech/utils/mlx_lock.py) serializes all GPU inference. When running multiple pipelines (--num_pipelines > 1) with live transcription enabled, STT calls block waiting for the lock, creating cascading delays. The repository automatically disables live transcription in this scenario (lines 22-29 of s2s_pipeline.py), but manual verification is essential.

Speculative Turn Overhead The SpeculativeTurnTracker (src/speech_to_speech/pipeline/speculative_turns.py) pre-emptively starts LLM inference based on partial STT results. Misaligned speculation triggers cancellation overhead, adding jitter to the response pipeline.

Model Warm-Up Delays Heavy STT models like Whisper or Faster-Whisper exhibit significant first-inference latency. The VADHandler emits spoken_prompt_queue items only after detecting voice activity, but subsequent STT initialization delays create perceptible gaps before transcription begins.

Step-by-Step Debugging Procedure

1. Enable Queue Size Monitoring

Add this instrumentation to track real-time queue depths:

import threading
import time
from queue import Queue

def monitor_queues(queues: dict, stop_event: Event):
    """Monitor all queue sizes every 500ms to identify bottlenecks."""
    while not stop_event.is_set():
        for name, q in queues.items():
            if isinstance(q, Queue):
                logger.debug(f"Queue {name} size={q.qsize()}")
        time.sleep(0.5)

# Start monitoring thread after pipeline initialization

threading.Thread(
    target=monitor_queues, 
    args=(queues_and_events, queues_and_events["stop_event"]), 
    daemon=True
).start()

Expected thresholds:

  • recv_audio_chunks_queue: ≤ 5 items
  • stt_output_queue: ≤ 1 items
  • lm_response_queue: ≤ 2 items
  • lm_processed_queue: ≤ 1 items

2. Isolate the Pipeline Configuration

Run a minimal single-pipeline configuration to eliminate cross-pipeline contention:

python -m speech_to_speech.s2s_pipeline \
    --mode local \
    --stt whisper \
    --tts pocket \
    --llm_backend mlx-lm \
    --log_level debug \
    --num_pipelines 1 \
    --enable_live_transcription False

The --num_pipelines 1 flag prevents MLX lock contention, while --enable_live_transcription False eliminates speculative STT calls that compete with the LLM for GPU resources.

3. Measure Handler Execution Time

Wrap handler calls in src/speech_to_speech/s2s_pipeline.py to identify slow stages:

import time

# Inside the pipeline processing loop

start = time.time()
handler.process(item)
elapsed = time.time() - start
logger.debug(f"{handler.__class__.__name__} elapsed={elapsed:.3f}s")

Target benchmarks: Each handler should complete within 200ms for lightweight models. Values exceeding this indicate the bottleneck stage.

4. Check Platform-Specific Resources

macOS (MLX): Monitor the global lock using Activity Monitor. If CPU usage spikes while GPU utilization drops, the process is waiting on the MLX lock in src/speech_to_speech/utils/mlx_lock.py.

Linux (CUDA): Use nvidia-smi to verify GPU memory isn't saturated. TTS models like Qwen3TTSHandler (src/speech_to_speech/TTS/qwen3_tts_handler.py) can exhaust VRAM, causing fallback to CPU processing.

Code Solutions for Latency Reduction

Converting Unbounded Queues to Bounded Queues

Modify initialize_queues_and_events() to surface back-pressure immediately:

def initialize_queues_and_events() -> dict[str, Any]:
    return {
        "recv_audio_chunks_queue": Queue(maxsize=10),
        "send_audio_chunks_queue": Queue(maxsize=10),
        "spoken_prompt_queue": Queue(maxsize=5),
        "stt_output_queue": Queue(maxsize=2),
        "lm_response_queue": Queue(maxsize=3),
        "lm_processed_queue": Queue(maxsize=2),
        # ... remaining queues

    }

When a bounded queue fills, the upstream producer blocks, making the latency source explicit in logs rather than consuming unlimited memory.

Disabling Speculative Turns

If the SpeculativeTurnTracker causes jitter, disable it by passing None to the pipeline builder:


# In your pipeline initialization

speculative_turns = None  # Instead of SpeculativeTurnTracker()

# Or via CLI flags for supported configurations

--speculative_turns false

This eliminates pre-emptive LLM inference and cancellation overhead at the cost of increased perceived response time.

Implementing End-to-End Latency Measurement

import time
from speech_to_speech.s2s_pipeline import (
    parse_arguments, 
    initialize_queues_and_events, 
    build_pipeline
)

def benchmark_pipeline():
    args = parse_arguments()
    args.module_kwargs.num_pipelines = 1
    args.module_kwargs.enable_live_transcription = False
    
    queues = initialize_queues_and_events()
    pipeline = build_pipeline(
        args.module_kwargs,
        args.socket_receiver_kwargs,
        args.socket_sender_kwargs,
        args.websocket_streamer_kwargs,
        args.vad_handler_kwargs,
        args.whisper_stt_handler_kwargs,
        args.faster_whisper_stt_handler_kwargs,
        args.paraformer_stt_handler_kwargs,
        args.mlx_audio_whisper_stt_handler_kwargs,
        args.parakeet_tdt_stt_handler_kwargs,
        args.language_model_handler_kwargs,
        args.responses_api_language_model_handler_kwargs,
        args.chat_tts_handler_kwargs,
        args.facebook_mms_tts_handler_kwargs,
        args.pocket_tts_handler_kwargs,
        args.kokoro_tts_handler_kwargs,
        args.qwen3_tts_handler_kwargs,
        queues,
    )
    
    start = time.perf_counter()
    pipeline.start()
    
    # Run test audio through recv_audio_chunks_queue

    # ... test logic ...

    
    pipeline.stop()
    elapsed = time.perf_counter() - start
    print(f"End-to-end latency: {elapsed:.3f}s")

Summary

  • Queue monitoring is essential: Unbounded queues in initialize_queues_and_events() hide latency until memory exhaustion occurs; use bounded queues with maxsize parameters to surface back-pressure immediately.
  • MLX lock contention on Apple Silicon requires single-pipeline operation (--num_pipelines 1) and disabling live transcription to prevent GPU serialization bottlenecks.
  • Speculative turns in src/speech_to_speech/pipeline/speculative_turns.py can add jitter; disable them by setting speculative_turns=None if cancellation overhead exceeds benefits.
  • Handler benchmarks should target <200ms per stage; wrap handler.process() calls with timers in src/speech_to_speech/s2s_pipeline.py to identify slow components.
  • Lightweight model selection (e.g., pocket TTS instead of qwen3, parakeet-tdt STT via --local_mac_optimal_settings) reduces queue buildup and processing delays.

Frequently Asked Questions

How do I identify which pipeline stage is causing latency?

Add timing instrumentation around each handler call in src/speech_to_speech/s2s_pipeline.py and monitor queue sizes using the monitor_queues function provided above. The stage with the longest execution time or the queue that consistently grows indicates your bottleneck. For example, if lm_processed_queue.qsize() consistently exceeds 2 items while stt_output_queue remains empty, the TTS handler is the limiting factor.

Why does macOS show higher latency when running multiple pipelines?

Apple Silicon uses a global MLX lock implemented in src/speech_to_speech/utils/mlx_lock.py to prevent GPU conflicts. When --num_pipelines > 1, multiple STT and LLM instances compete for this single lock, causing serial execution rather than parallel processing. The repository automatically disables live transcription in this configuration (lines 22-29 of s2s_pipeline.py), but you should explicitly use --num_pipelines 1 for optimal latency.

How do I reduce memory usage from queue buildup?

Convert the default unbounded queues in initialize_queues_and_events() to bounded queues by adding maxsize parameters (e.g., Queue(maxsize=10)). When a bounded queue fills, the upstream producer blocks, which prevents unlimited memory consumption and makes the bottleneck visible in your logs. Combine this with lighter model selection (e.g., --tts pocket instead of --tts qwen3) to increase throughput.

What is the difference between live transcription and speculative turns?

Live transcription (controlled by --enable_live_transcription) forwards partial STT results to the LLM before the user finishes speaking, but requires additional GPU resources that may conflict with the MLX lock on macOS. Speculative turns (managed by SpeculativeTurnTracker in src/speech_to_speech/pipeline/speculative_turns.py) pre-emptively start LLM inference based on partial results, then cancel if the final transcription differs. Live transcription affects STT→LLM latency, while speculative turns affect LLM processing efficiency and cancellation overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →