# How to Debug Latency Issues in the VAD-STT-LLM-TTS Pipeline: A Step-by-Step Guide

> Debug latency issues in VAD-STT-LLM-TTS pipelines with perf_counter timing and debug logging. Isolate delays by disabling optional features to optimize speech processing.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**To debug latency issues in the VAD-STT-LLM-TTS pipeline, instrument each handler’s `process` method with `perf_counter()` timing wrappers, enable debug logging at the library root, and progressively disable optional features like audio enhancement and torch compilation to isolate the specific stage causing delay.**

The huggingface/speech-to-speech repository implements a real-time conversational AI system as a linear processing chain. While the modular architecture separates concerns across Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) components, latency can accumulate at any stage. Understanding the exact code paths in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py), [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py), and the pipeline orchestrator is essential for targeted optimization.

## Understanding the VAD-STT-LLM-TTS Pipeline Architecture

The system processes audio through four sequential stages, each with specific latency characteristics:

- **Voice Activity Detection (VAD)** – The `VADHandler` class in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) isolates spoken segments and manages turn boundaries through the `_progressive_processing_pause` method.

- **Speech-to-Text (STT)** – The `WhisperSTTHandler` in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) handles model loading, optional torch compilation, and transcription inference.

- **Language Model (LLM)** – The `BaseOpenAICompatibleLanguageModel` in [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py) processes transcripts and generates responses, potentially streaming tokens.

- **Text-to-Speech (TTS)** – Handlers like `Qwen3TTSHandler` in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) synthesize audio from LLM outputs, including post-processing steps like resampling.

## Identifying Latency Bottlenecks by Stage

Each stage introduces specific delay patterns you can trace in the source code:

| Stage | Typical Sources of Delay | Key Code Locations |
|-------|------------------------|-------------------|
| **VAD** | Chunk size configuration, `realtime_processing_pause` calculations, DeepFilterNet audio enhancement | `VADHandler.process()` → `_process_realtime()`; lines 267-376 for `_progressive_processing_pause` |
| **STT** | Model warm-up routines, torch compilation overhead, feature extraction | `WhisperSTTHandler.setup()` for compilation; `WhisperSTTHandler.process()` for inference |
| **LLM** | Tokenization overhead, network round-trips for remote models, generation parameters | `BaseOpenAICompatibleLanguageModel._call()`; `ChatCompletionsLanguageModel` implementations |
| **TTS** | Model loading latency, audio post-processing and resampling | `Qwen3TTSHandler._synthesize()` and equivalent methods in other TTS handlers |

## Step-by-Step Debugging Workflow

### Enable Fine-Grained Logging

Start by configuring the Python logging module to capture timestamped events from each handler. The VAD handler already emits debug statements that indicate state changes and processing times.

```python
import logging

logging.basicConfig(level=logging.DEBUG)

```

Look for log lines prefixed with `VAD:` to observe chunk processing rates and identify where the pipeline stalls.

### Measure Per-Stage Timing with Overrides

Create timing wrappers around each handler’s `process` method to pinpoint exactly which stage dominates your latency budget. This pattern works for any handler in the pipeline.

```python
from time import perf_counter
import logging
from speech_to_speech.VAD.vad_handler import VADHandler

class TimedVADHandler(VADHandler):
    def process(self, audio_chunk):
        start = perf_counter()
        for out in super().process(audio_chunk):
            elapsed = perf_counter() - start
            logging.debug(f"VAD stage took {elapsed:.3f}s")
            start = perf_counter()
            yield out

```

Apply the same instrumentation to STT, LLM, and TTS handlers to generate a per-turn latency breakdown.

### Profile the VAD Progressive Pause

The VAD handler dynamically calculates pauses based on speech segment length through `_progressive_processing_pause`. If you observe delays between speech completion and processing initiation, modify the pause calculation logic.

```python

# Force a constant 200ms pause instead of progressive calculation

TimedVADHandler._progressive_processing_pause = lambda self, _: 0.2

```

Alternatively, adjust the `realtime_processing_pause` parameter (default 0.5s) in your handler initialization to reduce wait times.

### Check Audio Enhancement Overhead

When `audio_enhancement=True`, the VAD handler invokes DeepFilterNet via `_apply_audio_enhancement`, which adds resampling and model inference overhead. Temporarily disable this to isolate its impact.

```python
vad_handler = VADHandler()
vad_handler.setup(..., audio_enhancement=False)

```

Compare timing logs with enhancement enabled versus disabled to quantify the exact millisecond cost.

### Verify STT Model Warm-up and Compilation

The `WhisperSTTHandler` performs warm-up runs during setup, and torch compilation incurs CUDA graph capture costs. If latency spikes occur only on the first utterance, compilation is the culprit.

```python
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler

stt = WhisperSTTHandler()
stt.setup(model_name="distil-whisper/distil-large-v3", compile_mode=None)

```

Setting `compile_mode=None` skips graph compilation, trading warmup time for slightly slower per-inference performance. Check `WhisperSTTHandler.warmup()` for the dummy generation logic that prepares the model.

### Inspect Generation Kwargs

Large `max_new_tokens` values or aggressive sampling parameters increase inference time in both STT and LLM stages. Constrain these during profiling.

```python
from speech_to_speech.LLM.base_openai_compatible_language_model import BaseOpenAICompatibleLanguageModel

llm = BaseOpenAICompatibleLanguageModel()
llm.setup(gen_kwargs={"max_new_tokens": 64, "temperature": 0.7})

```

### Monitor GPU Utilization

Low GPU utilization during STT or LLM steps indicates CPU-side bottlenecks, typically in audio feature extraction or tokenization. Run `nvidia-smi` or use PyTorch diagnostics while processing audio.

```python
import torch

# Check allocated memory to confirm GPU engagement

print(f"GPU memory allocated: {torch.cuda.memory_allocated() / 1e6:.2f} MB")

```

### Review Turn-Handling Logic

The `SpeculativeTurnTracker` may delay turn finalization based on `speculative_reopen_ms` and `unanswered_reopen_ms` parameters. If latency occurs between speech end and LLM invocation, reduce these time windows in your pipeline configuration.

### End-to-End Timing

Add high-level instrumentation around the complete pipeline run in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) to measure total latency from audio capture to synthesized output.

```python
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
import time
import logging

pipeline = SpeechToSpeechPipeline(
    vad_handler=debug_vad,
    stt_handler=stt,
    llm_handler=llm,
    tts_handler=your_tts_handler,
)

start = time.time()
pipeline.run()
logging.info(f"Total pipeline latency: {time.time() - start:.3f}s")

```

## Practical Debugging Implementation

Combine the above techniques into a single debugging configuration:

```python
import logging
from time import perf_counter
from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler
from speech_to_speech.LLM.base_openai_compatible_language_model import BaseOpenAICompatibleLanguageModel

# Enable verbose logging

logging.basicConfig(level=logging.DEBUG)
logger = logging.getLogger(__name__)

# 1. Debug VAD with timing and reduced pause

class DebugVADHandler(VADHandler):
    def process(self, audio_chunk):
        start = perf_counter()
        for out in super().process(audio_chunk):
            logger.debug(f"VAD chunk latency: {perf_counter() - start:.4f}s")
            start = perf_counter()
            yield out

# Override progressive pause to constant 200ms

DebugVADHandler._progressive_processing_pause = lambda self, _: 0.2

vad = DebugVADHandler()
vad.setup(..., audio_enhancement=False)

# 2. Skip compilation in STT to eliminate first-run latency

stt = WhisperSTTHandler()
stt.setup(model_name="distil-whisper/distil-large-v3", compile_mode=None)

# 3. Constrain LLM generation

llm = BaseOpenAICompatibleLanguageModel()
llm.setup(gen_kwargs={"max_new_tokens": 64, "temperature": 0.7})

# 4. Instantiate pipeline and measure end-to-end time

from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline

pipeline = SpeechToSpeechPipeline(
    vad_handler=vad,
    stt_handler=stt,
    llm_handler=llm,
    tts_handler=...,
)

t0 = time.time()
pipeline.run()
logger.info(f"Pipeline completed in {time.time() - t0:.3f}s")

```

## Summary

- **Instrument each stage** by subclassing handlers and wrapping `process()` methods with `perf_counter()` to identify which component dominates latency.
- **Disable optional processing** such as DeepFilterNet audio enhancement and torch compilation to isolate their specific overhead costs.
- **Adjust VAD timing** by overriding `_progressive_processing_pause` or reducing `realtime_processing_pause` from its default 0.5s.
- **Monitor resource utilization** using GPU diagnostics to distinguish between model inference delays and CPU preprocessing bottlenecks.
- **Iterate with controlled tests** using the end-to-end timer in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) to validate improvements after each configuration change.

## Frequently Asked Questions

### How do I identify which stage is causing the most delay?

Subclass each handler (`VADHandler`, `WhisperSTTHandler`, etc.) and override the `process` method to log `perf_counter()` deltas before yielding results. Compare the logged timings across stages to identify the bottleneck. The VAD stage often hides latency in `_progressive_processing_pause` calculations, while STT delays typically occur during the first inference due to torch compilation.

### Why does the first utterance take longer than subsequent ones?

The STT handler performs model warm-up and optional torch compilation during initialization. According to the source code in [`whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_handler.py), the first real generation incurs CUDA graph capture costs when `torch.compile` is enabled. Set `compile_mode=None` during setup to eliminate this initial spike, or accept the one-time cost for better sustained throughput.

### Can I reduce latency without changing the underlying models?

Yes. Disable audio enhancement in the VAD handler by setting `audio_enhancement=False` to remove DeepFilterNet processing overhead. Reduce `realtime_processing_pause` below the default 0.5s, and constrain `max_new_tokens` in both STT and LLM handlers. These configuration changes reduce processing time without requiring model swaps or quantization.

### Where should I add logging to debug turn-handling delays?

Add instrumentation around the speculative turn tracker in the pipeline orchestration code. If you observe delays between speech completion and the LLM request, examine the `speculative_reopen_ms` and `unanswered_reopen_ms` parameters. These values control how long the system waits to confirm a turn boundary before forwarding text to the language model.