How to Debug Latency Issues in the VAD-STT-LLM-TTS Pipeline: A Step-by-Step Guide

To debug latency issues in the VAD-STT-LLM-TTS pipeline, instrument each handler’s process method with perf_counter() timing wrappers, enable debug logging at the library root, and progressively disable optional features like audio enhancement and torch compilation to isolate the specific stage causing delay.

The huggingface/speech-to-speech repository implements a real-time conversational AI system as a linear processing chain. While the modular architecture separates concerns across Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) components, latency can accumulate at any stage. Understanding the exact code paths in src/speech_to_speech/VAD/vad_handler.py, src/speech_to_speech/STT/whisper_stt_handler.py, and the pipeline orchestrator is essential for targeted optimization.

Understanding the VAD-STT-LLM-TTS Pipeline Architecture

The system processes audio through four sequential stages, each with specific latency characteristics:

Identifying Latency Bottlenecks by Stage

Each stage introduces specific delay patterns you can trace in the source code:

Stage Typical Sources of Delay Key Code Locations
VAD Chunk size configuration, realtime_processing_pause calculations, DeepFilterNet audio enhancement VADHandler.process() → _process_realtime(); lines 267-376 for _progressive_processing_pause
STT Model warm-up routines, torch compilation overhead, feature extraction WhisperSTTHandler.setup() for compilation; WhisperSTTHandler.process() for inference
LLM Tokenization overhead, network round-trips for remote models, generation parameters BaseOpenAICompatibleLanguageModel._call(); ChatCompletionsLanguageModel implementations
TTS Model loading latency, audio post-processing and resampling Qwen3TTSHandler._synthesize() and equivalent methods in other TTS handlers

Step-by-Step Debugging Workflow

Enable Fine-Grained Logging

Start by configuring the Python logging module to capture timestamped events from each handler. The VAD handler already emits debug statements that indicate state changes and processing times.

import logging

logging.basicConfig(level=logging.DEBUG)

Look for log lines prefixed with VAD: to observe chunk processing rates and identify where the pipeline stalls.

Measure Per-Stage Timing with Overrides

Create timing wrappers around each handler’s process method to pinpoint exactly which stage dominates your latency budget. This pattern works for any handler in the pipeline.

from time import perf_counter
import logging
from speech_to_speech.VAD.vad_handler import VADHandler

class TimedVADHandler(VADHandler):
    def process(self, audio_chunk):
        start = perf_counter()
        for out in super().process(audio_chunk):
            elapsed = perf_counter() - start
            logging.debug(f"VAD stage took {elapsed:.3f}s")
            start = perf_counter()
            yield out

Apply the same instrumentation to STT, LLM, and TTS handlers to generate a per-turn latency breakdown.

Profile the VAD Progressive Pause

The VAD handler dynamically calculates pauses based on speech segment length through _progressive_processing_pause. If you observe delays between speech completion and processing initiation, modify the pause calculation logic.


# Force a constant 200ms pause instead of progressive calculation

TimedVADHandler._progressive_processing_pause = lambda self, _: 0.2

Alternatively, adjust the realtime_processing_pause parameter (default 0.5s) in your handler initialization to reduce wait times.

Check Audio Enhancement Overhead

When audio_enhancement=True, the VAD handler invokes DeepFilterNet via _apply_audio_enhancement, which adds resampling and model inference overhead. Temporarily disable this to isolate its impact.

vad_handler = VADHandler()
vad_handler.setup(..., audio_enhancement=False)

Compare timing logs with enhancement enabled versus disabled to quantify the exact millisecond cost.

Verify STT Model Warm-up and Compilation

The WhisperSTTHandler performs warm-up runs during setup, and torch compilation incurs CUDA graph capture costs. If latency spikes occur only on the first utterance, compilation is the culprit.

from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler

stt = WhisperSTTHandler()
stt.setup(model_name="distil-whisper/distil-large-v3", compile_mode=None)

Setting compile_mode=None skips graph compilation, trading warmup time for slightly slower per-inference performance. Check WhisperSTTHandler.warmup() for the dummy generation logic that prepares the model.

Inspect Generation Kwargs

Large max_new_tokens values or aggressive sampling parameters increase inference time in both STT and LLM stages. Constrain these during profiling.

from speech_to_speech.LLM.base_openai_compatible_language_model import BaseOpenAICompatibleLanguageModel

llm = BaseOpenAICompatibleLanguageModel()
llm.setup(gen_kwargs={"max_new_tokens": 64, "temperature": 0.7})

Monitor GPU Utilization

Low GPU utilization during STT or LLM steps indicates CPU-side bottlenecks, typically in audio feature extraction or tokenization. Run nvidia-smi or use PyTorch diagnostics while processing audio.

import torch

# Check allocated memory to confirm GPU engagement

print(f"GPU memory allocated: {torch.cuda.memory_allocated() / 1e6:.2f} MB")

Review Turn-Handling Logic

The SpeculativeTurnTracker may delay turn finalization based on speculative_reopen_ms and unanswered_reopen_ms parameters. If latency occurs between speech end and LLM invocation, reduce these time windows in your pipeline configuration.

End-to-End Timing

Add high-level instrumentation around the complete pipeline run in s2s_pipeline.py to measure total latency from audio capture to synthesized output.

from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
import time
import logging

pipeline = SpeechToSpeechPipeline(
    vad_handler=debug_vad,
    stt_handler=stt,
    llm_handler=llm,
    tts_handler=your_tts_handler,
)

start = time.time()
pipeline.run()
logging.info(f"Total pipeline latency: {time.time() - start:.3f}s")

Practical Debugging Implementation

Combine the above techniques into a single debugging configuration:

import logging
from time import perf_counter
from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.STT.whisper_stt_handler import WhisperSTTHandler
from speech_to_speech.LLM.base_openai_compatible_language_model import BaseOpenAICompatibleLanguageModel

# Enable verbose logging

logging.basicConfig(level=logging.DEBUG)
logger = logging.getLogger(__name__)

# 1. Debug VAD with timing and reduced pause

class DebugVADHandler(VADHandler):
    def process(self, audio_chunk):
        start = perf_counter()
        for out in super().process(audio_chunk):
            logger.debug(f"VAD chunk latency: {perf_counter() - start:.4f}s")
            start = perf_counter()
            yield out

# Override progressive pause to constant 200ms

DebugVADHandler._progressive_processing_pause = lambda self, _: 0.2

vad = DebugVADHandler()
vad.setup(..., audio_enhancement=False)

# 2. Skip compilation in STT to eliminate first-run latency

stt = WhisperSTTHandler()
stt.setup(model_name="distil-whisper/distil-large-v3", compile_mode=None)

# 3. Constrain LLM generation

llm = BaseOpenAICompatibleLanguageModel()
llm.setup(gen_kwargs={"max_new_tokens": 64, "temperature": 0.7})

# 4. Instantiate pipeline and measure end-to-end time

from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline

pipeline = SpeechToSpeechPipeline(
    vad_handler=vad,
    stt_handler=stt,
    llm_handler=llm,
    tts_handler=...,
)

t0 = time.time()
pipeline.run()
logger.info(f"Pipeline completed in {time.time() - t0:.3f}s")

Summary

  • Instrument each stage by subclassing handlers and wrapping process() methods with perf_counter() to identify which component dominates latency.
  • Disable optional processing such as DeepFilterNet audio enhancement and torch compilation to isolate their specific overhead costs.
  • Adjust VAD timing by overriding _progressive_processing_pause or reducing realtime_processing_pause from its default 0.5s.
  • Monitor resource utilization using GPU diagnostics to distinguish between model inference delays and CPU preprocessing bottlenecks.
  • Iterate with controlled tests using the end-to-end timer in s2s_pipeline.py to validate improvements after each configuration change.

Frequently Asked Questions

How do I identify which stage is causing the most delay?

Subclass each handler (VADHandler, WhisperSTTHandler, etc.) and override the process method to log perf_counter() deltas before yielding results. Compare the logged timings across stages to identify the bottleneck. The VAD stage often hides latency in _progressive_processing_pause calculations, while STT delays typically occur during the first inference due to torch compilation.

Why does the first utterance take longer than subsequent ones?

The STT handler performs model warm-up and optional torch compilation during initialization. According to the source code in whisper_stt_handler.py, the first real generation incurs CUDA graph capture costs when torch.compile is enabled. Set compile_mode=None during setup to eliminate this initial spike, or accept the one-time cost for better sustained throughput.

Can I reduce latency without changing the underlying models?

Yes. Disable audio enhancement in the VAD handler by setting audio_enhancement=False to remove DeepFilterNet processing overhead. Reduce realtime_processing_pause below the default 0.5s, and constrain max_new_tokens in both STT and LLM handlers. These configuration changes reduce processing time without requiring model swaps or quantization.

Where should I add logging to debug turn-handling delays?

Add instrumentation around the speculative turn tracker in the pipeline orchestration code. If you observe delays between speech completion and the LLM request, examine the speculative_reopen_ms and unanswered_reopen_ms parameters. These values control how long the system waits to confirm a turn boundary before forwarding text to the language model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →