How to Benchmark and Profile the Different Pipeline Components for Optimization in Hugging Face Speech-to-Speech

The huggingface/speech-to-speech repository provides ready-to-use scripts (scripts/benchmark_tts.py and scripts/benchmark_stt.py) that measure latency, throughput, and Real-Time Factor (RTF) for each pipeline stage, and can be extended with cProfile and torch.profiler for deep performance analysis.

To systematically optimize a real-time speech-to-speech system, you must isolate and measure the cost of each modular component. The speech-to-speech repository implements a queue-based architecture where VAD, STT, LLM, and TTS handlers run as independent threads, allowing you to benchmark individual stages or the full pipeline using built-in tools and standard Python profilers.

Understanding the Modular Pipeline Architecture

The pipeline is assembled at runtime by build_pipeline() in speech_to_speech/s2s_pipeline.py, which creates typed queues (Queue[AudioInItem], Queue[STTOutItem], etc.) and events that connect independent handlers. This design allows each component to be instantiated and measured in isolation.

Component Responsibility Core Implementation
VAD Detects speech boundaries and streams partial audio speech_to_speech/VAD/vad_handler.py
STT Converts audio to text (Whisper, Faster-Whisper, MLX-Audio) speech_to_speech/s2s_pipeline.py → get_stt_handler()
LLM Generates responses via transformers, MLX-LM, or APIs speech_to_speech/s2s_pipeline.py → get_llm_handler()
TTS Synthesizes audio (Kokoro, Qwen3-TTS, Pocket-TTS, etc.) speech_to_speech/s2s_pipeline.py → get_tts_handler()

Because each handler runs in its own thread managed by the ThreadManager abstraction, you can benchmark only TTS to measure synthesis latency, or only STT to measure transcription overhead, without initializing the full pipeline.

Benchmarking Individual Components

The repository ships with two primary benchmarking scripts that follow a standardized pattern: warm-up, timed iteration loop, and statistical aggregation.

Text-to-Speech Benchmarking with benchmark_tts.py

Located at scripts/benchmark_tts.py, this script measures wall-clock time, warm-up latency, and Real-Time Factor (RTF)—calculated as audio_duration / inference_time—for any TTS handler.

python scripts/benchmark_tts.py \
    --text "Hello from the speech-to-speech benchmark. This is a latency test." \
    --handlers kokoro qwen3 pocket_tts \
    --iterations 5 \
    --output tts_bench.json

The script instantiates handlers like Qwen3TTSHandler, runs a warm-up phase to load weights and allocate tensors, then processes the input repeatedly while recording time.perf_counter() values. Results are saved as a BenchmarkResult object (defined at lines 34-44 of the script) containing mean, min, max, and standard deviation for inference time and Time-to-First-Chunk (TTFC).

Speech-to-Text Benchmarking with benchmark_stt.py

Located at scripts/benchmark_stt.py, this script loads audio via soundfile and optional librosa resampling (see the load_audio() function at lines 84-107), then benchmarks STT handlers such as Whisper, Faster-Whisper, or MLX-Audio-Whisper.

python scripts/benchmark_stt.py \
    --audio_file samples/short.wav \
    --handlers whisper mlx-audio-whisper faster-whisper \
    --iterations 3 \
    --output stt_bench.json

Both scripts output a comparison table ranking handlers by average inference time, making it easy to identify the fastest configuration for your specific hardware.

Deep Profiling for Optimization

The benchmark scripts provide high-level latency numbers, but optimization requires identifying specific bottlenecks in Python code, GPU kernels, or memory allocation. You can wrap the existing benchmark loops with standard profilers without modifying the handler implementations.

CPU Profiling with cProfile

To diagnose queue overhead or transcription post-processing bottlenecks, wrap the benchmark_handler function with cProfile:

import cProfile, pstats, io
from scripts.benchmark_tts import benchmark_handler

def run_profiled(handler_name, text, iterations):
    pr = cProfile.Profile()
    pr.enable()
    result = benchmark_handler(handler_name, text, iterations)
    pr.disable()
    
    s = io.StringIO()
    ps = pstats.Stats(pr, stream=s).sort_stats('cumtime')
    ps.print_stats(20)
    print(f"Profiling report for {handler_name}:\n", s.getvalue())
    return result

This prints the top 20 functions by cumulative time, highlighting costly operations like torch.nn.functional.linear or numpy.mean that may not be obvious from end-to-end latency alone.

GPU and MLX Profiling with torch.profiler

For CUDA or Apple Silicon optimization, use torch.profiler to capture kernel launch overhead and memory bottlenecks:

import torch.profiler as profiler
from scripts.benchmark_tts import benchmark_handler

def run_gpu_profiled(handler_name, text, iterations):
    with profiler.profile(
        schedule=profiler.schedule(wait=1, warmup=1, active=3, repeat=1),
        on_trace_ready=profiler.tensorboard_trace_handler("./tb_logs"),
        record_shapes=True,
        profile_memory=True,
        with_stack=True,
    ) as prof:
        result = benchmark_handler(handler_name, text, iterations)
        prof.step()
    
    print(f"GPU profile for {handler_name} stored in ./tb_logs")
    return result

Open the trace in TensorBoard with tensorboard --logdir ./tb_logs to visualize kernel-level timings. This works for both CUDA and MLX backends, as the repository uses a torch shim for MLX compatibility.

Practical Workflow for Optimization

Follow this systematic approach to benchmark and profile the different pipeline components for optimization:

  1. Establish baseline measurements – Run the vanilla benchmark_tts.py or benchmark_stt.py on your target hardware (Apple Silicon, RTX 3090, or CPU-only).
  2. Identify the bottleneck – Use the comparison table to determine which stage (VAD, STT, LLM, or TTS) contributes most to end-to-end latency.
  3. Instrument with profilers – Wrap the heavy handler with cProfile (for Python overhead) or torch.profiler (for GPU/MLX kernels).
  4. Iterate on configuration – Test device arguments (device="cuda", device="mps", device="cpu"), quantization levels (e.g., normalize_qwen3_mlx_quantizations() in benchmark_tts.py supports 4-bit and 8-bit for Qwen3-TTS), and live transcription toggles.
  5. Version-control results – The scripts automatically write JSON summaries (tts_benchmark_results.json). Compare these files across commits to catch performance regressions.
  6. Apply hardware-specific optimizations – Enable local_mac_optimal_settings in s2s_pipeline.py (see optimal_mac_settings() around line 31) or adjust the realtime pool size with --num_pipelines.

Complete Code Examples

End-to-End Pipeline Benchmark

To measure the full VAD→STT→LLM→TTS chain, use the argument parsing logic from s2s_pipeline.py:

python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --num_pipelines 2 \
    --stt whisper \
    --tts qwen3 \
    --llm_backend responses-api \
    --iterations 5 \
    --output full_pipeline_bench.json

This builds a realtime pool with two parallel pipelines and outputs the same statistical format as the individual benchmark scripts.

Profiling Qwen3-TTS with PyTorch Profiler

Create a dedicated profiling script to test quantization effects on Apple Silicon:


# profile_qwen3_tts.py

from scripts.benchmark_tts import benchmark_handler, normalize_qwen3_mlx_quantizations
import torch.profiler as profiler

handler_name = "qwen3"
text = "Benchmarking Qwen3-TTS on Apple Silicon"
iterations = 3

quantizations = normalize_qwen3_mlx_quantizations(["4bit"])
handler_kwargs = {"mlx_quantization": quantizations[0]}

with profiler.profile(
    schedule=profiler.schedule(wait=1, warmup=1, active=2),
    on_trace_ready=profiler.tensorboard_trace_handler("./tensorboard_qwen3"),
    record_shapes=True,
    profile_memory=True,
):
    result = benchmark_handler(
        handler_name, text, iterations, handler_kwargs=handler_kwargs
    )

print("Result:", result.get_stats())

Run with python profile_qwen3_tts.py and visualize with tensorboard --logdir ./tensorboard_qwen3.

Diagnosing STT Queue Overhead with cProfile

To isolate queue handling costs from model inference in speech_to_speech/STT/whisper_stt_handler.py:

import cProfile, pstats, io
from scripts.benchmark_stt import benchmark_handler

def profile_stt(handler_name, audio_path, iterations):
    pr = cProfile.Profile()
    pr.enable()
    result = benchmark_handler(handler_name, audio_path, iterations)
    pr.disable()
    
    s = io.StringIO()
    ps = pstats.Stats(pr, stream=s).sort_stats('cumtime')
    ps.print_stats(15)
    print(f"cProfile for {handler_name}:\n", s.getvalue())
    return result

Replace the direct call in benchmark_stt.py with profile_stt(...) to see exactly how much time is spent in queue operations versus the Whisper model forward pass.

Summary

  • Use the built-in scripts (scripts/benchmark_tts.py and scripts/benchmark_stt.py) to generate baseline latency, RTF, and TTFC metrics for individual handlers.
  • Leverage modularity – The queue-based architecture in speech_to_speech/s2s_pipeline.py allows you to benchmark get_tts_handler() or get_stt_handler() in isolation without initializing the full pipeline.
  • Add granular profiling – Wrap benchmark calls with cProfile for Python-level bottlenecks or torch.profiler for GPU/MLX kernel analysis.
  • Optimize via configuration – Test device arguments, mlx_quantization levels (via normalize_qwen3_mlx_quantizations()), and num_pipelines to find the optimal throughput for your hardware.
  • Monitor regressions – Commit the JSON output files (tts_benchmark_results.json) to track performance changes across code updates.

Frequently Asked Questions

How do I measure the Real-Time Factor (RTF) for TTS handlers?

The benchmark_tts.py script calculates RTF automatically by dividing the audio duration by the inference time. After running the benchmark with --iterations, check the output table or JSON file for the RTF column. Values below 1.0 indicate real-time capability (faster than playback), while values above 1.0 indicate the synthesis is slower than real-time.

Can I benchmark individual pipeline stages without running the full speech-to-speech system?

Yes. Because the pipeline uses a modular architecture with independent handlers connected via queues, you can import and benchmark individual components directly. Use get_stt_handler() or get_tts_handler() from speech_to_speech/s2s_pipeline.py to instantiate a single handler, then pass it to the benchmark functions in scripts/benchmark_tts.py or scripts/benchmark_stt.py without building the full pipeline.

What is the difference between using cProfile and torch.profiler for optimization?

Use cProfile when you need to identify slow Python functions, data preprocessing, or queue management overhead in the CPU-bound sections of the code. Use torch.profiler when optimizing GPU or Apple MLX execution, as it captures kernel launch times, memory bandwidth bottlenecks, and tensor shapes. For comprehensive optimization, run cProfile first to find the slow stage, then use torch.profiler on that specific handler to analyze its compute kernels.

How do I optimize for Apple Silicon specifically?

Enable Apple-specific optimizations by using the mlx-audio-whisper handler for STT and qwen3 handler with MLX quantization for TTS. In benchmark_tts.py, pass quantization levels through normalize_qwen3_mlx_quantizations() (e.g., ["4bit"] or ["8bit"]). Additionally, set local_mac_optimal_settings=True when building the pipeline in s2s_pipeline.py, which configures optimal thread counts and memory settings for M1/M2/M3 chips.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →