How to Benchmark and Profile the Different Pipeline Components for Optimization in Hugging Face Speech-to-Speech
The huggingface/speech-to-speech repository provides ready-to-use scripts (scripts/benchmark_tts.py and scripts/benchmark_stt.py) that measure latency, throughput, and Real-Time Factor (RTF) for each pipeline stage, and can be extended with cProfile and torch.profiler for deep performance analysis.
To systematically optimize a real-time speech-to-speech system, you must isolate and measure the cost of each modular component. The speech-to-speech repository implements a queue-based architecture where VAD, STT, LLM, and TTS handlers run as independent threads, allowing you to benchmark individual stages or the full pipeline using built-in tools and standard Python profilers.
Understanding the Modular Pipeline Architecture
The pipeline is assembled at runtime by build_pipeline() in speech_to_speech/s2s_pipeline.py, which creates typed queues (Queue[AudioInItem], Queue[STTOutItem], etc.) and events that connect independent handlers. This design allows each component to be instantiated and measured in isolation.
| Component | Responsibility | Core Implementation |
|---|---|---|
| VAD | Detects speech boundaries and streams partial audio | speech_to_speech/VAD/vad_handler.py |
| STT | Converts audio to text (Whisper, Faster-Whisper, MLX-Audio) | speech_to_speech/s2s_pipeline.py → get_stt_handler() |
| LLM | Generates responses via transformers, MLX-LM, or APIs | speech_to_speech/s2s_pipeline.py → get_llm_handler() |
| TTS | Synthesizes audio (Kokoro, Qwen3-TTS, Pocket-TTS, etc.) | speech_to_speech/s2s_pipeline.py → get_tts_handler() |
Because each handler runs in its own thread managed by the ThreadManager abstraction, you can benchmark only TTS to measure synthesis latency, or only STT to measure transcription overhead, without initializing the full pipeline.
Benchmarking Individual Components
The repository ships with two primary benchmarking scripts that follow a standardized pattern: warm-up, timed iteration loop, and statistical aggregation.
Text-to-Speech Benchmarking with benchmark_tts.py
Located at scripts/benchmark_tts.py, this script measures wall-clock time, warm-up latency, and Real-Time Factor (RTF)—calculated as audio_duration / inference_time—for any TTS handler.
python scripts/benchmark_tts.py \
--text "Hello from the speech-to-speech benchmark. This is a latency test." \
--handlers kokoro qwen3 pocket_tts \
--iterations 5 \
--output tts_bench.json
The script instantiates handlers like Qwen3TTSHandler, runs a warm-up phase to load weights and allocate tensors, then processes the input repeatedly while recording time.perf_counter() values. Results are saved as a BenchmarkResult object (defined at lines 34-44 of the script) containing mean, min, max, and standard deviation for inference time and Time-to-First-Chunk (TTFC).
Speech-to-Text Benchmarking with benchmark_stt.py
Located at scripts/benchmark_stt.py, this script loads audio via soundfile and optional librosa resampling (see the load_audio() function at lines 84-107), then benchmarks STT handlers such as Whisper, Faster-Whisper, or MLX-Audio-Whisper.
python scripts/benchmark_stt.py \
--audio_file samples/short.wav \
--handlers whisper mlx-audio-whisper faster-whisper \
--iterations 3 \
--output stt_bench.json
Both scripts output a comparison table ranking handlers by average inference time, making it easy to identify the fastest configuration for your specific hardware.
Deep Profiling for Optimization
The benchmark scripts provide high-level latency numbers, but optimization requires identifying specific bottlenecks in Python code, GPU kernels, or memory allocation. You can wrap the existing benchmark loops with standard profilers without modifying the handler implementations.
CPU Profiling with cProfile
To diagnose queue overhead or transcription post-processing bottlenecks, wrap the benchmark_handler function with cProfile:
import cProfile, pstats, io
from scripts.benchmark_tts import benchmark_handler
def run_profiled(handler_name, text, iterations):
pr = cProfile.Profile()
pr.enable()
result = benchmark_handler(handler_name, text, iterations)
pr.disable()
s = io.StringIO()
ps = pstats.Stats(pr, stream=s).sort_stats('cumtime')
ps.print_stats(20)
print(f"Profiling report for {handler_name}:\n", s.getvalue())
return result
This prints the top 20 functions by cumulative time, highlighting costly operations like torch.nn.functional.linear or numpy.mean that may not be obvious from end-to-end latency alone.
GPU and MLX Profiling with torch.profiler
For CUDA or Apple Silicon optimization, use torch.profiler to capture kernel launch overhead and memory bottlenecks:
import torch.profiler as profiler
from scripts.benchmark_tts import benchmark_handler
def run_gpu_profiled(handler_name, text, iterations):
with profiler.profile(
schedule=profiler.schedule(wait=1, warmup=1, active=3, repeat=1),
on_trace_ready=profiler.tensorboard_trace_handler("./tb_logs"),
record_shapes=True,
profile_memory=True,
with_stack=True,
) as prof:
result = benchmark_handler(handler_name, text, iterations)
prof.step()
print(f"GPU profile for {handler_name} stored in ./tb_logs")
return result
Open the trace in TensorBoard with tensorboard --logdir ./tb_logs to visualize kernel-level timings. This works for both CUDA and MLX backends, as the repository uses a torch shim for MLX compatibility.
Practical Workflow for Optimization
Follow this systematic approach to benchmark and profile the different pipeline components for optimization:
- Establish baseline measurements – Run the vanilla
benchmark_tts.pyorbenchmark_stt.pyon your target hardware (Apple Silicon, RTX 3090, or CPU-only). - Identify the bottleneck – Use the comparison table to determine which stage (VAD, STT, LLM, or TTS) contributes most to end-to-end latency.
- Instrument with profilers – Wrap the heavy handler with
cProfile(for Python overhead) ortorch.profiler(for GPU/MLX kernels). - Iterate on configuration – Test device arguments (
device="cuda",device="mps",device="cpu"), quantization levels (e.g.,normalize_qwen3_mlx_quantizations()inbenchmark_tts.pysupports 4-bit and 8-bit for Qwen3-TTS), and live transcription toggles. - Version-control results – The scripts automatically write JSON summaries (
tts_benchmark_results.json). Compare these files across commits to catch performance regressions. - Apply hardware-specific optimizations – Enable
local_mac_optimal_settingsins2s_pipeline.py(seeoptimal_mac_settings()around line 31) or adjust the realtime pool size with--num_pipelines.
Complete Code Examples
End-to-End Pipeline Benchmark
To measure the full VAD→STT→LLM→TTS chain, use the argument parsing logic from s2s_pipeline.py:
python -m speech_to_speech.s2s_pipeline \
--mode realtime \
--num_pipelines 2 \
--stt whisper \
--tts qwen3 \
--llm_backend responses-api \
--iterations 5 \
--output full_pipeline_bench.json
This builds a realtime pool with two parallel pipelines and outputs the same statistical format as the individual benchmark scripts.
Profiling Qwen3-TTS with PyTorch Profiler
Create a dedicated profiling script to test quantization effects on Apple Silicon:
# profile_qwen3_tts.py
from scripts.benchmark_tts import benchmark_handler, normalize_qwen3_mlx_quantizations
import torch.profiler as profiler
handler_name = "qwen3"
text = "Benchmarking Qwen3-TTS on Apple Silicon"
iterations = 3
quantizations = normalize_qwen3_mlx_quantizations(["4bit"])
handler_kwargs = {"mlx_quantization": quantizations[0]}
with profiler.profile(
schedule=profiler.schedule(wait=1, warmup=1, active=2),
on_trace_ready=profiler.tensorboard_trace_handler("./tensorboard_qwen3"),
record_shapes=True,
profile_memory=True,
):
result = benchmark_handler(
handler_name, text, iterations, handler_kwargs=handler_kwargs
)
print("Result:", result.get_stats())
Run with python profile_qwen3_tts.py and visualize with tensorboard --logdir ./tensorboard_qwen3.
Diagnosing STT Queue Overhead with cProfile
To isolate queue handling costs from model inference in speech_to_speech/STT/whisper_stt_handler.py:
import cProfile, pstats, io
from scripts.benchmark_stt import benchmark_handler
def profile_stt(handler_name, audio_path, iterations):
pr = cProfile.Profile()
pr.enable()
result = benchmark_handler(handler_name, audio_path, iterations)
pr.disable()
s = io.StringIO()
ps = pstats.Stats(pr, stream=s).sort_stats('cumtime')
ps.print_stats(15)
print(f"cProfile for {handler_name}:\n", s.getvalue())
return result
Replace the direct call in benchmark_stt.py with profile_stt(...) to see exactly how much time is spent in queue operations versus the Whisper model forward pass.
Summary
- Use the built-in scripts (
scripts/benchmark_tts.pyandscripts/benchmark_stt.py) to generate baseline latency, RTF, and TTFC metrics for individual handlers. - Leverage modularity – The queue-based architecture in
speech_to_speech/s2s_pipeline.pyallows you to benchmarkget_tts_handler()orget_stt_handler()in isolation without initializing the full pipeline. - Add granular profiling – Wrap benchmark calls with
cProfilefor Python-level bottlenecks ortorch.profilerfor GPU/MLX kernel analysis. - Optimize via configuration – Test
devicearguments,mlx_quantizationlevels (vianormalize_qwen3_mlx_quantizations()), andnum_pipelinesto find the optimal throughput for your hardware. - Monitor regressions – Commit the JSON output files (
tts_benchmark_results.json) to track performance changes across code updates.
Frequently Asked Questions
How do I measure the Real-Time Factor (RTF) for TTS handlers?
The benchmark_tts.py script calculates RTF automatically by dividing the audio duration by the inference time. After running the benchmark with --iterations, check the output table or JSON file for the RTF column. Values below 1.0 indicate real-time capability (faster than playback), while values above 1.0 indicate the synthesis is slower than real-time.
Can I benchmark individual pipeline stages without running the full speech-to-speech system?
Yes. Because the pipeline uses a modular architecture with independent handlers connected via queues, you can import and benchmark individual components directly. Use get_stt_handler() or get_tts_handler() from speech_to_speech/s2s_pipeline.py to instantiate a single handler, then pass it to the benchmark functions in scripts/benchmark_tts.py or scripts/benchmark_stt.py without building the full pipeline.
What is the difference between using cProfile and torch.profiler for optimization?
Use cProfile when you need to identify slow Python functions, data preprocessing, or queue management overhead in the CPU-bound sections of the code. Use torch.profiler when optimizing GPU or Apple MLX execution, as it captures kernel launch times, memory bandwidth bottlenecks, and tensor shapes. For comprehensive optimization, run cProfile first to find the slow stage, then use torch.profiler on that specific handler to analyze its compute kernels.
How do I optimize for Apple Silicon specifically?
Enable Apple-specific optimizations by using the mlx-audio-whisper handler for STT and qwen3 handler with MLX quantization for TTS. In benchmark_tts.py, pass quantization levels through normalize_qwen3_mlx_quantizations() (e.g., ["4bit"] or ["8bit"]). Additionally, set local_mac_optimal_settings=True when building the pipeline in s2s_pipeline.py, which configures optimal thread counts and memory settings for M1/M2/M3 chips.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →