How to Evaluate the Performance of a Speech-to-Speech Model
Benchmark speech-to-speech models using the huggingface/speech-to-speech repository's built-in scripts to measure latency, real-time factor, and time-to-first-chunk across TTS and STT components.
Evaluating speech-to-speech (S2S) systems requires more than accuracy metrics—you need precise latency measurements to ensure real-time voice-agent responsiveness. The huggingface/speech-to-speech repository provides dedicated benchmark scripts for measuring the performance of each pipeline stage.
Understanding S2S Pipeline Components
Speech-to-speech pipelines chain four interchangeable components: VAD → STT → LLM → TTS. Performance bottlenecks can occur at any stage, so the repository isolates the two most latency-sensitive modules—STT and TTS—for standalone benchmarking.
| Component | Critical Metrics |
|---|---|
| STT | Warm-up time, inference latency, time-to-first-token (TTFT) |
| TTS | Warm-up time, inference latency, time-to-first-chunk (TTFC), real-time factor (RTF) |
Benchmarking TTS Performance
The scripts/benchmark_tts.py script measures synthesis speed across any supported TTS handler: Kokoro-82M, Qwen3-TTS, Pocket-TTS, or others.
Key TTS Metrics Explained
- Warm-up time: Seconds to load the model and allocate buffers before first synthesis (
result.warmup_time = time.perf_counter() - start_setupinbenchmark_tts.pylines 70-71) - Inference time: End-to-end latency for a single request (
time_taken = end_time - start_time, lines 92-94) - Time-to-first-chunk (TTFC): Critical for UX—how long until audio starts streaming (
time_to_first_chunk = time.perf_counter() - start_time, lines 82-84) - Real-time factor (RTF):
audio_duration / inference_time—values ≤ 1 indicate faster-than-real-time synthesis (computed inBenchmarkResult.get_stats(), lines 64-66)
Running TTS Benchmarks
# Basic usage — benchmark default handlers (kokoro, qwen3, pocket_tts)
python scripts/benchmark_tts.py \
--text "Hello, this is a latency test." \
--iterations 5 \
--output tts_perf.json
# Compare Qwen3-TTS across all MLX quantizations (Apple Silicon)
python scripts/benchmark_tts.py \
--handlers qwen3 \
--qwen3_mlx_quantizations bf16 4bit 6bit 8bit \
--iterations 3 \
--output qwen3_mlx_bench.json
Benchmarking STT Performance
The scripts/benchmark_stt.py script evaluates speech recognition latency using Whisper, Faster-Whisper, Lightning-Whisper-MLX, Paraformer, or Parakeet-TDT backends.
Key STT Metrics
- Warm-up time: Model initialization before first transcription
- Inference time: Full utterance transcription latency
- Time-to-first-token (TTFT): Streaming responsiveness metric
- Sample transcription: Sanity-check output for correctness validation
Running STT Benchmarks
# Compare Whisper (CUDA) and Parakeet-TDT on identical audio
python scripts/benchmark_stt.py \
--audio_file samples/sample.wav \
--handlers whisper parakeet-tdt \
--iterations 4 \
--output stt_perf.json
# Benchmark Lightning-Whisper-MLX on Apple Silicon
python scripts/benchmark_stt.py \
--audio_file samples/sample.wav \
--handlers whisper-mlx \
--iterations 3 \
--output mlx_whisper.json
Interpreting Benchmark Results
Both scripts follow an identical execution pattern:
- Setup — create
stop_event, input/output queues, instantiate handler - Warm-up — measure initialization time
- Benchmark loop — run
iterationstimes, collecting per-call latency - Aggregation — compute mean, min, max, standard deviation, and derived metrics
- Reporting — print summary table and optionally save JSON
Sample JSON Output Structure
{
"results": [
{
"handler": "qwen3[bf16]",
"warmup_time": 2.14,
"avg_inference_time": 0.87,
"min_inference_time": 0.81,
"max_inference_time": 0.93,
"std_inference_time": 0.04,
"avg_audio_duration": 1.23,
"avg_rtf": 1.41,
"avg_time_to_first_chunk": 0.12,
"total_iterations": 3,
"errors": []
}
],
"timestamp": "2026-07-07 12:34:56"
}
How to Evaluate the Performance of a Speech-to-Speech Model in Practice
Complete evaluation requires testing both directions of the pipeline:
- Establish baseline — run TTS benchmark with
--iterations 10minimum for statistical stability - Measure STT separately — use representative audio samples matching your target domain
- Combine results — sum
avg_inference_timefrom both stages plus estimated LLM latency for total pipeline latency - Validate RTF < 1 — ensure synthesis keeps pace with real-time requirements
For handler-specific documentation, refer to:
src/speech_to_speech/TTS/README.md— TTS handler configurationssrc/speech_to_speech/STT/README.md— supported STT models and parameters
Summary
- Use
scripts/benchmark_tts.pyto measure synthesis latency, TTFC, and RTF across any TTS backend - Use
scripts/benchmark_stt.pyto evaluate recognition speed and TTFT for STT models - Target RTF ≤ 1 for real-time voice agents—higher values cause perceptible delays
- Run multiple iterations (5-10 minimum) to obtain stable statistics with standard deviation
- Compare quantizations using built-in flags like
--qwen3_mlx_quantizationsfor deployment optimization
Frequently Asked Questions
What is Real-Time Factor (RTF) in TTS evaluation?
RTF measures synthesis speed relative to audio duration. Calculated as audio_duration / inference_time in BenchmarkResult.get_stats(), an RTF of 1.0 means synthesis takes exactly as long as the utterance itself. Values below 1.0 indicate faster-than-real-time generation. For interactive voice agents, maintain RTF well under 1.0 to prevent cumulative latency.
How does time-to-first-chunk (TTFC) differ from inference time?
TTFC captures streaming latency; inference time measures total duration. TTFC represents when the first audio bytes are ready (time.perf_counter() - start_time at first chunk arrival), crucial for perceived responsiveness. Inference time spans the complete synthesis. A model with low inference time but high TTFC feels sluggish to users.
Can I benchmark custom TTS or STT handlers?
Yes—any handler conforming to the repository's interface is compatible. The benchmark scripts accept arbitrary handler names via --handlers and instantiate them through the standard factory pattern. Ensure your handler implements the required __call__ signature accepting (queue, stop_event) and yielding results progressively.
What hardware considerations affect benchmark validity?
Apple Silicon optimizations (MLX) and CUDA backends show significant divergence. The repository provides quantization-specific flags like --qwen3_mlx_quantizations because bf16, 4bit, 6bit, and 8bit variants exhibit different latency-accuracy tradeoffs. Always benchmark on target deployment hardware—results from CUDA systems do not transfer to MLX or CPU deployments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →