How to Evaluate the Performance of a Speech-to-Speech Model: Benchmarking Guide
Evaluate speech-to-speech model performance by measuring warm-up time, inference latency, time-to-first-chunk/token, and real-time factor using the dedicated benchmark scripts in the huggingface/speech-to-speech repository.
The huggingface/speech-to-speech repository implements a modular Speech-to-Speech (S2S) pipeline composed of four interchangeable components: VAD → STT → LLM → TTS. To evaluate the performance of a speech-to-speech model, the repository provides specialized benchmark scripts that quantify latency, throughput, and real-time capability for the TTS and STT stages.
Understanding the Evaluation Metrics
The benchmark scripts capture five critical performance indicators for each component. These metrics determine whether a voice agent can deliver real-time conversational experiences.
TTS Performance Metrics
In scripts/benchmark_tts.py, the following metrics are collected per handler:
- Warm-up time: Seconds required to load the model and allocate buffers before the first synthesis. Measured immediately after handler construction using
result.warmup_time = time.perf_counter() - start_setup(lines 70-71). - Inference time: End-to-end latency for a single synthesis request, calculated as
time_taken = end_time - start_timewithin the benchmark loop (lines 92-94). - Time-to-first-chunk (TTFC): Latency until the first audio chunk is produced, critical for real-time user experience. Captured when the first chunk arrives via
time_to_first_chunk = time.perf_counter() - start_time(lines 82-84). - Audio duration: Length of the generated waveform in seconds, computed as
audio_duration = total_samples / DEFAULT_SAMPLE_RATE(lines 94-95). - Real-Time Factor (RTF): Ratio of
audio_duration / inference_time. An RTF ≤ 1 indicates faster-than-real-time synthesis, computed inBenchmarkResult.get_stats()(lines 64-66).
STT Performance Metrics
The scripts/benchmark_stt.py script tracks analogous metrics for speech recognition:
- Warm-up time: Initialization duration before the first transcription, measured using the same pattern as TTS during handler creation.
- Inference time: Time required to transcribe a full utterance, recorded per iteration as
time_taken = end_time - start_time. - Time-to-first-token (TTFT): Latency until the first token of the transcript is emitted, stored in
self.time_to_first_tokenwhen the handler yields its first result. - Sample transcription: One example output saved as
sample_transcriptionin the stats dictionary to sanity-check correctness.
The Benchmarking Methodology
Both benchmark_tts.py and benchmark_stt.py follow an identical five-step structure:
- Setup: Create a
stop_event, input/output queues, and instantiate the handler (handler = …Handler(...)). - Warm-up: Measure initialization time captured as
warmup_time. - Benchmark loop: Execute the handler for a specified number of
iterations, measuring per-iteration latency and auxiliary data (audio length, TTFC/TTFT). - Result aggregation: Compute mean, min, max, standard deviation, and RTF/TTFC statistics via
BenchmarkResult.get_stats(). - Reporting: Pretty-print a summary table and optionally export JSON via
save_results.
Running the TTS Benchmark
Benchmark TTS handlers to compare synthesis speed across different models and quantization levels.
Basic Usage
Benchmark the default handlers (kokoro, qwen3, pocket_tts) using a test prompt:
python scripts/benchmark_tts.py \
--text "Hello, this is a latency test." \
--iterations 5 \
--output tts_perf.json
Apple Silicon MLX Quantization
For Apple Silicon devices, benchmark Qwen3-TTS across multiple MLX quantization levels:
python scripts/benchmark_tts.py \
--handlers qwen3 \
--qwen3_mlx_quantizations bf16 4bit 6bit 8bit \
--iterations 3 \
--output qwen3_mlx_bench.json
Running the STT Benchmark
Evaluate speech recognition latency using different back-ends such as Whisper, Faster-Whisper, Lightning-Whisper-MLX, or Paraformer.
Multi-Handler Comparison
Benchmark Whisper (CUDA) and Parakeet-TDT on the same audio file:
python scripts/benchmark_stt.py \
--audio_file samples/sample.wav \
--handlers whisper parakeet-tdt \
--iterations 4 \
--output stt_perf.json
Apple Silicon Specific
Benchmark Lightning-Whisper-MLX optimized for Apple Silicon:
python scripts/benchmark_stt.py \
--audio_file samples/sample.wav \
--handlers whisper-mlx \
--iterations 3 \
--output mlx_whisper.json
Interpreting the Results
The benchmark scripts generate JSON output containing detailed statistics. Here is an excerpt from a TTS benchmark showing Qwen3 performance:
{
"results": [
{
"handler": "qwen3[bf16]",
"warmup_time": 2.14,
"avg_inference_time": 0.87,
"min_inference_time": 0.81,
"max_inference_time": 0.93,
"std_inference_time": 0.04,
"avg_audio_duration": 1.23,
"avg_rtf": 1.41,
"avg_time_to_first_chunk": 0.12,
"total_iterations": 3,
"errors": []
}
],
"timestamp": "2026-07-07 12:34:56"
}
Key fields map directly to the implementation: warmup_time reflects the initialization measured at lines 70-71 of benchmark_tts.py, while avg_rtf derives from the BenchmarkResult.get_stats() method at lines 64-66. An RTF value of 1.41 indicates the model takes 1.41 seconds to generate 1 second of audio—slower than real-time—whereas values below 1.0 indicate real-time capability.
Summary
- Measure component-specific metrics: Track warm-up time, inference latency, and time-to-first-output for both TTS (TTFC) and STT (TTFT) stages.
- Use provided benchmark scripts: Execute
scripts/benchmark_tts.pyandscripts/benchmark_stt.pyto evaluate any combination of handlers. - Monitor Real-Time Factor: Ensure RTF ≤ 1 for real-time conversational agents; higher values indicate latency that will degrade user experience.
- Compare quantizations: Test different model sizes and quantization levels (e.g., Qwen3 MLX 4-bit vs. 8-bit) to find the optimal speed-quality trade-off.
- Validate correctness: Check
sample_transcriptionin STT results to ensure speed optimizations do not compromise accuracy.
Frequently Asked Questions
What is the Real-Time Factor (RTF) and why does it matter?
The Real-Time Factor (RTF) is the ratio of audio duration to inference time. An RTF of 1.0 means the model generates audio exactly as fast as it plays; values below 1.0 indicate faster-than-real-time generation. For conversational speech-to-speech systems, maintaining RTF ≤ 1 is essential to prevent perceptible delays in responses.
How do I benchmark specific handlers rather than the defaults?
Pass the --handlers flag followed by the specific handler names you want to evaluate. For example, --handlers qwen3 will benchmark only the Qwen3-TTS handler, while --handlers whisper parakeet-tdt will compare those two STT back-ends. Refer to src/speech_to_speech/TTS/README.md and src/speech_to_speech/STT/README.md for the complete list of supported handlers.
What is the difference between TTFC and TTFT?
Time-to-first-chunk (TTFC) applies to TTS and measures the latency until the first audio byte is produced, crucial for perceived responsiveness in voice agents. Time-to-first-token (TTFT) applies to STT and measures the latency until the first transcription character is emitted. Both metrics determine how quickly a speech-to-speech system can begin responding to user input.
Can I evaluate the LLM component using these scripts?
The current benchmark scripts in scripts/benchmark_tts.py and scripts/benchmark_stt.py focus specifically on the TTS and STT components. The LLM evaluation is handled separately through the function calling utilities in src/speech_to_speech/LLM/tool_call/function_call.py, which validates parsing accuracy rather than inference latency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →