How to Evaluate the Performance of a Speech-to-Speech Model: Benchmarking Guide

Evaluate speech-to-speech model performance by measuring warm-up time, inference latency, time-to-first-chunk/token, and real-time factor using the dedicated benchmark scripts in the huggingface/speech-to-speech repository.

The huggingface/speech-to-speech repository implements a modular Speech-to-Speech (S2S) pipeline composed of four interchangeable components: VAD → STT → LLM → TTS. To evaluate the performance of a speech-to-speech model, the repository provides specialized benchmark scripts that quantify latency, throughput, and real-time capability for the TTS and STT stages.

Understanding the Evaluation Metrics

The benchmark scripts capture five critical performance indicators for each component. These metrics determine whether a voice agent can deliver real-time conversational experiences.

TTS Performance Metrics

In scripts/benchmark_tts.py, the following metrics are collected per handler:

  • Warm-up time: Seconds required to load the model and allocate buffers before the first synthesis. Measured immediately after handler construction using result.warmup_time = time.perf_counter() - start_setup (lines 70-71).
  • Inference time: End-to-end latency for a single synthesis request, calculated as time_taken = end_time - start_time within the benchmark loop (lines 92-94).
  • Time-to-first-chunk (TTFC): Latency until the first audio chunk is produced, critical for real-time user experience. Captured when the first chunk arrives via time_to_first_chunk = time.perf_counter() - start_time (lines 82-84).
  • Audio duration: Length of the generated waveform in seconds, computed as audio_duration = total_samples / DEFAULT_SAMPLE_RATE (lines 94-95).
  • Real-Time Factor (RTF): Ratio of audio_duration / inference_time. An RTF ≤ 1 indicates faster-than-real-time synthesis, computed in BenchmarkResult.get_stats() (lines 64-66).

STT Performance Metrics

The scripts/benchmark_stt.py script tracks analogous metrics for speech recognition:

  • Warm-up time: Initialization duration before the first transcription, measured using the same pattern as TTS during handler creation.
  • Inference time: Time required to transcribe a full utterance, recorded per iteration as time_taken = end_time - start_time.
  • Time-to-first-token (TTFT): Latency until the first token of the transcript is emitted, stored in self.time_to_first_token when the handler yields its first result.
  • Sample transcription: One example output saved as sample_transcription in the stats dictionary to sanity-check correctness.

The Benchmarking Methodology

Both benchmark_tts.py and benchmark_stt.py follow an identical five-step structure:

  1. Setup: Create a stop_event, input/output queues, and instantiate the handler (handler = …Handler(...)).
  2. Warm-up: Measure initialization time captured as warmup_time.
  3. Benchmark loop: Execute the handler for a specified number of iterations, measuring per-iteration latency and auxiliary data (audio length, TTFC/TTFT).
  4. Result aggregation: Compute mean, min, max, standard deviation, and RTF/TTFC statistics via BenchmarkResult.get_stats().
  5. Reporting: Pretty-print a summary table and optionally export JSON via save_results.

Running the TTS Benchmark

Benchmark TTS handlers to compare synthesis speed across different models and quantization levels.

Basic Usage

Benchmark the default handlers (kokoro, qwen3, pocket_tts) using a test prompt:

python scripts/benchmark_tts.py \
    --text "Hello, this is a latency test." \
    --iterations 5 \
    --output tts_perf.json

Apple Silicon MLX Quantization

For Apple Silicon devices, benchmark Qwen3-TTS across multiple MLX quantization levels:

python scripts/benchmark_tts.py \
    --handlers qwen3 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit \
    --iterations 3 \
    --output qwen3_mlx_bench.json

Running the STT Benchmark

Evaluate speech recognition latency using different back-ends such as Whisper, Faster-Whisper, Lightning-Whisper-MLX, or Paraformer.

Multi-Handler Comparison

Benchmark Whisper (CUDA) and Parakeet-TDT on the same audio file:

python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper parakeet-tdt \
    --iterations 4 \
    --output stt_perf.json

Apple Silicon Specific

Benchmark Lightning-Whisper-MLX optimized for Apple Silicon:

python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper-mlx \
    --iterations 3 \
    --output mlx_whisper.json

Interpreting the Results

The benchmark scripts generate JSON output containing detailed statistics. Here is an excerpt from a TTS benchmark showing Qwen3 performance:

{
  "results": [
    {
      "handler": "qwen3[bf16]",
      "warmup_time": 2.14,
      "avg_inference_time": 0.87,
      "min_inference_time": 0.81,
      "max_inference_time": 0.93,
      "std_inference_time": 0.04,
      "avg_audio_duration": 1.23,
      "avg_rtf": 1.41,
      "avg_time_to_first_chunk": 0.12,
      "total_iterations": 3,
      "errors": []
    }
  ],
  "timestamp": "2026-07-07 12:34:56"
}

Key fields map directly to the implementation: warmup_time reflects the initialization measured at lines 70-71 of benchmark_tts.py, while avg_rtf derives from the BenchmarkResult.get_stats() method at lines 64-66. An RTF value of 1.41 indicates the model takes 1.41 seconds to generate 1 second of audio—slower than real-time—whereas values below 1.0 indicate real-time capability.

Summary

  • Measure component-specific metrics: Track warm-up time, inference latency, and time-to-first-output for both TTS (TTFC) and STT (TTFT) stages.
  • Use provided benchmark scripts: Execute scripts/benchmark_tts.py and scripts/benchmark_stt.py to evaluate any combination of handlers.
  • Monitor Real-Time Factor: Ensure RTF ≤ 1 for real-time conversational agents; higher values indicate latency that will degrade user experience.
  • Compare quantizations: Test different model sizes and quantization levels (e.g., Qwen3 MLX 4-bit vs. 8-bit) to find the optimal speed-quality trade-off.
  • Validate correctness: Check sample_transcription in STT results to ensure speed optimizations do not compromise accuracy.

Frequently Asked Questions

What is the Real-Time Factor (RTF) and why does it matter?

The Real-Time Factor (RTF) is the ratio of audio duration to inference time. An RTF of 1.0 means the model generates audio exactly as fast as it plays; values below 1.0 indicate faster-than-real-time generation. For conversational speech-to-speech systems, maintaining RTF ≤ 1 is essential to prevent perceptible delays in responses.

How do I benchmark specific handlers rather than the defaults?

Pass the --handlers flag followed by the specific handler names you want to evaluate. For example, --handlers qwen3 will benchmark only the Qwen3-TTS handler, while --handlers whisper parakeet-tdt will compare those two STT back-ends. Refer to src/speech_to_speech/TTS/README.md and src/speech_to_speech/STT/README.md for the complete list of supported handlers.

What is the difference between TTFC and TTFT?

Time-to-first-chunk (TTFC) applies to TTS and measures the latency until the first audio byte is produced, crucial for perceived responsiveness in voice agents. Time-to-first-token (TTFT) applies to STT and measures the latency until the first transcription character is emitted. Both metrics determine how quickly a speech-to-speech system can begin responding to user input.

Can I evaluate the LLM component using these scripts?

The current benchmark scripts in scripts/benchmark_tts.py and scripts/benchmark_stt.py focus specifically on the TTS and STT components. The LLM evaluation is handled separately through the function calling utilities in src/speech_to_speech/LLM/tool_call/function_call.py, which validates parsing accuracy rather than inference latency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →