How to Evaluate the Performance of a Speech-to-Speech Model: A Complete Benchmarking Guide

You can evaluate speech-to-speech model performance using the dedicated benchmark scripts benchmark_tts.py and benchmark_stt.py in the Hugging Face speech-to-speech repository, which measure latency, real-time factor, and time-to-first-chunk/token across the VAD → STT → LLM → TTS pipeline components.

The huggingface/speech-to-speech repository implements a modular pipeline architecture where audio flows through four interchangeable stages: Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS). To evaluate the performance of a speech-to-speech model effectively, the codebase provides specialized benchmarking tools that quantify latency and throughput for the TTS and STT components specifically.

Understanding the Speech-to-Speech Pipeline Components

The repository structures speech-to-speech inference as a chain of handlers. When you evaluate the performance of a speech-to-speech model, you are essentially measuring how quickly each stage processes data and delivers results to the next component. The benchmark scripts focus on the two most latency-critical stages: STT (input processing) and TTS (output generation).

Key Performance Metrics for Speech-to-Speech Models

The evaluation framework captures distinct metrics for synthesis and recognition tasks.

TTS Benchmark Metrics

In scripts/benchmark_tts.py, the following measurements are collected:

  • Warm-up time: Seconds required to load the model and allocate buffers before the first synthesis. This is captured immediately after handler construction using result.warmup_time = time.perf_counter() - start_setup (lines 70-71).
  • Inference time: End-to-end latency for a single synthesis request, calculated as time_taken = end_time - start_time within the benchmark loop (lines 92-94).
  • Time-to-first-chunk (TTFC): Latency until the first audio chunk is produced, critical for real-time user experience. Measured when the first chunk arrives via time_to_first_chunk = time.perf_counter() - start_time (lines 82-84).
  • Audio duration: Length of the generated waveform calculated as audio_duration = total_samples / DEFAULT_SAMPLE_RATE (lines 94-95).
  • Real-Time Factor (RTF): The ratio of audio_duration / inference_time, computed in BenchmarkResult.get_stats() (lines 64-66). An RTF ≤ 1 indicates faster-than-real-time synthesis.

STT Benchmark Metrics

In scripts/benchmark_stt.py, the evaluation focuses on:

  • Warm-up time: Initialization duration before the first transcription, measured using the same pattern as TTS during handler creation.
  • Inference time: Duration required to transcribe a full utterance, recorded per iteration.
  • Time-to-first-token (TTFT): Latency until the first token of the transcript is emitted, stored in self.time_to_first_token when the handler yields its first result.
  • Sample transcription: One example output saved as sample_transcription to sanity-check correctness alongside performance metrics.

Benchmarking Methodology in the Codebase

Both benchmark scripts follow an identical five-phase structure:

  1. Setup: Create a stop_event, input/output queues, and instantiate the handler (handler = …Handler(...)).
  2. Warm-up: Measure initialization time immediately after handler construction.
  3. Benchmark loop: Execute the handler for a specified number of iterations, capturing per-iteration latency and auxiliary data.
  4. Result aggregation: Compute mean, min, max, standard deviation, and derived metrics like RTF using BenchmarkResult.get_stats().
  5. Reporting: Output a formatted table to stdout and optionally write JSON results via the --output parameter.

This methodology allows you to compare any combination of handlers, from Kokoro-82M and Qwen3-TTS to Whisper, Faster-Whisper, and Lightning-Whisper-MLX.

Running TTS Benchmarks

To evaluate text-to-speech performance, use scripts/benchmark_tts.py with the --text parameter to specify input prompts and --iterations to control statistical stability.

Benchmark the default three handlers (kokoro, qwen3, pocket_tts):

python scripts/benchmark_tts.py \
    --text "Hello, this is a latency test." \
    --iterations 5 \
    --output tts_perf.json

Evaluate Qwen3-TTS across all supported MLX quantizations on Apple Silicon:

python scripts/benchmark_tts.py \
    --handlers qwen3 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit \
    --iterations 3 \
    --output qwen3_mlx_bench.json

Running STT Benchmarks

For speech-to-text evaluation, use scripts/benchmark_stt.py with the --audio_file parameter pointing to a test waveform.

Compare Whisper (CUDA) and Parakeet-TDT on the same audio file:

python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper parakeet-tdt \
    --iterations 4 \
    --output stt_perf.json

Benchmark Lightning-Whisper-MLX on Apple Silicon:

python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper-mlx \
    --iterations 3 \
    --output mlx_whisper.json

Interpreting Benchmark Results

The scripts output JSON files containing aggregated statistics. Here is an example result from a TTS benchmark:

{
  "results": [
    {
      "handler": "qwen3[bf16]",
      "warmup_time": 2.14,
      "avg_inference_time": 0.87,
      "min_inference_time": 0.81,
      "max_inference_time": 0.93,
      "std_inference_time": 0.04,
      "avg_audio_duration": 1.23,
      "avg_rtf": 1.41,
      "avg_time_to_first_chunk": 0.12,
      "total_iterations": 3,
      "errors": []
    }
  ],
  "timestamp": "2026-07-07 12:34:56"
}

Key fields map directly to the metrics described above. The avg_rtf value of 1.41 indicates that synthesis takes 1.41 times longer than the audio duration (slower than real-time), while avg_time_to_first_chunk shows the initial latency before audio streaming begins.

Summary

  • The huggingface/speech-to-speech repository provides scripts/benchmark_tts.py and scripts/benchmark_stt.py to evaluate the performance of a speech-to-speech model.
  • TTS evaluation measures warm-up time, inference latency, time-to-first-chunk (TTFC), and real-time factor (RTF) in BenchmarkResult.get_stats().
  • STT evaluation tracks warm-up time, transcription latency, and time-to-first-token (TTFT).
  • Both scripts support multiple handlers via the --handlers argument and output JSON for statistical analysis.
  • An RTF ≤ 1 indicates the model can process audio faster than real-time, essential for interactive voice applications.

Frequently Asked Questions

What is Real-Time Factor (RTF) in speech synthesis?

Real-Time Factor (RTF) is the ratio of audio duration to inference time, calculated as audio_duration / inference_time in the BenchmarkResult.get_stats() method. An RTF value less than or equal to 1.0 means the model generates audio faster than it takes to play it, which is the minimum requirement for real-time conversational agents. Values greater than 1.0 indicate latency that would cause noticeable delays in live interactions.

How do I benchmark speech-to-speech models on Apple Silicon?

To evaluate models on Apple Silicon devices, use MLX-optimized handlers. For TTS, run scripts/benchmark_tts.py with --handlers qwen3 and specify --qwen3_mlx_quantizations with values like bf16, 4bit, 6bit, or 8bit. For STT, use --handlers whisper-mlx with scripts/benchmark_stt.py. These handlers leverage the MLX framework for accelerated inference on M-series chips.

What is the difference between TTFC and TTFT in speech models?

Time-to-first-chunk (TTFC) applies to TTS models and measures the delay before the first audio byte is produced after receiving text input, captured at lines 82-84 in benchmark_tts.py. Time-to-first-token (TTFT) applies to STT models and measures the delay before the first text character is transcribed from audio input. Both metrics are critical for perceived responsiveness in voice agents, as they represent the initial latency before streaming output begins.

Where are the benchmark scripts located in the repository?

The benchmark scripts reside in the scripts/ directory at the repository root. Specifically, scripts/benchmark_tts.py handles text-to-speech evaluation, while scripts/benchmark_stt.py handles speech-to-text evaluation. Both scripts can be invoked directly from the command line and support the --help flag to view all available parameters, including --handlers, --iterations, and --output.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →