# How to Evaluate the Performance of a Speech-to-Speech Model: A Complete Benchmarking Guide

> Learn how to evaluate speech-to-speech model performance with Hugging Face's benchmark scripts. Measure latency, real-time factor, and more across the entire pipeline.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: tutorial
- Published: 2026-08-01

---

**You can evaluate speech-to-speech model performance using the dedicated benchmark scripts [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) and [`benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_stt.py) in the Hugging Face speech-to-speech repository, which measure latency, real-time factor, and time-to-first-chunk/token across the VAD → STT → LLM → TTS pipeline components.**

The `huggingface/speech-to-speech` repository implements a modular pipeline architecture where audio flows through four interchangeable stages: Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS). To evaluate the performance of a speech-to-speech model effectively, the codebase provides specialized benchmarking tools that quantify latency and throughput for the TTS and STT components specifically.

## Understanding the Speech-to-Speech Pipeline Components

The repository structures speech-to-speech inference as a chain of handlers. When you evaluate the performance of a speech-to-speech model, you are essentially measuring how quickly each stage processes data and delivers results to the next component. The benchmark scripts focus on the two most latency-critical stages: **STT** (input processing) and **TTS** (output generation).

## Key Performance Metrics for Speech-to-Speech Models

The evaluation framework captures distinct metrics for synthesis and recognition tasks.

### TTS Benchmark Metrics

In [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py), the following measurements are collected:

- **Warm-up time**: Seconds required to load the model and allocate buffers before the first synthesis. This is captured immediately after handler construction using `result.warmup_time = time.perf_counter() - start_setup` (lines 70-71).
- **Inference time**: End-to-end latency for a single synthesis request, calculated as `time_taken = end_time - start_time` within the benchmark loop (lines 92-94).
- **Time-to-first-chunk (TTFC)**: Latency until the first audio chunk is produced, critical for real-time user experience. Measured when the first chunk arrives via `time_to_first_chunk = time.perf_counter() - start_time` (lines 82-84).
- **Audio duration**: Length of the generated waveform calculated as `audio_duration = total_samples / DEFAULT_SAMPLE_RATE` (lines 94-95).
- **Real-Time Factor (RTF)**: The ratio of `audio_duration / inference_time`, computed in `BenchmarkResult.get_stats()` (lines 64-66). An RTF ≤ 1 indicates faster-than-real-time synthesis.

### STT Benchmark Metrics

In [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py), the evaluation focuses on:

- **Warm-up time**: Initialization duration before the first transcription, measured using the same pattern as TTS during handler creation.
- **Inference time**: Duration required to transcribe a full utterance, recorded per iteration.
- **Time-to-first-token (TTFT)**: Latency until the first token of the transcript is emitted, stored in `self.time_to_first_token` when the handler yields its first result.
- **Sample transcription**: One example output saved as `sample_transcription` to sanity-check correctness alongside performance metrics.

## Benchmarking Methodology in the Codebase

Both benchmark scripts follow an identical five-phase structure:

1. **Setup**: Create a `stop_event`, input/output queues, and instantiate the handler (`handler = …Handler(...)`).
2. **Warm-up**: Measure initialization time immediately after handler construction.
3. **Benchmark loop**: Execute the handler for a specified number of `iterations`, capturing per-iteration latency and auxiliary data.
4. **Result aggregation**: Compute mean, min, max, standard deviation, and derived metrics like RTF using `BenchmarkResult.get_stats()`.
5. **Reporting**: Output a formatted table to stdout and optionally write JSON results via the `--output` parameter.

This methodology allows you to compare any combination of handlers, from **Kokoro-82M** and **Qwen3-TTS** to **Whisper**, **Faster-Whisper**, and **Lightning-Whisper-MLX**.

## Running TTS Benchmarks

To evaluate text-to-speech performance, use [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) with the `--text` parameter to specify input prompts and `--iterations` to control statistical stability.

Benchmark the default three handlers (kokoro, qwen3, pocket_tts):

```bash
python scripts/benchmark_tts.py \
    --text "Hello, this is a latency test." \
    --iterations 5 \
    --output tts_perf.json

```

Evaluate Qwen3-TTS across all supported MLX quantizations on Apple Silicon:

```bash
python scripts/benchmark_tts.py \
    --handlers qwen3 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit \
    --iterations 3 \
    --output qwen3_mlx_bench.json

```

## Running STT Benchmarks

For speech-to-text evaluation, use [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py) with the `--audio_file` parameter pointing to a test waveform.

Compare Whisper (CUDA) and Parakeet-TDT on the same audio file:

```bash
python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper parakeet-tdt \
    --iterations 4 \
    --output stt_perf.json

```

Benchmark Lightning-Whisper-MLX on Apple Silicon:

```bash
python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper-mlx \
    --iterations 3 \
    --output mlx_whisper.json

```

## Interpreting Benchmark Results

The scripts output JSON files containing aggregated statistics. Here is an example result from a TTS benchmark:

```json
{
  "results": [
    {
      "handler": "qwen3[bf16]",
      "warmup_time": 2.14,
      "avg_inference_time": 0.87,
      "min_inference_time": 0.81,
      "max_inference_time": 0.93,
      "std_inference_time": 0.04,
      "avg_audio_duration": 1.23,
      "avg_rtf": 1.41,
      "avg_time_to_first_chunk": 0.12,
      "total_iterations": 3,
      "errors": []
    }
  ],
  "timestamp": "2026-07-07 12:34:56"
}

```

Key fields map directly to the metrics described above. The `avg_rtf` value of 1.41 indicates that synthesis takes 1.41 times longer than the audio duration (slower than real-time), while `avg_time_to_first_chunk` shows the initial latency before audio streaming begins.

## Summary

- The `huggingface/speech-to-speech` repository provides [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) and [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py) to evaluate the performance of a speech-to-speech model.
- **TTS evaluation** measures warm-up time, inference latency, time-to-first-chunk (TTFC), and real-time factor (RTF) in `BenchmarkResult.get_stats()`.
- **STT evaluation** tracks warm-up time, transcription latency, and time-to-first-token (TTFT).
- Both scripts support multiple handlers via the `--handlers` argument and output JSON for statistical analysis.
- An RTF ≤ 1 indicates the model can process audio faster than real-time, essential for interactive voice applications.

## Frequently Asked Questions

### What is Real-Time Factor (RTF) in speech synthesis?

**Real-Time Factor (RTF)** is the ratio of audio duration to inference time, calculated as `audio_duration / inference_time` in the `BenchmarkResult.get_stats()` method. An RTF value less than or equal to 1.0 means the model generates audio faster than it takes to play it, which is the minimum requirement for real-time conversational agents. Values greater than 1.0 indicate latency that would cause noticeable delays in live interactions.

### How do I benchmark speech-to-speech models on Apple Silicon?

To evaluate models on Apple Silicon devices, use MLX-optimized handlers. For TTS, run [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) with `--handlers qwen3` and specify `--qwen3_mlx_quantizations` with values like `bf16`, `4bit`, `6bit`, or `8bit`. For STT, use `--handlers whisper-mlx` with [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py). These handlers leverage the MLX framework for accelerated inference on M-series chips.

### What is the difference between TTFC and TTFT in speech models?

**Time-to-first-chunk (TTFC)** applies to TTS models and measures the delay before the first audio byte is produced after receiving text input, captured at lines 82-84 in [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py). **Time-to-first-token (TTFT)** applies to STT models and measures the delay before the first text character is transcribed from audio input. Both metrics are critical for perceived responsiveness in voice agents, as they represent the initial latency before streaming output begins.

### Where are the benchmark scripts located in the repository?

The benchmark scripts reside in the `scripts/` directory at the repository root. Specifically, [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) handles text-to-speech evaluation, while [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py) handles speech-to-text evaluation. Both scripts can be invoked directly from the command line and support the `--help` flag to view all available parameters, including `--handlers`, `--iterations`, and `--output`.