# How to Evaluate the Performance of a Speech-to-Speech Model: Benchmarking Guide

> Learn how to evaluate speech-to-speech model performance with our benchmarking guide. Measure warm-up time, latency, and real-time factor using huggingface/speech-to-speech scripts.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: benchmarking-guide
- Published: 2026-07-07

---

**Evaluate speech-to-speech model performance by measuring warm-up time, inference latency, time-to-first-chunk/token, and real-time factor using the dedicated benchmark scripts in the `huggingface/speech-to-speech` repository.**

The `huggingface/speech-to-speech` repository implements a modular Speech-to-Speech (S2S) pipeline composed of four interchangeable components: **VAD → STT → LLM → TTS**. To evaluate the performance of a speech-to-speech model, the repository provides specialized benchmark scripts that quantify latency, throughput, and real-time capability for the TTS and STT stages.

## Understanding the Evaluation Metrics

The benchmark scripts capture five critical performance indicators for each component. These metrics determine whether a voice agent can deliver real-time conversational experiences.

### TTS Performance Metrics

In [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py), the following metrics are collected per handler:

- **Warm-up time**: Seconds required to load the model and allocate buffers before the first synthesis. Measured immediately after handler construction using `result.warmup_time = time.perf_counter() - start_setup` (lines 70-71).
- **Inference time**: End-to-end latency for a single synthesis request, calculated as `time_taken = end_time - start_time` within the benchmark loop (lines 92-94).
- **Time-to-first-chunk (TTFC)**: Latency until the first audio chunk is produced, critical for real-time user experience. Captured when the first chunk arrives via `time_to_first_chunk = time.perf_counter() - start_time` (lines 82-84).
- **Audio duration**: Length of the generated waveform in seconds, computed as `audio_duration = total_samples / DEFAULT_SAMPLE_RATE` (lines 94-95).
- **Real-Time Factor (RTF)**: Ratio of `audio_duration / inference_time`. An RTF ≤ 1 indicates faster-than-real-time synthesis, computed in `BenchmarkResult.get_stats()` (lines 64-66).

### STT Performance Metrics

The [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py) script tracks analogous metrics for speech recognition:

- **Warm-up time**: Initialization duration before the first transcription, measured using the same pattern as TTS during handler creation.
- **Inference time**: Time required to transcribe a full utterance, recorded per iteration as `time_taken = end_time - start_time`.
- **Time-to-first-token (TTFT)**: Latency until the first token of the transcript is emitted, stored in `self.time_to_first_token` when the handler yields its first result.
- **Sample transcription**: One example output saved as `sample_transcription` in the stats dictionary to sanity-check correctness.

## The Benchmarking Methodology

Both [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) and [`benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_stt.py) follow an identical five-step structure:

1. **Setup**: Create a `stop_event`, input/output queues, and instantiate the handler (`handler = …Handler(...)`).
2. **Warm-up**: Measure initialization time captured as `warmup_time`.
3. **Benchmark loop**: Execute the handler for a specified number of `iterations`, measuring per-iteration latency and auxiliary data (audio length, TTFC/TTFT).
4. **Result aggregation**: Compute mean, min, max, standard deviation, and RTF/TTFC statistics via `BenchmarkResult.get_stats()`.
5. **Reporting**: Pretty-print a summary table and optionally export JSON via `save_results`.

## Running the TTS Benchmark

Benchmark TTS handlers to compare synthesis speed across different models and quantization levels.

### Basic Usage

Benchmark the default handlers (`kokoro`, `qwen3`, `pocket_tts`) using a test prompt:

```bash
python scripts/benchmark_tts.py \
    --text "Hello, this is a latency test." \
    --iterations 5 \
    --output tts_perf.json

```

### Apple Silicon MLX Quantization

For Apple Silicon devices, benchmark Qwen3-TTS across multiple MLX quantization levels:

```bash
python scripts/benchmark_tts.py \
    --handlers qwen3 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit \
    --iterations 3 \
    --output qwen3_mlx_bench.json

```

## Running the STT Benchmark

Evaluate speech recognition latency using different back-ends such as Whisper, Faster-Whisper, Lightning-Whisper-MLX, or Paraformer.

### Multi-Handler Comparison

Benchmark Whisper (CUDA) and Parakeet-TDT on the same audio file:

```bash
python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper parakeet-tdt \
    --iterations 4 \
    --output stt_perf.json

```

### Apple Silicon Specific

Benchmark Lightning-Whisper-MLX optimized for Apple Silicon:

```bash
python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper-mlx \
    --iterations 3 \
    --output mlx_whisper.json

```

## Interpreting the Results

The benchmark scripts generate JSON output containing detailed statistics. Here is an excerpt from a TTS benchmark showing Qwen3 performance:

```json
{
  "results": [
    {
      "handler": "qwen3[bf16]",
      "warmup_time": 2.14,
      "avg_inference_time": 0.87,
      "min_inference_time": 0.81,
      "max_inference_time": 0.93,
      "std_inference_time": 0.04,
      "avg_audio_duration": 1.23,
      "avg_rtf": 1.41,
      "avg_time_to_first_chunk": 0.12,
      "total_iterations": 3,
      "errors": []
    }
  ],
  "timestamp": "2026-07-07 12:34:56"
}

```

Key fields map directly to the implementation: `warmup_time` reflects the initialization measured at lines 70-71 of [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py), while `avg_rtf` derives from the `BenchmarkResult.get_stats()` method at lines 64-66. An RTF value of 1.41 indicates the model takes 1.41 seconds to generate 1 second of audio—slower than real-time—whereas values below 1.0 indicate real-time capability.

## Summary

- **Measure component-specific metrics**: Track warm-up time, inference latency, and time-to-first-output for both TTS (TTFC) and STT (TTFT) stages.
- **Use provided benchmark scripts**: Execute [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) and [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py) to evaluate any combination of handlers.
- **Monitor Real-Time Factor**: Ensure RTF ≤ 1 for real-time conversational agents; higher values indicate latency that will degrade user experience.
- **Compare quantizations**: Test different model sizes and quantization levels (e.g., Qwen3 MLX 4-bit vs. 8-bit) to find the optimal speed-quality trade-off.
- **Validate correctness**: Check `sample_transcription` in STT results to ensure speed optimizations do not compromise accuracy.

## Frequently Asked Questions

### What is the Real-Time Factor (RTF) and why does it matter?

The **Real-Time Factor (RTF)** is the ratio of audio duration to inference time. An RTF of 1.0 means the model generates audio exactly as fast as it plays; values below 1.0 indicate faster-than-real-time generation. For conversational speech-to-speech systems, maintaining RTF ≤ 1 is essential to prevent perceptible delays in responses.

### How do I benchmark specific handlers rather than the defaults?

Pass the `--handlers` flag followed by the specific handler names you want to evaluate. For example, `--handlers qwen3` will benchmark only the Qwen3-TTS handler, while `--handlers whisper parakeet-tdt` will compare those two STT back-ends. Refer to [`src/speech_to_speech/TTS/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/README.md) and [`src/speech_to_speech/STT/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/README.md) for the complete list of supported handlers.

### What is the difference between TTFC and TTFT?

**Time-to-first-chunk (TTFC)** applies to TTS and measures the latency until the first audio byte is produced, crucial for perceived responsiveness in voice agents. **Time-to-first-token (TTFT)** applies to STT and measures the latency until the first transcription character is emitted. Both metrics determine how quickly a speech-to-speech system can begin responding to user input.

### Can I evaluate the LLM component using these scripts?

The current benchmark scripts in [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) and [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py) focus specifically on the TTS and STT components. The LLM evaluation is handled separately through the function calling utilities in [`src/speech_to_speech/LLM/tool_call/function_call.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_call.py), which validates parsing accuracy rather than inference latency.