# How to Evaluate the Performance of a Speech-to-Speech Model

> Evaluate speech-to-speech model performance with huggingface tools. Measure latency, real-time factor, and TTS/STT component speed for better results.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-08-02

---

**Benchmark speech-to-speech models using the huggingface/speech-to-speech repository's built-in scripts to measure latency, real-time factor, and time-to-first-chunk across TTS and STT components.**

Evaluating speech-to-speech (S2S) systems requires more than accuracy metrics—you need precise latency measurements to ensure real-time voice-agent responsiveness. The `huggingface/speech-to-speech` repository provides dedicated benchmark scripts for measuring the performance of each pipeline stage.

## Understanding S2S Pipeline Components

Speech-to-speech pipelines chain four interchangeable components: **VAD → STT → LLM → TTS**. Performance bottlenecks can occur at any stage, so the repository isolates the two most latency-sensitive modules—STT and TTS—for standalone benchmarking.

| Component | Critical Metrics |
|-----------|----------------|
| **STT** | Warm-up time, inference latency, time-to-first-token (TTFT) |
| **TTS** | Warm-up time, inference latency, time-to-first-chunk (TTFC), real-time factor (RTF) |

## Benchmarking TTS Performance

The **[`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py)** script measures synthesis speed across any supported TTS handler: Kokoro-82M, Qwen3-TTS, Pocket-TTS, or others.

### Key TTS Metrics Explained

- **Warm-up time**: Seconds to load the model and allocate buffers before first synthesis (`result.warmup_time = time.perf_counter() - start_setup` in [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) lines 70-71)
- **Inference time**: End-to-end latency for a single request (`time_taken = end_time - start_time`, lines 92-94)
- **Time-to-first-chunk (TTFC)**: Critical for UX—how long until audio starts streaming (`time_to_first_chunk = time.perf_counter() - start_time`, lines 82-84)
- **Real-time factor (RTF)**: `audio_duration / inference_time`—values ≤ 1 indicate faster-than-real-time synthesis (computed in `BenchmarkResult.get_stats()`, lines 64-66)

### Running TTS Benchmarks

```bash

# Basic usage — benchmark default handlers (kokoro, qwen3, pocket_tts)

python scripts/benchmark_tts.py \
    --text "Hello, this is a latency test." \
    --iterations 5 \
    --output tts_perf.json

```

```bash

# Compare Qwen3-TTS across all MLX quantizations (Apple Silicon)

python scripts/benchmark_tts.py \
    --handlers qwen3 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit \
    --iterations 3 \
    --output qwen3_mlx_bench.json

```

## Benchmarking STT Performance

The **[`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py)** script evaluates speech recognition latency using Whisper, Faster-Whisper, Lightning-Whisper-MLX, Paraformer, or Parakeet-TDT backends.

### Key STT Metrics

- **Warm-up time**: Model initialization before first transcription
- **Inference time**: Full utterance transcription latency
- **Time-to-first-token (TTFT)**: Streaming responsiveness metric
- **Sample transcription**: Sanity-check output for correctness validation

### Running STT Benchmarks

```bash

# Compare Whisper (CUDA) and Parakeet-TDT on identical audio

python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper parakeet-tdt \
    --iterations 4 \
    --output stt_perf.json

```

```bash

# Benchmark Lightning-Whisper-MLX on Apple Silicon

python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper-mlx \
    --iterations 3 \
    --output mlx_whisper.json

```

## Interpreting Benchmark Results

Both scripts follow an identical execution pattern:

1. **Setup** — create `stop_event`, input/output queues, instantiate handler
2. **Warm-up** — measure initialization time
3. **Benchmark loop** — run `iterations` times, collecting per-call latency
4. **Aggregation** — compute mean, min, max, standard deviation, and derived metrics
5. **Reporting** — print summary table and optionally save JSON

### Sample JSON Output Structure

```json
{
  "results": [
    {
      "handler": "qwen3[bf16]",
      "warmup_time": 2.14,
      "avg_inference_time": 0.87,
      "min_inference_time": 0.81,
      "max_inference_time": 0.93,
      "std_inference_time": 0.04,
      "avg_audio_duration": 1.23,
      "avg_rtf": 1.41,
      "avg_time_to_first_chunk": 0.12,
      "total_iterations": 3,
      "errors": []
    }
  ],
  "timestamp": "2026-07-07 12:34:56"
}

```

## How to Evaluate the Performance of a Speech-to-Speech Model in Practice

Complete evaluation requires testing both directions of the pipeline:

1. **Establish baseline** — run TTS benchmark with `--iterations 10` minimum for statistical stability
2. **Measure STT separately** — use representative audio samples matching your target domain
3. **Combine results** — sum `avg_inference_time` from both stages plus estimated LLM latency for total pipeline latency
4. **Validate RTF < 1** — ensure synthesis keeps pace with real-time requirements

For handler-specific documentation, refer to:
- [`src/speech_to_speech/TTS/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/README.md) — TTS handler configurations
- [`src/speech_to_speech/STT/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/README.md) — supported STT models and parameters

## Summary

- **Use [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py)** to measure synthesis latency, TTFC, and RTF across any TTS backend
- **Use [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py)** to evaluate recognition speed and TTFT for STT models
- **Target RTF ≤ 1** for real-time voice agents—higher values cause perceptible delays
- **Run multiple iterations** (5-10 minimum) to obtain stable statistics with standard deviation
- **Compare quantizations** using built-in flags like `--qwen3_mlx_quantizations` for deployment optimization

## Frequently Asked Questions

### What is Real-Time Factor (RTF) in TTS evaluation?

**RTF measures synthesis speed relative to audio duration.** Calculated as `audio_duration / inference_time` in `BenchmarkResult.get_stats()`, an RTF of 1.0 means synthesis takes exactly as long as the utterance itself. Values below 1.0 indicate faster-than-real-time generation. For interactive voice agents, maintain RTF well under 1.0 to prevent cumulative latency.

### How does time-to-first-chunk (TTFC) differ from inference time?

**TTFC captures streaming latency; inference time measures total duration.** TTFC represents when the first audio bytes are ready (`time.perf_counter() - start_time` at first chunk arrival), crucial for perceived responsiveness. Inference time spans the complete synthesis. A model with low inference time but high TTFC feels sluggish to users.

### Can I benchmark custom TTS or STT handlers?

**Yes—any handler conforming to the repository's interface is compatible.** The benchmark scripts accept arbitrary handler names via `--handlers` and instantiate them through the standard factory pattern. Ensure your handler implements the required `__call__` signature accepting `(queue, stop_event)` and yielding results progressively.

### What hardware considerations affect benchmark validity?

**Apple Silicon optimizations (MLX) and CUDA backends show significant divergence.** The repository provides quantization-specific flags like `--qwen3_mlx_quantizations` because `bf16`, `4bit`, `6bit`, and `8bit` variants exhibit different latency-accuracy tradeoffs. Always benchmark on target deployment hardware—results from CUDA systems do not transfer to MLX or CPU deployments.