How to Evaluate the Performance of a Speech-to-Speech Model

Benchmark speech-to-speech models using the huggingface/speech-to-speech repository's built-in scripts to measure latency, real-time factor, and time-to-first-chunk across TTS and STT components.

Evaluating speech-to-speech (S2S) systems requires more than accuracy metrics—you need precise latency measurements to ensure real-time voice-agent responsiveness. The huggingface/speech-to-speech repository provides dedicated benchmark scripts for measuring the performance of each pipeline stage.

Understanding S2S Pipeline Components

Speech-to-speech pipelines chain four interchangeable components: VAD → STT → LLM → TTS. Performance bottlenecks can occur at any stage, so the repository isolates the two most latency-sensitive modules—STT and TTS—for standalone benchmarking.

Component Critical Metrics
STT Warm-up time, inference latency, time-to-first-token (TTFT)
TTS Warm-up time, inference latency, time-to-first-chunk (TTFC), real-time factor (RTF)

Benchmarking TTS Performance

The scripts/benchmark_tts.py script measures synthesis speed across any supported TTS handler: Kokoro-82M, Qwen3-TTS, Pocket-TTS, or others.

Key TTS Metrics Explained

  • Warm-up time: Seconds to load the model and allocate buffers before first synthesis (result.warmup_time = time.perf_counter() - start_setup in benchmark_tts.py lines 70-71)
  • Inference time: End-to-end latency for a single request (time_taken = end_time - start_time, lines 92-94)
  • Time-to-first-chunk (TTFC): Critical for UX—how long until audio starts streaming (time_to_first_chunk = time.perf_counter() - start_time, lines 82-84)
  • Real-time factor (RTF): audio_duration / inference_time—values ≤ 1 indicate faster-than-real-time synthesis (computed in BenchmarkResult.get_stats(), lines 64-66)

Running TTS Benchmarks


# Basic usage — benchmark default handlers (kokoro, qwen3, pocket_tts)

python scripts/benchmark_tts.py \
    --text "Hello, this is a latency test." \
    --iterations 5 \
    --output tts_perf.json

# Compare Qwen3-TTS across all MLX quantizations (Apple Silicon)

python scripts/benchmark_tts.py \
    --handlers qwen3 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit \
    --iterations 3 \
    --output qwen3_mlx_bench.json

Benchmarking STT Performance

The scripts/benchmark_stt.py script evaluates speech recognition latency using Whisper, Faster-Whisper, Lightning-Whisper-MLX, Paraformer, or Parakeet-TDT backends.

Key STT Metrics

  • Warm-up time: Model initialization before first transcription
  • Inference time: Full utterance transcription latency
  • Time-to-first-token (TTFT): Streaming responsiveness metric
  • Sample transcription: Sanity-check output for correctness validation

Running STT Benchmarks


# Compare Whisper (CUDA) and Parakeet-TDT on identical audio

python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper parakeet-tdt \
    --iterations 4 \
    --output stt_perf.json

# Benchmark Lightning-Whisper-MLX on Apple Silicon

python scripts/benchmark_stt.py \
    --audio_file samples/sample.wav \
    --handlers whisper-mlx \
    --iterations 3 \
    --output mlx_whisper.json

Interpreting Benchmark Results

Both scripts follow an identical execution pattern:

  1. Setup — create stop_event, input/output queues, instantiate handler
  2. Warm-up — measure initialization time
  3. Benchmark loop — run iterations times, collecting per-call latency
  4. Aggregation — compute mean, min, max, standard deviation, and derived metrics
  5. Reporting — print summary table and optionally save JSON

Sample JSON Output Structure

{
  "results": [
    {
      "handler": "qwen3[bf16]",
      "warmup_time": 2.14,
      "avg_inference_time": 0.87,
      "min_inference_time": 0.81,
      "max_inference_time": 0.93,
      "std_inference_time": 0.04,
      "avg_audio_duration": 1.23,
      "avg_rtf": 1.41,
      "avg_time_to_first_chunk": 0.12,
      "total_iterations": 3,
      "errors": []
    }
  ],
  "timestamp": "2026-07-07 12:34:56"
}

How to Evaluate the Performance of a Speech-to-Speech Model in Practice

Complete evaluation requires testing both directions of the pipeline:

  1. Establish baseline — run TTS benchmark with --iterations 10 minimum for statistical stability
  2. Measure STT separately — use representative audio samples matching your target domain
  3. Combine results — sum avg_inference_time from both stages plus estimated LLM latency for total pipeline latency
  4. Validate RTF < 1 — ensure synthesis keeps pace with real-time requirements

For handler-specific documentation, refer to:

Summary

  • Use scripts/benchmark_tts.py to measure synthesis latency, TTFC, and RTF across any TTS backend
  • Use scripts/benchmark_stt.py to evaluate recognition speed and TTFT for STT models
  • Target RTF ≤ 1 for real-time voice agents—higher values cause perceptible delays
  • Run multiple iterations (5-10 minimum) to obtain stable statistics with standard deviation
  • Compare quantizations using built-in flags like --qwen3_mlx_quantizations for deployment optimization

Frequently Asked Questions

What is Real-Time Factor (RTF) in TTS evaluation?

RTF measures synthesis speed relative to audio duration. Calculated as audio_duration / inference_time in BenchmarkResult.get_stats(), an RTF of 1.0 means synthesis takes exactly as long as the utterance itself. Values below 1.0 indicate faster-than-real-time generation. For interactive voice agents, maintain RTF well under 1.0 to prevent cumulative latency.

How does time-to-first-chunk (TTFC) differ from inference time?

TTFC captures streaming latency; inference time measures total duration. TTFC represents when the first audio bytes are ready (time.perf_counter() - start_time at first chunk arrival), crucial for perceived responsiveness. Inference time spans the complete synthesis. A model with low inference time but high TTFC feels sluggish to users.

Can I benchmark custom TTS or STT handlers?

Yes—any handler conforming to the repository's interface is compatible. The benchmark scripts accept arbitrary handler names via --handlers and instantiate them through the standard factory pattern. Ensure your handler implements the required __call__ signature accepting (queue, stop_event) and yielding results progressively.

What hardware considerations affect benchmark validity?

Apple Silicon optimizations (MLX) and CUDA backends show significant divergence. The repository provides quantization-specific flags like --qwen3_mlx_quantizations because bf16, 4bit, 6bit, and 8bit variants exhibit different latency-accuracy tradeoffs. Always benchmark on target deployment hardware—results from CUDA systems do not transfer to MLX or CPU deployments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →