# How to Benchmark and Profile the Different Pipeline Components for Optimization in Hugging Face Speech-to-Speech

> Learn to benchmark and profile Hugging Face Speech-to-Speech pipeline components for optimization. Use provided scripts and profilers to measure latency, throughput, and RTF for peak performance.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-07-08

---

**The huggingface/speech-to-speech repository provides ready-to-use scripts ([`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) and [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py)) that measure latency, throughput, and Real-Time Factor (RTF) for each pipeline stage, and can be extended with `cProfile` and `torch.profiler` for deep performance analysis.**

To systematically optimize a real-time speech-to-speech system, you must isolate and measure the cost of each modular component. The `speech-to-speech` repository implements a queue-based architecture where **VAD**, **STT**, **LLM**, and **TTS** handlers run as independent threads, allowing you to benchmark individual stages or the full pipeline using built-in tools and standard Python profilers.

## Understanding the Modular Pipeline Architecture

The pipeline is assembled at runtime by `build_pipeline()` in [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py), which creates typed queues (`Queue[AudioInItem]`, `Queue[STTOutItem]`, etc.) and events that connect independent handlers. This design allows each component to be instantiated and measured in isolation.

| Component | Responsibility | Core Implementation |
|-----------|----------------|----------------------|
| **VAD** | Detects speech boundaries and streams partial audio | [`speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/VAD/vad_handler.py) |
| **STT** | Converts audio to text (Whisper, Faster-Whisper, MLX-Audio) | [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py) → `get_stt_handler()` |
| **LLM** | Generates responses via transformers, MLX-LM, or APIs | [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py) → `get_llm_handler()` |
| **TTS** | Synthesizes audio (Kokoro, Qwen3-TTS, Pocket-TTS, etc.) | [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py) → `get_tts_handler()` |

Because each handler runs in its own thread managed by the `ThreadManager` abstraction, you can benchmark **only TTS** to measure synthesis latency, or **only STT** to measure transcription overhead, without initializing the full pipeline.

## Benchmarking Individual Components

The repository ships with two primary benchmarking scripts that follow a standardized pattern: warm-up, timed iteration loop, and statistical aggregation.

### Text-to-Speech Benchmarking with [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py)

Located at [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py), this script measures wall-clock time, warm-up latency, and **Real-Time Factor (RTF)**—calculated as `audio_duration / inference_time`—for any TTS handler.

```bash
python scripts/benchmark_tts.py \
    --text "Hello from the speech-to-speech benchmark. This is a latency test." \
    --handlers kokoro qwen3 pocket_tts \
    --iterations 5 \
    --output tts_bench.json

```

The script instantiates handlers like `Qwen3TTSHandler`, runs a warm-up phase to load weights and allocate tensors, then processes the input repeatedly while recording `time.perf_counter()` values. Results are saved as a `BenchmarkResult` object (defined at lines 34-44 of the script) containing mean, min, max, and standard deviation for inference time and **Time-to-First-Chunk (TTFC)**.

### Speech-to-Text Benchmarking with [`benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_stt.py)

Located at [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py), this script loads audio via `soundfile` and optional `librosa` resampling (see the `load_audio()` function at lines 84-107), then benchmarks STT handlers such as Whisper, Faster-Whisper, or MLX-Audio-Whisper.

```bash
python scripts/benchmark_stt.py \
    --audio_file samples/short.wav \
    --handlers whisper mlx-audio-whisper faster-whisper \
    --iterations 3 \
    --output stt_bench.json

```

Both scripts output a comparison table ranking handlers by average inference time, making it easy to identify the fastest configuration for your specific hardware.

## Deep Profiling for Optimization

The benchmark scripts provide high-level latency numbers, but optimization requires identifying specific bottlenecks in Python code, GPU kernels, or memory allocation. You can wrap the existing benchmark loops with standard profilers without modifying the handler implementations.

### CPU Profiling with `cProfile`

To diagnose queue overhead or transcription post-processing bottlenecks, wrap the `benchmark_handler` function with `cProfile`:

```python
import cProfile, pstats, io
from scripts.benchmark_tts import benchmark_handler

def run_profiled(handler_name, text, iterations):
    pr = cProfile.Profile()
    pr.enable()
    result = benchmark_handler(handler_name, text, iterations)
    pr.disable()
    
    s = io.StringIO()
    ps = pstats.Stats(pr, stream=s).sort_stats('cumtime')
    ps.print_stats(20)
    print(f"Profiling report for {handler_name}:\n", s.getvalue())
    return result

```

This prints the top 20 functions by cumulative time, highlighting costly operations like `torch.nn.functional.linear` or `numpy.mean` that may not be obvious from end-to-end latency alone.

### GPU and MLX Profiling with `torch.profiler`

For CUDA or Apple Silicon optimization, use `torch.profiler` to capture kernel launch overhead and memory bottlenecks:

```python
import torch.profiler as profiler
from scripts.benchmark_tts import benchmark_handler

def run_gpu_profiled(handler_name, text, iterations):
    with profiler.profile(
        schedule=profiler.schedule(wait=1, warmup=1, active=3, repeat=1),
        on_trace_ready=profiler.tensorboard_trace_handler("./tb_logs"),
        record_shapes=True,
        profile_memory=True,
        with_stack=True,
    ) as prof:
        result = benchmark_handler(handler_name, text, iterations)
        prof.step()
    
    print(f"GPU profile for {handler_name} stored in ./tb_logs")
    return result

```

Open the trace in TensorBoard with `tensorboard --logdir ./tb_logs` to visualize kernel-level timings. This works for both **CUDA** and **MLX** backends, as the repository uses a `torch` shim for MLX compatibility.

## Practical Workflow for Optimization

Follow this systematic approach to benchmark and profile the different pipeline components for optimization:

1. **Establish baseline measurements** – Run the vanilla [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) or [`benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_stt.py) on your target hardware (Apple Silicon, RTX 3090, or CPU-only).
2. **Identify the bottleneck** – Use the comparison table to determine which stage (VAD, STT, LLM, or TTS) contributes most to end-to-end latency.
3. **Instrument with profilers** – Wrap the heavy handler with `cProfile` (for Python overhead) or `torch.profiler` (for GPU/MLX kernels).
4. **Iterate on configuration** – Test device arguments (`device="cuda"`, `device="mps"`, `device="cpu"`), quantization levels (e.g., `normalize_qwen3_mlx_quantizations()` in [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) supports 4-bit and 8-bit for Qwen3-TTS), and live transcription toggles.
5. **Version-control results** – The scripts automatically write JSON summaries ([`tts_benchmark_results.json`](https://github.com/huggingface/speech-to-speech/blob/main/tts_benchmark_results.json)). Compare these files across commits to catch performance regressions.
6. **Apply hardware-specific optimizations** – Enable `local_mac_optimal_settings` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (see `optimal_mac_settings()` around line 31) or adjust the realtime pool size with `--num_pipelines`.

## Complete Code Examples

### End-to-End Pipeline Benchmark

To measure the full VAD→STT→LLM→TTS chain, use the argument parsing logic from [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py):

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --num_pipelines 2 \
    --stt whisper \
    --tts qwen3 \
    --llm_backend responses-api \
    --iterations 5 \
    --output full_pipeline_bench.json

```

This builds a realtime pool with two parallel pipelines and outputs the same statistical format as the individual benchmark scripts.

### Profiling Qwen3-TTS with PyTorch Profiler

Create a dedicated profiling script to test quantization effects on Apple Silicon:

```python

# profile_qwen3_tts.py

from scripts.benchmark_tts import benchmark_handler, normalize_qwen3_mlx_quantizations
import torch.profiler as profiler

handler_name = "qwen3"
text = "Benchmarking Qwen3-TTS on Apple Silicon"
iterations = 3

quantizations = normalize_qwen3_mlx_quantizations(["4bit"])
handler_kwargs = {"mlx_quantization": quantizations[0]}

with profiler.profile(
    schedule=profiler.schedule(wait=1, warmup=1, active=2),
    on_trace_ready=profiler.tensorboard_trace_handler("./tensorboard_qwen3"),
    record_shapes=True,
    profile_memory=True,
):
    result = benchmark_handler(
        handler_name, text, iterations, handler_kwargs=handler_kwargs
    )

print("Result:", result.get_stats())

```

Run with `python profile_qwen3_tts.py` and visualize with `tensorboard --logdir ./tensorboard_qwen3`.

### Diagnosing STT Queue Overhead with `cProfile`

To isolate queue handling costs from model inference in [`speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/STT/whisper_stt_handler.py):

```python
import cProfile, pstats, io
from scripts.benchmark_stt import benchmark_handler

def profile_stt(handler_name, audio_path, iterations):
    pr = cProfile.Profile()
    pr.enable()
    result = benchmark_handler(handler_name, audio_path, iterations)
    pr.disable()
    
    s = io.StringIO()
    ps = pstats.Stats(pr, stream=s).sort_stats('cumtime')
    ps.print_stats(15)
    print(f"cProfile for {handler_name}:\n", s.getvalue())
    return result

```

Replace the direct call in [`benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_stt.py) with `profile_stt(...)` to see exactly how much time is spent in queue operations versus the Whisper model forward pass.

## Summary

- **Use the built-in scripts** ([`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) and [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py)) to generate baseline latency, RTF, and TTFC metrics for individual handlers.
- **Leverage modularity** – The queue-based architecture in [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py) allows you to benchmark `get_tts_handler()` or `get_stt_handler()` in isolation without initializing the full pipeline.
- **Add granular profiling** – Wrap benchmark calls with `cProfile` for Python-level bottlenecks or `torch.profiler` for GPU/MLX kernel analysis.
- **Optimize via configuration** – Test `device` arguments, `mlx_quantization` levels (via `normalize_qwen3_mlx_quantizations()`), and `num_pipelines` to find the optimal throughput for your hardware.
- **Monitor regressions** – Commit the JSON output files ([`tts_benchmark_results.json`](https://github.com/huggingface/speech-to-speech/blob/main/tts_benchmark_results.json)) to track performance changes across code updates.

## Frequently Asked Questions

### How do I measure the Real-Time Factor (RTF) for TTS handlers?

The [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) script calculates RTF automatically by dividing the audio duration by the inference time. After running the benchmark with `--iterations`, check the output table or JSON file for the RTF column. Values below 1.0 indicate real-time capability (faster than playback), while values above 1.0 indicate the synthesis is slower than real-time.

### Can I benchmark individual pipeline stages without running the full speech-to-speech system?

Yes. Because the pipeline uses a modular architecture with independent handlers connected via queues, you can import and benchmark individual components directly. Use `get_stt_handler()` or `get_tts_handler()` from [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py) to instantiate a single handler, then pass it to the benchmark functions in [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) or [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py) without building the full pipeline.

### What is the difference between using `cProfile` and `torch.profiler` for optimization?

Use **`cProfile`** when you need to identify slow Python functions, data preprocessing, or queue management overhead in the CPU-bound sections of the code. Use **`torch.profiler`** when optimizing GPU or Apple MLX execution, as it captures kernel launch times, memory bandwidth bottlenecks, and tensor shapes. For comprehensive optimization, run `cProfile` first to find the slow stage, then use `torch.profiler` on that specific handler to analyze its compute kernels.

### How do I optimize for Apple Silicon specifically?

Enable Apple-specific optimizations by using the `mlx-audio-whisper` handler for STT and `qwen3` handler with MLX quantization for TTS. In [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py), pass quantization levels through `normalize_qwen3_mlx_quantizations()` (e.g., `["4bit"]` or `["8bit"]`). Additionally, set `local_mac_optimal_settings=True` when building the pipeline in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), which configures optimal thread counts and memory settings for M1/M2/M3 chips.