# How to Improve Speech-to-Speech Inference Latency: A Technical Guide for the Hugging Face Pipeline

> Learn how to reduce speech-to-speech inference latency with efficient device selection speculative turns and queue tuning in the Hugging Face pipeline Optimize your system now

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-07-07

---

**Reduce end-to-end latency in the huggingface/speech-to-speech system by optimizing device selection, enabling speculative turns, and tuning the queue architecture.**

The huggingface/speech-to-speech repository provides a modular real-time conversation pipeline that chains speech-to-text (STT), large language models (LLM), and text-to-speech (TTS) handlers. Because these components operate in separate threads communicating via queues, latency optimization requires careful attention to hardware acceleration, parallel processing via speculative turns, and minimizing unnecessary computational overhead.

## Select Optimal Hardware Acceleration

The pipeline’s latency is fundamentally bounded by device selection. In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the `prepare_all_args` function calls `overwrite_device_argument` (lines 263–277) to enforce GPU utilization when available.

**Use CUDA for NVIDIA GPUs.** The default device mapping routes STT, LLM, and TTS to `cuda` when detected, eliminating costly CPU-to-GPU memory copies.

**Use MLX for Apple Silicon.** When `--local_mac_optimal_settings` is enabled, the `optimal_mac_settings` helper (lines 31–38) automatically switches the LLM backend to `mlx` and the TTS model to `qwen3` with MPS device targeting. This leverages Apple’s unified memory architecture to avoid buffer transfers.

**Avoid CPU fallback.** The `module_kwargs` device handling in `prepare_all_args` explicitly sets `mps` for Apple Silicon and `cuda` for NVIDIA cards; overriding these to `cpu` increases latency by 5–10×.

## Enable Speculative Turn Processing

The **SpeculativeTurnTracker** class (instantiated in `build_pipeline` → `_build_pipeline_handlers`) enables parallel processing by streaming LLM output tokens to the TTS handler before the full response completes.

**How speculative turns work.** Rather than waiting for the LLM to finish generating the entire text, the tracker dispatches partial outputs to the TTS queue immediately. This overlaps LLM inference with TTS synthesis, effectively hiding the TTS generation time within the LLM’s token generation.

**Activate in realtime mode.** When running with `--mode realtime`, the pipeline automatically constructs multiple parallel units via `_build_realtime_pipeline_unit` (lines 45–63). Each unit maintains its own `SpeculativeTurnTracker`, allowing concurrent processing of multiple conversation turns:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --num_pipelines 2 \
    --enable_speculative_turns

```

## Optimize Queue Architecture and Buffering

The pipeline uses Python’s `queue.Queue` objects to pass data between STT, LLM, and TTS threads. Tuning these queues prevents bottlenecks that introduce latency jitter.

**Monitor the STT-to-LLM bridge.** The `stt_output_queue` feeds into `text_prompt_queue` via the `TranscriptionNotifier`. Keep these queues unbounded (`maxsize=0`) for real-time streaming, but ensure sufficient RAM to prevent memory pressure.

**Reduce LM output buffering.** The `LMOutputProcessor` (located in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py)) buffers LLM tokens before passing them to the TTS handler. Lowering the internal batch size in this processor reduces latency at the cost of slightly higher TTS invocation frequency.

**Adjust audio chunk sizes.** In `SocketReceiverArguments`, reduce the default `chunk_size` (≈1600 samples) to decrease the time spent accumulating audio before STT processing begins.

## Choose Low-Latency STT Backends

The `get_stt_handler` function (lines 71–84) supports multiple STT implementations with varying latency characteristics.

**Use faster-whisper for production.** The `FasterWhisperSTTHandler` (accessed via `--stt faster-whisper`) provides 30–50% speedup over vanilla Whisper with comparable word error rates, as it runs Whisper in C++ with optimized transformers.

**Use parakeet-tdt on Apple Silicon.** When `module_kwargs.stt` is set to `"parakeet-tdt"`, the pipeline loads NVIDIA’s Parakeet model optimized for streaming inference on CUDA and Apple Silicon, offering lower latency than encoder-decoder architectures.

**Use mlx-audio-whisper for MLX devices.** If running on M1/M2/M3 chips, specify `--stt mlx-audio-whisper` to run the STT inference through the MLX backend, avoiding PyTorch overhead entirely.

## Disable Non-Essential Features

**Turn off live transcription.** The VAD handler runs a secondary loop for incremental transcription updates when `enable_live_transcription` is true. Disable this via `--enable_live_transcription false` (see `VADHandlerArguments` processing at lines 229–238) to reduce per-chunk computation overhead when only final transcripts are needed.

**Remove debug logging.** The pipeline sets `torch._logging.set_logs` when `log_level == "debug"`. Running with `--log_level info` eliminates the synchronization overhead of detailed PyTorch inductor logging.

## Leverage Caching and Compilation

**Cache model weights on fast storage.** The pipeline sets `TORCHINDUCTOR_CACHE_DIR` to a local `tmp` directory derived from `Path(__file__).resolve().parent` (line 84). Ensure this directory resides on an NVMe or SSD to reduce model loading and compilation stalls by approximately 50% on first runs.

**Pre-compile with torch.compile.** While the standard pipeline uses eager mode, you can manually apply `torch.compile` to the LLM and TTS models before loading them into the pipeline handlers. This eliminates just-in-time compilation warm-up latency during the first inference request.

## Complete Configuration Example

The following command combines all latency optimizations for an NVIDIA GPU deployment:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --num_pipelines 2 \
    --stt faster-whisper \
    --llm_backend transformers \
    --tts qwen3 \
    --device cuda \
    --enable_speculative_turns true \
    --enable_live_transcription false \
    --log_level info

```

This configuration:
1. Runs two parallel pipelines in realtime mode
2. Uses CUDA for all components
3. Streams tokens via speculative turns
4. Disables live transcription overhead
5. Employs the faster-whisper backend for minimal STT latency

## Summary

- **Select native accelerators** (CUDA for NVIDIA, MLX for Apple Silicon) via `overwrite_device_argument` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) to eliminate device transfer overhead.
- **Enable speculative turns** to overlap LLM generation with TTS synthesis, instantiated in `_build_pipeline_handlers`.
- **Tune queue depths** and reduce `LMOutputProcessor` buffering to minimize inter-thread latency.
- **Choose faster-whisper, parakeet-tdt, or mlx-audio-whisper** depending on your hardware for optimal STT speed.
- **Disable live transcription** and debug logging to reduce per-chunk computational overhead.
- **Cache models on NVMe storage** and consider `torch.compile` for eliminating warm-up latency.

## Frequently Asked Questions

### What is speculative turn processing in speech-to-speech inference?

**Speculative turn processing** is a parallelism technique where the pipeline streams partial LLM outputs to the TTS handler before the full text generation completes. Implemented via the `SpeculativeTurnTracker` class in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py), this approach overlaps the LLM’s token generation time with the TTS synthesis process, effectively reducing total end-to-end latency by the duration of the TTS generation.

### How do I reduce latency on Apple Silicon devices?

On Apple Silicon (M1/M2/M3), enable the `--local_mac_optimal_settings` flag. This triggers the `optimal_mac_settings` helper (lines 31–38 in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)) to switch the LLM backend to MLX and the TTS model to Qwen3 running on MPS. Additionally, use `--stt mlx-audio-whisper` to run speech recognition through the MLX framework rather than PyTorch, avoiding cross-framework memory copies.

### Which STT model offers the lowest latency in this pipeline?

For NVIDIA GPUs, **faster-whisper** provides the best latency-to-accuracy ratio, offering 30–50% speedup over standard Whisper. For Apple Silicon, **mlx-audio-whisper** or **parakeet-tdt** provide lower latency by leveraging specialized backends (MLX for Whisper, optimized CUDA kernels for Parakeet) that minimize the overhead of the PyTorch runtime.

### Why does the pipeline use queues between components, and how do I tune them?

The pipeline uses `queue.Queue` objects to enable asynchronous communication between the STT, LLM, and TTS threads, preventing blocking I/O from stalling the audio stream. To tune them, ensure the `stt_output_queue` and `lm_response_queue` are unbounded (`maxsize=0`) for realtime applications, and reduce the `chunk_size` in `SocketReceiverArguments` to decrease audio accumulation delays. If memory becomes constrained, set explicit `maxsize` limits and monitor the `LMOutputProcessor` batch size in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py).