# Optimizing TTS Streaming Latency in the Hugging Face Speech-to-Speech Library

> Optimize TTS streaming latency with small chunk sizes, streaming mode, minimal audio buffers, and speculative turns to achieve sub-100ms first-audio latency on GPUs. Learn best practices.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: best-practices
- Published: 2026-07-10

---

**The best practices for optimizing TTS streaming latency involve configuring small chunk sizes (256–512 tokens), forcing streaming mode over non-streaming buffers, minimizing audio output buffers, and enabling speculative turns to achieve sub-100ms first-audio latency on modern GPUs.**

Real-time speech-to-speech systems demand near-instantaneous text-to-speech generation to maintain natural conversational flow. The huggingface/speech-to-speech repository achieves low-latency TTS by streaming audio chunks incrementally rather than waiting for full utterance synthesis. By tuning specific parameters in the TTS handlers and pipeline configuration, you can reduce end-to-end latency from text input to audible output below 100 milliseconds on CUDA-enabled hardware and under 200 milliseconds on Apple Silicon.

## Configure Streaming Chunk Size

The primary lever for controlling latency is the **streaming chunk size**, which determines how much audio generates before emitting to the output stream. Smaller chunks deliver audio to the listener faster but increase CPU overhead and context-switching.

The default value of `512` tokens produces approximately 30ms of audio per chunk. For minimal latency, reduce this to `256` tokens (~15ms) by setting the `streaming_chunk_size` parameter in `Qwen3TTSArguments`. In [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), the `_resolve_streaming_chunk_size` method validates this configuration, ensuring the value meets hardware constraints while maintaining real-time performance.

## Force Streaming Mode Over Non-Streaming

Avoid the **non-streaming mode**, which buffers the entire utterance before emitting any audio. This mode adds the full synthesis duration to your first-audio latency.

Set `non_streaming_mode=False` (the default) to enable true incremental streaming. In [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), the initialization logic checks this flag; when disabled, the handler yields audio chunks as the model generates them rather than waiting for end-of-sequence markers. Note that on Apple Silicon, MLX-Audio does not expose the `non_streaming_mode` flag (as documented in lines 147-148 of [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py)), so the system automatically relies on the streaming API.

## Select Hardware-Specific Backends

Different hardware paths offer distinct latency characteristics:

- **CUDA/GPU**: Use the `faster-qwen3-tts` backend (enabled by default) for maximum token throughput and minimal per-chunk processing time.
- **Apple Silicon**: Keep `non_streaming_mode=False` and rely on the native streaming implementation, as the MLX-Audio integration requires the streaming path.

## Minimize Audio Output Buffer Latency

The audio sink introduces significant latency through internal buffering and OS mixing. Configure the `LocalAudioStreamer` in [`src/speech_to_speech/connections/local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/local_audio_streamer.py) with a minimal `blocksize` (e.g., `256`) to reduce the delay between chunk generation and speaker output.

Ensure the streamer's `sample_rate` matches the model output (typically 24000Hz) to eliminate resampling overhead, which can add 10-20ms of processing time per chunk.

## Enable Speculative Turn Processing

Implement **speculative turns** to begin TTS generation before the LLM completes its full response. The `SpeculativeTurns` class in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py) allows the pipeline to predict upcoming utterances based on partial LLM outputs, overlapping TTS computation with ongoing language model inference.

Combine this with a short VAD horizon (`vad_horizon=0.1s`) to trigger generation promptly when the user stops speaking, reducing the gap between turn completion and audio playback.

## Implementation Guide

### 1. Configure the TTS Handler for Low Latency

```python
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments

# Configure for minimal latency

tts_args = Qwen3TTSArguments(
    streaming_chunk_size=256,  # ~15ms audio per chunk

    non_streaming_mode=False,  # Force streaming mode

)
tts_handler = Qwen3TTSHandler(tts_args)

```

### 2. Set Up Low-Latency Audio Output

```python
from speech_to_speech.connections.local_audio_streamer import LocalAudioStreamer

audio_streamer = LocalAudioStreamer(
    blocksize=256,      # Minimal internal buffer

    sample_rate=24000,  # Match model output sample rate

)

```

### 3. Assemble the Pipeline with Speculative Turns

```python
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.pipeline.speculative_turns import SpeculativeTurns

pipeline = SpeechToSpeechPipeline(
    tts_handler=tts_handler,
    audio_streamer=audio_streamer,
    speculative_turns=SpeculativeTurns(enable=True),
)

```

## Monitor Latency Metrics

Validate your optimizations using the built-in telemetry in [`src/speech_to_speech/api/openai_realtime/service.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/service.py) (line 138), which logs `tts_latency`, `llm_latency`, `vad_latency`, and `stt_latency` for each interaction. The `_log_first_audio_latency` method in [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) (lines 740-749) specifically tracks the time from text input to first audio emission, allowing you to verify that your configuration achieves the target sub-100ms latency on GPU or sub-200ms on Apple Silicon.

## Summary

- Set `streaming_chunk_size=256` in `Qwen3TTSArguments` to reduce audio chunks to approximately 15ms
- Keep `non_streaming_mode=False` to enable true streaming in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)
- Instantiate `LocalAudioStreamer` with `blocksize=256` to minimize audio output buffering
- Enable `SpeculativeTurns` to overlap TTS generation with LLM inference
- Monitor `tts_latency` and `first_audio_latency` metrics in the service logs to validate sub-100ms performance on GPUs

## Frequently Asked Questions

### What is the optimal chunk size for TTS streaming latency?

A chunk size of **256 tokens** (~15ms of audio) provides the optimal balance between latency and CPU overhead for most real-time applications. However, ensure your target hardware can process each chunk in under 10ms to prevent backlog and audio stuttering. The default of 512 tokens (~30ms) offers a safer margin for slower hardware.

### Why does non-streaming mode increase latency?

Non-streaming mode buffers the entire text generation before emitting audio, adding the full synthesis duration to your first-audio latency. Streaming mode emits chunks as they are generated, typically reducing latency by an order of magnitude from seconds to milliseconds.

### How do I verify my latency optimizations are working?

The pipeline logs `tts_latency` and `first_audio_latency` metrics in [`src/speech_to_speech/api/openai_realtime/service.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/service.py). Check these values in your console output; optimized configurations should show first-audio latency under 0.1 seconds on GPU and under 0.2 seconds on Apple Silicon.

### Does speculative turn processing affect audio quality?

No, speculative turns only affect **when** generation begins, not the synthesis process itself. The TTS model still processes the full text token-by-token; it simply starts earlier based on predicted utterances. This reduces perceived latency without altering the audio output or introducing artifacts.