# Qwen3-TTS vs Pocket TTS vs ChatTTS vs MMS TTS: Voice Quality and Latency Comparison

> Compare Qwen3-TTS, Pocket TTS, ChatTTS, and MMS TTS. Discover voice quality and latency differences to choose the best TTS for your needs. See which offers low latency and high fidelity.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: comparison
- Published: 2026-07-30

---

**Qwen3-TTS delivers the highest voice quality with true token-level streaming latency under 200 ms, while Pocket TTS runs CPU-only with frame-wise generation, ChatTTS uses chunk-based streaming, and MMS TTS generates full utterances before playback, creating a quality-latency tradeoff spectrum from high-fidelity/low-latency to resource-efficient/high-latency.**

The [huggingface/speech-to-speech](https://github.com/huggingface/speech-to-speech) repository implements four distinct text-to-speech handlers, each with unique architectures that directly impact voice naturalness and time-to-first-audio (TTFA). Understanding the differences between **Qwen3-TTS**, **Pocket TTS**, **ChatTTS**, and **Facebook MMS TTS** helps you select the optimal backend for real-time conversational AI or batch synthesis workloads.

## Voice Quality Characteristics

### Qwen3-TTS: High-Fidelity Voice Cloning and Design

**Qwen3-TTS** utilizes 1.7 billion parameter models including `CustomVoice`, `Voice-Design`, and `Voice-Clone` variants. According to the implementation in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), this handler supports multilingual synthesis with automatic language detection via `QWEN3_LANGUAGE_ALIASES` mapping.

The model architecture enables three distinct voice customization modes:
- **Custom-voice**: Pre-defined speaker profiles
- **Voice-design**: Free-form text prompts for speaker characteristics
- **Voice-clone**: Reference audio + text input for speaker replication

These modes produce the most natural prosody and expressivity among the four options, with support for dynamic token budgets via `_estimate_max_new_tokens`.

### Pocket TTS: Lightweight English Synthesis

**Pocket TTS** from Kyutai Labs operates with fewer than 500 million parameters, loading models through `pocket_tts.TTSModel` as implemented in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py). The handler provides a preset catalog including voices like `alba`, `jean`, and `fantine`.

Quality tends toward "synthetic" clarity for short utterances due to the smaller backbone and limited speaker diversity. Unlike Qwen3-TTS, Pocket TTS restricts output to English only and lacks explicit quantization options.

### ChatTTS: Random Speaker Embeddings

**ChatTTS** employs a VITS-style architecture with approximately 300 million parameters. The `ChatTTSHandler` in [`src/speech_to_speech/TTS/chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/chatTTS_handler.py) initializes with `ChatTTS.Chat()` and generates a random speaker embedding at startup via `sample_random_speaker()`.

While capable of natural-sounding multilingual synthesis, the non-deterministic speaker embedding can introduce inconsistency across conversation turns. The model generates full waveform chunks at 24 kHz before resampling to 16 kHz.

### Facebook MMS TTS: Single-Speaker Per Language

**Facebook MMS TTS** uses language-specific VITS models with approximately 700 million parameters each. The `FacebookMMSTTSHandler` in [`src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/facebookmms_handler.py) maps Whisper language codes to MMS suffixes through `WHISPER_LANGUAGE_TO_FACEBOOK_LANGUAGE`.

Each language model defines exactly one voice, resulting in a "robotic" quality compared to Qwen3-TTS. The architecture lacks speaker diversity, producing consistent but less expressive output across supported languages including English, French, Spanish, German, and Russian.

## Latency and Streaming Architecture

Streaming implementation determines time-to-first-audio (TTFA) more than model size alone.

### True Token Streaming: Qwen3-TTS

**Qwen3-TTS** implements genuine token-level streaming with the lowest latency:
- **Non-macOS**: Uses `faster-qwen3-tts` backend with configurable `DEFAULT_FASTER_STREAMING_CHUNK_SIZE` of 8 tokens
- **macOS**: Uses **mlx-audio** at approximately 12.5 tokens per second with default chunk size of 4 tokens

This architecture achieves **< 200 ms on CUDA ≥ 12 GPUs** or **< 150 ms on Apple Silicon**, emitting audio continuously rather than buffering complete chunks.

### Frame-Wise Generation: Pocket TTS

**Pocket TTS** generates audio frame-by-frame (approximately 10-20 ms per frame) in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py). The handler buffers these frames until accumulating enough samples for the pipeline's 512-sample blocksize.

This CPU-only approach yields **~300-400 ms TTFA**, with latency dominated by the per-frame generation loop and resampling from 24 kHz to 16 kHz.

### Chunk-Based Generation: ChatTTS

**ChatTTS** produces complete waveform chunks via `model.infer(..., stream=True)` before yielding any audio. As implemented in [`src/speech_to_speech/TTS/chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/chatTTS_handler.py), each chunk undergoes resampling from 24 kHz to 16 kHz before transmission.

This strategy results in **~350-500 ms latency on GPU**, higher than Qwen3-TTS because the model must synthesize entire chunks before playback begins.

### Full Utterance Generation: MMS TTS

**Facebook MMS TTS** exhibits the highest latency, generating the **complete waveform** in one forward pass via `model(..., return_dict=False)`. When `stream=True`, the handler in [`src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/facebookmms_handler.py) merely slices the pre-generated waveform into 512-sample blocks.

This architecture produces **~600-800 ms TTFA** on GPU, as the entire utterance synthesizes before any audio emission occurs.

## Hardware Requirements and Quantization

Hardware flexibility varies significantly across implementations:

- **Qwen3-TTS**: Supports GPU via `ggml` or `torch` backends with CUDA graphs, plus Apple Silicon acceleration through **mlx-audio** with optional `bf16/4bit/6bit/8bit` quantization
- **Pocket TTS**: CPU-only execution with no quantization options, suitable for edge deployment without GPU
- **ChatTTS**: Requires GPU for real-time performance; no explicit quantization flags available
- **MMS TTS**: GPU-optimized; CPU execution possible but impractical for real-time applications

## Implementation Examples

The [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) CLI instantiates each handler with backend-specific arguments. Below are minimal configurations for each TTS system:

**Qwen3-TTS (GPU with GGML):**

```python
python s2s_pipeline.py \
  --tts qwen3 \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
  --qwen3_tts_device cuda \
  --qwen3_tts_backend ggml \
  --qwen3_tts_speaker Aiden \
  --qwen3_tts_language auto

```

**Pocket TTS (CPU):**

```python
python s2s_pipeline.py \
  --tts pocket \
  --pocket_tts_voice jean \
  --pocket_tts_device cpu \
  --pocket_tts_sample_rate 16000

```

**ChatTTS (GPU):**

```python
python s2s_pipeline.py \
  --tts chatTTS \
  --chat_tts_device cuda \
  --chat_tts_stream true \
  --chat_tts_chunk_size 512

```

**Facebook MMS TTS (GPU, English):**

```python
python s2s_pipeline.py \
  --tts facebookMMS \
  --facebook_mms_device cuda \
  --tts_language en

```

### Benchmarking Latency

The repository includes [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) for empirical comparison:

```bash
python benchmark_tts.py \
  --handlers qwen3 pocket chatTTS facebookMMS \
  --iterations 5

```

This script reports TTFA and RTF (real-time factor) metrics, confirming the theoretical latency hierarchy: Qwen3-TTS < Pocket TTS < ChatTTS < MMS TTS.

## Summary

- **Voice quality** improves with model flexibility: **Qwen3-TTS > ChatTTS ≈ MMS TTS > Pocket TTS**, with Qwen3-TTS uniquely supporting voice cloning and design prompts.
- **Latency** correlates with streaming granularity: Qwen3-TTS achieves **< 200 ms** via token streaming, while MMS TTS requires **~600-800 ms** for full utterance generation.
- **Hardware requirements** range from CPU-only (Pocket TTS) to GPU-optimized with Apple Silicon support (Qwen3-TTS).
- **Implementation files** driving these differences include [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py), [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py), [`chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/chatTTS_handler.py), and [`facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/facebookmms_handler.py) in the `src/speech_to_speech/TTS/` directory.

## Frequently Asked Questions

### Which TTS handler offers the lowest latency for real-time conversations?

**Qwen3-TTS** provides the lowest latency through true token-level streaming, achieving **< 200 ms on modern GPUs** and **< 150 ms on Apple Silicon** via the MLX backend. This contrasts with Facebook MMS TTS, which buffers the entire utterance before playback, resulting in **~600-800 ms** delays.

### Can I run these TTS models without a GPU?

Only **Pocket TTS** is optimized for CPU-only operation, though this limits output to English and increases latency to **~300-400 ms**. Qwen3-TTS supports Apple Silicon (M1/M2/M3) via MLX, while ChatTTS and MMS TTS require GPUs for usable real-time performance.

### How does voice customization differ between Qwen3-TTS and ChatTTS?

**Qwen3-TTS** in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) supports explicit voice cloning from reference audio and text-based voice design prompts. **ChatTTS** generates random speaker embeddings at startup via `sample_random_speaker()`, offering less consistency across sessions but requiring no reference audio.

### Why does MMS TTS have higher latency despite using streaming parameters?

The `FacebookMMSTTSHandler` generates the **complete waveform** in `model(..., return_dict=False)` before slicing it into blocks. Setting `stream=True` only affects post-generation chunking, not synthesis latency. This differs from Qwen3-TTS, which emits audio tokens continuously during generation.