Qwen3-TTS vs Pocket TTS vs ChatTTS vs MMS TTS: Voice Quality and Latency Comparison
Qwen3-TTS delivers the highest voice quality with true token-level streaming latency under 200 ms, while Pocket TTS runs CPU-only with frame-wise generation, ChatTTS uses chunk-based streaming, and MMS TTS generates full utterances before playback, creating a quality-latency tradeoff spectrum from high-fidelity/low-latency to resource-efficient/high-latency.
The huggingface/speech-to-speech repository implements four distinct text-to-speech handlers, each with unique architectures that directly impact voice naturalness and time-to-first-audio (TTFA). Understanding the differences between Qwen3-TTS, Pocket TTS, ChatTTS, and Facebook MMS TTS helps you select the optimal backend for real-time conversational AI or batch synthesis workloads.
Voice Quality Characteristics
Qwen3-TTS: High-Fidelity Voice Cloning and Design
Qwen3-TTS utilizes 1.7 billion parameter models including CustomVoice, Voice-Design, and Voice-Clone variants. According to the implementation in src/speech_to_speech/TTS/qwen3_tts_handler.py, this handler supports multilingual synthesis with automatic language detection via QWEN3_LANGUAGE_ALIASES mapping.
The model architecture enables three distinct voice customization modes:
- Custom-voice: Pre-defined speaker profiles
- Voice-design: Free-form text prompts for speaker characteristics
- Voice-clone: Reference audio + text input for speaker replication
These modes produce the most natural prosody and expressivity among the four options, with support for dynamic token budgets via _estimate_max_new_tokens.
Pocket TTS: Lightweight English Synthesis
Pocket TTS from Kyutai Labs operates with fewer than 500 million parameters, loading models through pocket_tts.TTSModel as implemented in src/speech_to_speech/TTS/pocket_tts_handler.py. The handler provides a preset catalog including voices like alba, jean, and fantine.
Quality tends toward "synthetic" clarity for short utterances due to the smaller backbone and limited speaker diversity. Unlike Qwen3-TTS, Pocket TTS restricts output to English only and lacks explicit quantization options.
ChatTTS: Random Speaker Embeddings
ChatTTS employs a VITS-style architecture with approximately 300 million parameters. The ChatTTSHandler in src/speech_to_speech/TTS/chatTTS_handler.py initializes with ChatTTS.Chat() and generates a random speaker embedding at startup via sample_random_speaker().
While capable of natural-sounding multilingual synthesis, the non-deterministic speaker embedding can introduce inconsistency across conversation turns. The model generates full waveform chunks at 24 kHz before resampling to 16 kHz.
Facebook MMS TTS: Single-Speaker Per Language
Facebook MMS TTS uses language-specific VITS models with approximately 700 million parameters each. The FacebookMMSTTSHandler in src/speech_to_speech/TTS/facebookmms_handler.py maps Whisper language codes to MMS suffixes through WHISPER_LANGUAGE_TO_FACEBOOK_LANGUAGE.
Each language model defines exactly one voice, resulting in a "robotic" quality compared to Qwen3-TTS. The architecture lacks speaker diversity, producing consistent but less expressive output across supported languages including English, French, Spanish, German, and Russian.
Latency and Streaming Architecture
Streaming implementation determines time-to-first-audio (TTFA) more than model size alone.
True Token Streaming: Qwen3-TTS
Qwen3-TTS implements genuine token-level streaming with the lowest latency:
- Non-macOS: Uses
faster-qwen3-ttsbackend with configurableDEFAULT_FASTER_STREAMING_CHUNK_SIZEof 8 tokens - macOS: Uses mlx-audio at approximately 12.5 tokens per second with default chunk size of 4 tokens
This architecture achieves < 200 ms on CUDA ≥ 12 GPUs or < 150 ms on Apple Silicon, emitting audio continuously rather than buffering complete chunks.
Frame-Wise Generation: Pocket TTS
Pocket TTS generates audio frame-by-frame (approximately 10-20 ms per frame) in src/speech_to_speech/TTS/pocket_tts_handler.py. The handler buffers these frames until accumulating enough samples for the pipeline's 512-sample blocksize.
This CPU-only approach yields ~300-400 ms TTFA, with latency dominated by the per-frame generation loop and resampling from 24 kHz to 16 kHz.
Chunk-Based Generation: ChatTTS
ChatTTS produces complete waveform chunks via model.infer(..., stream=True) before yielding any audio. As implemented in src/speech_to_speech/TTS/chatTTS_handler.py, each chunk undergoes resampling from 24 kHz to 16 kHz before transmission.
This strategy results in ~350-500 ms latency on GPU, higher than Qwen3-TTS because the model must synthesize entire chunks before playback begins.
Full Utterance Generation: MMS TTS
Facebook MMS TTS exhibits the highest latency, generating the complete waveform in one forward pass via model(..., return_dict=False). When stream=True, the handler in src/speech_to_speech/TTS/facebookmms_handler.py merely slices the pre-generated waveform into 512-sample blocks.
This architecture produces ~600-800 ms TTFA on GPU, as the entire utterance synthesizes before any audio emission occurs.
Hardware Requirements and Quantization
Hardware flexibility varies significantly across implementations:
- Qwen3-TTS: Supports GPU via
ggmlortorchbackends with CUDA graphs, plus Apple Silicon acceleration through mlx-audio with optionalbf16/4bit/6bit/8bitquantization - Pocket TTS: CPU-only execution with no quantization options, suitable for edge deployment without GPU
- ChatTTS: Requires GPU for real-time performance; no explicit quantization flags available
- MMS TTS: GPU-optimized; CPU execution possible but impractical for real-time applications
Implementation Examples
The s2s_pipeline.py CLI instantiates each handler with backend-specific arguments. Below are minimal configurations for each TTS system:
Qwen3-TTS (GPU with GGML):
python s2s_pipeline.py \
--tts qwen3 \
--qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
--qwen3_tts_device cuda \
--qwen3_tts_backend ggml \
--qwen3_tts_speaker Aiden \
--qwen3_tts_language auto
Pocket TTS (CPU):
python s2s_pipeline.py \
--tts pocket \
--pocket_tts_voice jean \
--pocket_tts_device cpu \
--pocket_tts_sample_rate 16000
ChatTTS (GPU):
python s2s_pipeline.py \
--tts chatTTS \
--chat_tts_device cuda \
--chat_tts_stream true \
--chat_tts_chunk_size 512
Facebook MMS TTS (GPU, English):
python s2s_pipeline.py \
--tts facebookMMS \
--facebook_mms_device cuda \
--tts_language en
Benchmarking Latency
The repository includes benchmark_tts.py for empirical comparison:
python benchmark_tts.py \
--handlers qwen3 pocket chatTTS facebookMMS \
--iterations 5
This script reports TTFA and RTF (real-time factor) metrics, confirming the theoretical latency hierarchy: Qwen3-TTS < Pocket TTS < ChatTTS < MMS TTS.
Summary
- Voice quality improves with model flexibility: Qwen3-TTS > ChatTTS ≈ MMS TTS > Pocket TTS, with Qwen3-TTS uniquely supporting voice cloning and design prompts.
- Latency correlates with streaming granularity: Qwen3-TTS achieves < 200 ms via token streaming, while MMS TTS requires ~600-800 ms for full utterance generation.
- Hardware requirements range from CPU-only (Pocket TTS) to GPU-optimized with Apple Silicon support (Qwen3-TTS).
- Implementation files driving these differences include
qwen3_tts_handler.py,pocket_tts_handler.py,chatTTS_handler.py, andfacebookmms_handler.pyin thesrc/speech_to_speech/TTS/directory.
Frequently Asked Questions
Which TTS handler offers the lowest latency for real-time conversations?
Qwen3-TTS provides the lowest latency through true token-level streaming, achieving < 200 ms on modern GPUs and < 150 ms on Apple Silicon via the MLX backend. This contrasts with Facebook MMS TTS, which buffers the entire utterance before playback, resulting in ~600-800 ms delays.
Can I run these TTS models without a GPU?
Only Pocket TTS is optimized for CPU-only operation, though this limits output to English and increases latency to ~300-400 ms. Qwen3-TTS supports Apple Silicon (M1/M2/M3) via MLX, while ChatTTS and MMS TTS require GPUs for usable real-time performance.
How does voice customization differ between Qwen3-TTS and ChatTTS?
Qwen3-TTS in src/speech_to_speech/TTS/qwen3_tts_handler.py supports explicit voice cloning from reference audio and text-based voice design prompts. ChatTTS generates random speaker embeddings at startup via sample_random_speaker(), offering less consistency across sessions but requiring no reference audio.
Why does MMS TTS have higher latency despite using streaming parameters?
The FacebookMMSTTSHandler generates the complete waveform in model(..., return_dict=False) before slicing it into blocks. Setting stream=True only affects post-generation chunking, not synthesis latency. This differs from Qwen3-TTS, which emits audio tokens continuously during generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →