What Models Are Supported by Hugging Face Speech-to-Speech: Complete Model Guide

Hugging Face Speech-to-Speech supports modular backends across four pipeline stages: Silero VAD for voice detection, six STT options including Parakeet TDT and Whisper variants, OpenAI-compatible APIs and Transformers for LLMs, and five TTS models including Qwen3-TTS and Kokoro.

The Hugging Face Speech-to-Speech (S2S) repository is a fully modular pipeline that lets you swap any of its four core stages—Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS). Understanding what models are supported by Hugging Face Speech-to-Speech is essential for building low-latency voice applications on CUDA, CPU, or Apple Silicon.

Voice Activity Detection (VAD) Models

The pipeline uses Silero VAD v5 as its sole built-in voice activity detection backend. According to the source code in src/speech_to_speech/VAD/vad_handler.py [lines 53-60], the handler wraps the Silero model and feeds start/stop timestamps to the pipeline queue. This backend runs on all platforms without additional dependencies.

Speech-to-Text (STT) Models

The STT layer supports six distinct backends, each implemented as a separate handler class inheriting from BaseSTTHandler in src/speech_to_speech/STT/base_stt_handler.py.

Parakeet TDT (Default)

Parakeet TDT serves as the default STT backend, optimized for both CUDA/CPU via nano-parakeet and Apple Silicon via MLX. The implementation lives in src/speech_to_speech/STT/parakeet_tdt_handler.py [lines 91-100], which handles audio chunking and model inference.

Whisper Variants

The pipeline supports multiple Whisper implementations:

Paraformer

Paraformer via FunASR supports CUDA/CPU and excels at Chinese speech recognition. The handler in src/speech_to_speech/STT/paraformer_handler.py [lines 22-30] requires the paraformer extra for installation.

Language Model (LLM) Backends

The LLM stage supports both API-based and local inference through handlers defined in src/speech_to_speech/LLM/language_model.py [lines 145-150].

OpenAI-Compatible APIs

The Responses API and Chat Completions handlers—src/speech_to_speech/LLM/responses_api_language_model.py and src/speech_to_speech/LLM/chat_completions_language_model.py—support any OpenAI-compatible endpoint including OpenAI, Hugging Face Inference Providers, OpenRouter, vLLM, and llama.cpp. These ship built-in and require only an API key or base URL.

Local Transformers and MLX

  • Transformers: Any text-generation model from the Hugging Face Hub runs locally via the built-in handler.
  • mlx-lm: Optimized Apple Silicon inference that ships built-in on macOS.

Text-to-Speech (TTS) Models

The TTS layer offers five backends, each implementing BaseHandler[TTSIn, TTSOut].

Qwen3-TTS (Default)

Qwen3-TTS serves as the default backend, automatically selecting GGML on Linux or mlx-audio on macOS. The handler in src/speech_to_speech/TTS/qwen3_tts_handler.py [lines 80-88] streams audio chunks back to the pipeline. Configuration arguments reside in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py [lines 5-12].

Kokoro-82M

Kokoro-82M runs on CUDA/CPU and Apple Silicon. The handler in src/speech_to_speech/TTS/kokoro_handler.py [lines 76-84] requires the kokoro extra on non-macOS systems but ships built-in on macOS.

Optional TTS Backends

Configuration Examples

Below are practical CLI configurations demonstrating different model combinations.

Default End-to-End Pipeline

Run the default stack (Parakeet TDT → OpenAI Responses API → Qwen3-TTS):

speech-to-speech \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-4o-mini"

Apple Silicon Local Deployment

Optimize for local Mac execution using MLX across all stages:

speech-to-speech \
    --local_mac_optimal_settings \
    --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16

This flag automatically selects Parakeet TDT, mlx-lm, and MLX Qwen3-TTS.

Custom STT and TTS Combination

Use Paraformer for Chinese STT with local MLX LLM and Qwen3-TTS:

speech-to-speech \
    --stt paraformer \
    --stt_model_name nvidia/parakeet-tdt-0.6b-v3 \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --language auto

Self-Hosted LLM with vLLM

Point to a local vLLM instance:

speech-to-speech \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream

Alternative TTS Backend

Configure Pocket TTS with specific voice and device settings:

speech-to-speech \
    --stt whisper-mlx \
    --llm_backend transformers \
    --tts pocket \
    --pocket_tts_voice jean \
    --pocket_tts_device cpu

Summary

  • VAD: Only Silero VAD v5 is supported, implemented in src/speech_to_speech/VAD/vad_handler.py.
  • STT: Six backends available including Parakeet TDT (default), four Whisper variants, and Paraformer, each with dedicated handlers in src/speech_to_speech/STT/.
  • LLM: Supports OpenAI-compatible APIs, Transformers, and mlx-lm through modular handlers in src/speech_to_speech/LLM/.
  • TTS: Five backends including Qwen3-TTS (default), Kokoro-82M, Pocket TTS, ChatTTS, and MMS TTS, implemented in src/speech_to_speech/TTS/.
  • Installation: Most models ship built-in; optional extras include faster-whisper, whisper-mlx, paraformer, kokoro, pocket, chattts, and facebook-mms.

Frequently Asked Questions

What is the default model configuration for Hugging Face Speech-to-Speech?

The default configuration uses Parakeet TDT for STT, OpenAI Responses API for LLM inference, and Qwen3-TTS for speech synthesis. This setup requires no additional installation beyond the base package and automatically optimizes for your hardware (GGML on Linux, mlx-audio on macOS).

Can I run Hugging Face Speech-to-Speech entirely with local models?

Yes. Use the --local_mac_optimal_settings flag on Apple Silicon to automatically select Parakeet TDT, mlx-lm, and MLX Qwen3-TTS. On CUDA/CPU systems, specify --llm_backend transformers with any Hugging Face text-generation model and --stt parakeet-tdt or --stt whisper for local STT.

Which STT model works best on Apple Silicon?

For Apple Silicon, you have three optimized options: Parakeet TDT with MLX (built-in), Lightning Whisper MLX (install via whisper-mlx extra), or MLX Audio Whisper (built-in on macOS). Parakeet TDT offers the best balance of speed and accuracy for English, while Lightning Whisper MLX excels at multilingual tasks.

How do I install optional models like Kokoro or ChatTTS?

Install optional TTS backends using pip extras: pip install speech-to-speech[kokoro] for Kokoro-82M, pip install speech-to-speech[chattts] for ChatTTS, and pip install speech-to-speech[pocket] for Pocket TTS. For STT extras, use pip install speech-to-speech[faster-whisper] or pip install speech-to-speech[paraformer].

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →