Speech-to-Speech Tasks Supported by the Hugging Face speech-to-speech Repository

The huggingface/speech-to-speech repository supports four modular speech-to-speech tasks—Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM) generation, and Text-to-Speech (TTS)—each with pluggable backends ranging from Silero VAD to Qwen3-TTS that can be mixed and matched via CLI arguments.

The huggingface/speech-to-speech package implements a production-ready pipeline for end-to-end voice conversations, decomposing speech-to-speech tasks into four independent stages that communicate via async queues. According to the source code in src/speech_to_speech/s2s_pipeline.py, each task runs in its own thread and can be swapped without modifying core pipeline logic, enabling deployment across CUDA, CPU, and Apple Silicon hardware.

The Four Core Speech-to-Speech Tasks

The repository organizes functionality into four distinct handlers, each defined in dedicated subdirectories under src/speech_to_speech/.

Voice Activity Detection (VAD)

VAD detects when a user starts and stops speaking, providing turn-taking information to downstream components. The default implementation uses Silero VAD v5 via the handler defined in src/speech_to_speech/VAD/vad_handler.py. This handler processes raw microphone streams and emits SpeechStartedEvent and SpeechStoppedEvent objects, forwarding audio chunks as VADAudio items to the STT stage.

Speech-to-Text (STT)

STT converts spoken audio into textual transcripts, optionally streaming partial results for real-time displays. The default backend is Parakeet TDT 0.6B v3, implemented in src/speech_to_speech/STT/whisper_stt_handler.py. Alternative backends include:

  • Whisper (via Transformers)
  • Faster Whisper
  • Lightning Whisper MLX (Apple Silicon)
  • MLX Audio Whisper (Apple Silicon)
  • Paraformer (FunASR)

The STT handler pushes transcribed output as a TextEventItem onto the LLM queue.

Language Model (LLM)

LLM generates the assistant’s response text, optionally streaming tool-call events. The default backend uses the OpenAI-compatible Responses API, with handler logic referenced in src/speech_to_speech/pipeline/handler_types.py. Alternative backends include:

  • Transformers (CUDA/CPU)
  • mlx-lm (Apple Silicon)
  • Self-hosted servers via Responses API or Chat Completions (e.g., vLLM, llama.cpp)

This stage consumes TextEventItem objects from the STT stage and produces new TextEventItem instances for the TTS stage.

Text-to-Speech (TTS)

TTS synthesizes the LLM’s textual output into audio waveforms streamed back to the client. The default is Qwen3-TTS, using GGML on Linux and mlx-audio on macOS, implemented in src/speech_to_speech/TTS/qwen3_tts_handler.py. Alternative backends include:

  • Kokoro-82M (CUDA/CPU, Apple Silicon)
  • Pocket TTS (CPU/CUDA)
  • ChatTTS (CUDA/CPU)
  • MMS TTS (Transformers)

Task Pipeline Architecture

The four speech-to-speech tasks execute as a cascade of independent handlers linked by async queues. As implemented in s2s_pipeline.py, the flow is:

  1. VAD processes raw audio, detects speech boundaries, and forwards VADAudio chunks.
  2. STT receives chunks, transcribes them, and pushes a TextEventItem to the LLM queue.
  3. LLM consumes the transcript, generates a response, and pushes a TextEventItem to the TTS queue.
  4. TTS synthesizes the final audio output for delivery via WebSocket or local playback.

Each handler inherits from the base Handler class defined in baseHandler.py, allowing seamless backend substitution through CLI flags like --stt, --llm_backend, and --tts.

Execution Modes for Speech-to-Speech Tasks

The repository supports multiple runtime configurations defined in src/speech_to_speech/api/openai_realtime/websocket_router.py:

  • Realtime mode (--mode realtime): Uses the OpenAI Realtime protocol over WebSocket/WebRTC, enabling live transcription (--enable_live_transcription) and speculative turn handling.
  • Local mode (--mode local): Runs the entire pipeline on a single machine without the Realtime protocol wrapper.
  • Raw-WebSocket (--mode raw-websocket): Streams raw PCM audio without protocol overhead.
  • Socket (--mode socket): Minimal TCP-socket streaming for embedded applications.

Configuring Speech-to-Speech Tasks: Code Examples

Running the Default Realtime Server

Install the core package and start the server:

pip install speech-to-speech

# Starts server on ws://localhost:8765/v1/realtime

speech-to-speech

Switching to Faster Whisper for STT

Select an alternative STT backend using the --stt flag, which loads the handler from src/speech_to_speech/STT/faster_whisper_handler.py:

speech-to-speech \
    --stt faster-whisper \
    --stt_model_name base \
    --mode realtime

Using Local LLM Inference with mlx-lm

Deploy Apple Silicon-optimized LLM inference via the --llm_backend flag:

speech-to-speech \
    --llm_backend mlx-lm \
    --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16 \
    --mode realtime

The mlx-lm backend is defined in src/speech_to_speech/pipeline/handler_types.py.

Configuring Pocket TTS for Speech Synthesis

Switch to the Pocket TTS backend implemented in src/speech_to_speech/TTS/pocket_tts_handler.py:

speech-to-speech \
    --tts pocket \
    --pocket_tts_voice jean \
    --pocket_tts_device cpu \
    --mode realtime

Summary

  • The huggingface/speech-to-speech repository implements four modular speech-to-speech tasks: VAD, STT, LLM, and TTS.
  • Each task supports multiple interchangeable backends selectable via CLI flags, including Silero VAD, Parakeet/Whisper STT, OpenAI/Transformers/MLX LLMs, and Qwen3/Kokoro TTS.
  • Tasks communicate via async queues in a multi-threaded pipeline orchestrated by s2s_pipeline.py.
  • The system supports four execution modes: realtime (OpenAI protocol), local, raw-websocket, and socket.

Frequently Asked Questions

What is the default STT backend in the speech-to-speech repository?

The default Speech-to-Text backend is Parakeet TDT 0.6B v3, implemented in src/speech_to_speech/STT/whisper_stt_handler.py. The system also supports Whisper, Faster Whisper, Lightning Whisper MLX, and Paraformer via the --stt CLI flag.

Can I run speech-to-speech tasks entirely on Apple Silicon without CUDA?

Yes. The repository supports Apple Silicon-optimized backends for all four tasks: Silero VAD for voice detection, Lightning Whisper MLX or MLX Audio Whisper for transcription, mlx-lm for language modeling, and mlx-audio for Qwen3-TTS synthesis.

How do I enable live transcription streaming during conversations?

Pass the --enable_live_transcription flag when starting in realtime mode (--mode realtime). This configures the STT handler to emit partial transcripts while speech is still ongoing, rather than waiting for the VAD to detect speech completion.

Where are the CLI arguments for each speech-to-speech task defined?

Argument definitions for VAD, STT, LLM, and TTS handlers are located in src/speech_to_speech/arguments_classes/, with files such as whisper_stt_arguments.py defining the parameters for specific backends.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →