Speech-to-Speech Model Architecture in the Hugging Face Speech-to-Speech Repository

The huggingface/speech-to-speech repository implements a modular, four-stage pipeline (VAD → STT → LLM → TTS) where each stage runs in its own thread and communicates via typed queues, enabling low-latency, pluggable voice assistants.

The speech-to-speech project provides a production-ready foundation for building real-time voice agents. Its architecture decouples voice activity detection, speech recognition, language generation, and speech synthesis into independent, swappable components. This design allows developers to mix open-source and proprietary models while maintaining thread-safe, low-latency audio streaming.

Four-Stage Pipeline Architecture

The core architecture follows a linear data flow with four sequential handlers. Each stage is implemented as a subclass of BaseHandler and operates in its own thread.

Stage Responsibility Default Backend Key File
Voice Activity Detection (VAD) Detects speech boundaries, splits audio into turns, emits turn metadata Silero VAD v5 src/speech_to_speech/VAD/vad_handler.py
Speech-to-Text (STT) Transcribes audio turns into text (with optional streaming) Parakeet TDT src/speech_to_speech/STT/parakeet_tdt_handler.py
Language Model (LLM) Generates assistant responses, optionally with tool calls OpenAI-compatible Responses API src/speech_to_speech/LLM/responses_api_model.py
Text-to-Speech (TTS) Synthesizes text back to audio and streams to client Qwen3-TTS src/speech_to_speech/TTS/qwen3_tts_handler.py

Data Flow Diagram


┌─────────────┐   audio   ┌───────┐   text   ┌───────┐   text   ┌───────┐
│  Audio      │ ───────► │ VAD   │ ─────► │ STT   │ ─────► │ LLM   │
│  Source     │          │       │          │       │          │
│ (mic/socket│          │       │          │       │          │
│  /file)     │          │       │          │       │          │
└─────────────┘          └───────┘          └───────┘          └───────┘
                                                                   │
                                                                   ▼
                                                            audio (synth)
                                                               │
                                                               ▼
                                                          ┌───────┐
                                                          │ TTS   │
                                                          └───────┘
                                                               │
                                                            output

Threading and Queue Communication

Inter-stage communication uses typed queues defined in src/speech_to_speech/pipeline/queue_types.py. Each handler reads from its input queue and writes to its output queue, enabling asynchronous, non-blocking data flow.

The build_pipeline function in src/speech_to_speech/s2s_pipeline.py (line 681) orchestrates this wiring:

  • Creates queue instances for each stage transition
  • Instantiates handler threads with appropriate kwargs
  • Returns a ThreadManager that starts and stops all threads

This design ensures that slow stages (like LLM generation) don't block audio capture or playback.

Realtime Mode and Pipeline Pooling

For production deployments, the repository provides realtime mode via RealtimeServer. This exposes an OpenAI Realtime-compatible WebSocket API at /v1/realtime.

Key implementation details from src/speech_to_speech/s2s_pipeline.py:

  • _build_realtime_pipeline_unit (line 480) creates isolated pipeline instances
  • Each WebSocket connection gets its own VAD→STT→LLM→TTS chain
  • The server pools these units to handle concurrent clients

# Start the realtime server with default backend stack

speech-to-speech \
    --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --enable_live_transcription

Turn-Taking and Speculative Interruptions

The VAD handler implements sophisticated turn management to minimize latency and handle interruptions gracefully.

Progressive Transcription

When enable_realtime_transcription is set, the VAD emits partial audio chunks before the turn completes. This allows the STT to begin transcription early, reducing perceived latency.

Speculative Turn Reopening

The speculative_turns feature (implemented in src/speech_to_speech/pipeline/speculative_turns.py) handles cases where users pause briefly then resume speaking:

  • SpeculativeTurnTracker monitors turn state
  • _should_reopen_current_turn evaluates whether to continue the current turn
  • _reopen_current_turn and _begin_pending_reopen_if_needed manage the transition

This avoids triggering a full LLM inference cycle for brief pauses, with configurable timeout via speculative_reopen_ms.

Backend Flexibility and Device Adaptation

Every stage supports multiple backend implementations. Selection happens at runtime through CLI flags parsed by ParsedArguments in s2s_pipeline.py.

STT Options

The get_stt_handler function (line 500) branches on the --stt flag:

  • parakeet-tdt — NVIDIA Parakeet TDT (default)
  • whisper — OpenAI Whisper
  • faster-whisper — Optimized Whisper implementation
  • lightning-whisper-mlx — Apple Silicon optimized
  • mlx-audio-whisper — MLX-native Whisper
  • paraformer — Alibaba Paraformer

LLM Options

get_llm_handler (line 549) supports:

  • responses-api — OpenAI-compatible Responses API (default)
  • transformers — Hugging Face Transformers
  • mlx-lm — Apple Silicon optimized inference
  • chat-completions — OpenAI Chat Completions compatibility

TTS Options

get_tts_handler (line 689) includes:

  • qwen3 — Qwen3-TTS (default)
  • kokoro — Kokoro TTS
  • pocket — Pocket TTS (lightweight)
  • chat-tts — ChatTTS
  • facebook-mms — Meta MMS

macOS and MLX Integration

The prepare_all_args helper automatically selects MLX backends when --device mps is detected. The --local_mac_optimal_settings flag applies tuned defaults for Apple Silicon.

Building a Custom Pipeline

The Python API provides fine-grained control over speech-to-speech model architecture composition:

from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse arguments programmatically

args = parse_arguments()

# Normalize device settings and backend kwargs

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    # ... additional handler kwargs

)

# Initialize communication queues

queues = initialize_queues_and_events()

# Construct the pipeline

pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.vad_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues,
)

pipeline_manager.start()
pipeline_manager.wait()

This mirrors the main() entry point but allows dynamic configuration without CLI invocation.

Swapping Backends: Practical Example

Replace the default STT with Faster-Whisper and use a local Transformers LLM:

speech-to-speech \
    --mode realtime \
    --stt faster-whisper \
    --faster_whisper_stt_model_name large-v2 \
    --llm_backend transformers \
    --model_name meta-llama/Meta-Llama-3.1-8B-Instruct \
    --tts qwen3 \
    --enable_live_transcription

The --stt faster-whisper branch in get_stt_handler instantiates FasterWhisperSTTHandler from src/speech_to_speech/STT/faster_whisper_handler.py.

Extending the Architecture

Adding new backends requires minimal changes:

  1. Implement a subclass of BaseHandler in the appropriate STT/, LLM/, or TTS/ directory
  2. Register it in the corresponding get_*_handler function
  3. Add argument dataclass in src/speech_to_speech/arguments_classes/
  4. Expose CLI flags through ParsedArguments

The arguments_classes/ pattern keeps configuration modular—each handler owns its parameter definitions without modifying core pipeline code.

Summary

  • The speech-to-speech model architecture uses four pluggable stages: VAD → STT → LLM → TTS
  • Each stage runs in its own thread with typed queue communication via queue_types.py
  • Realtime mode provides OpenAI-compatible WebSocket API with per-client pipeline isolation
  • Speculative turn reopening enables low-latency interruption handling without full pipeline resets
  • Multiple backends per stage (6 STT, 4 LLM, 5 TTS) support diverse hardware and latency requirements
  • Device adaptation automatically selects MLX backends on Apple Silicon

Frequently Asked Questions

What is the default speech-to-speech pipeline configuration?

The default stack uses Silero VAD v5 for voice detection, Parakeet TDT for speech-to-text, OpenAI Responses API for language modeling, and Qwen3-TTS for text-to-speech. This configuration prioritizes quality and low latency on GPU-equipped servers. Run with speech-to-speech --mode realtime to activate.

How does the pipeline handle user interruptions?

The speculative turn mechanism in VADHandler tracks provisional turn boundaries. If speech resumes within speculative_reopen_ms milliseconds, the turn continues without triggering a new LLM inference. This logic lives in _should_reopen_current_turn and uses SpeculativeTurnTracker from speculative_turns.py to manage state.

Can I run the pipeline on Apple Silicon Macs?

Yes. Set --device mps and optionally --local_mac_optimal_settings. The prepare_all_args helper automatically routes to MLX-compatible backends: mlx-lm for LLM inference, lightning-whisper-mlx or mlx-audio-whisper for STT, and Qwen3-TTS supports MPS directly.

How do I add a custom TTS model to the pipeline?

Implement a BaseHandler subclass in src/speech_to_speech/TTS/, add a factory branch in get_tts_handler (line 689), create an arguments dataclass in arguments_classes/, and register the CLI flag. The handler receives text via its input queue and should emit audio chunks to its output queue in the expected format.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →