How to Use the speech-to-speech Python API: A Complete Guide

The speech-to-speech Python API exposes a modular voice-assistant pipeline through the s2s_pipeline.py module, allowing developers to programmatically configure VAD, STT, LLM, and TTS components using factory functions and the ThreadManager class.

The huggingface/speech-to-speech repository implements a fully modular voice-assistant architecture that can be invoked from Python code or the provided CLI. Unlike monolithic speech systems, this framework separates concerns into four interchangeable handlers that communicate through typed queues, enabling low-latency streaming and easy backend swapping. Whether building a local offline assistant or an OpenAI Realtime-compatible server, the speech-to-speech Python API provides granular control over every pipeline stage.

Architecture Overview

The system consists of four interchangeable components that run in parallel threads and exchange data through typed queues:

  • Voice-Activity Detection (VAD) – Detects speech boundaries using Silero VAD and optionally streams live transcription events.
  • Speech-to-Text (STT) – Converts spoken input into text using backends like Parakeet TDT, Whisper, Faster-Whisper, MLX-Audio-Whisper, or Paraformer.
  • Large Language Model (LLM) – Generates responses via local models (Transformers, mlx-lm) or OpenAI-compatible APIs (responses-api or chat-completions).
  • Text-to-Speech (TTS) – Synthesizes audio using Qwen3-TTS (default), Pocket TTS, ChatTTS, Kokoro-82M, or Facebook MMS.

These handlers communicate through typed queues (AudioInItem, STTOutItem, LMOutItem, TTSInItem) and synchronize using Event flags (stop_event, should_listen, response_playing). The orchestration logic in src/speech_to_speech/s2s_pipeline.py uses HfArgumentParser to parse CLI arguments or JSON configs, normalizes argument prefixes (e.g., --stt_*, --tts_*) via rename_args, and builds concrete handler instances through factory functions.

Transport Modes

The pipeline supports four transport modes controlled via the --mode argument:

Mode Transport Use Case
realtime OpenAI Realtime API over WebSocket/WebRTC Build voice assistants compatible with OpenAI clients
local Direct microphone & speakers Quick prototyping on a single machine
raw-websocket Raw PCM over WebSocket Custom clients streaming 16 kHz int16 PCM
socket Raw PCM over TCP Simple LAN streaming without Realtime features

Programmatic API Usage

To start the pipeline programmatically instead of via CLI, import the core functions from src/speech_to_speech/s2s_pipeline.py and manage the lifecycle through ThreadManager:

from speech_to_speech.s2s_pipeline import (
    parse_arguments,
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# Parse CLI arguments or construct dataclasses manually

args = parse_arguments()

# Normalize argument prefixes and apply device defaults

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Initialize shared queues and synchronization primitives

queues_and_events = initialize_queues_and_events()

# Build handlers via factory functions (get_stt_handler, get_llm_handler, get_tts_handler)

pipeline_manager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues_and_events,
)

# Start and block until SIGINT/SIGTERM

pipeline_manager.start()
pipeline_manager.wait()

The build_pipeline function internally uses factory functions (get_stt_handler, get_llm_handler, get_tts_handler) to instantiate concrete handlers based on your configuration, then wires them together under a ThreadManager that handles graceful shutdown.

CLI Configuration Examples

While the Python API provides full control, you can also leverage the CLI entry point which calls main() in s2s_pipeline.py:

Start a realtime server (default VAD → Parakeet TDT → Responses-API → Qwen3-TTS):

export OPENAI_API_KEY=sk-...
speech-to-speech

Swap STT to Whisper:

speech-to-speech --stt whisper --stt_model_name openai/whisper-base

Run locally with MLX LLM on Apple Silicon:

speech-to-speech --mode local --llm_backend mlx-lm --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16

Use Pocket TTS:

speech-to-speech --tts pocket --pocket_tts_voice jean --pocket_tts_device cpu

Enable the LLM proxy with --enable_llm_proxy to expose the configured LLM as a stand-alone OpenAI-compatible endpoint, letting other services call the model while the voice pipeline runs uninterrupted.

Key Source Files

Understanding the repository structure helps when extending the API:

  • src/speech_to_speech/s2s_pipeline.py – Core orchestration containing parse_arguments(), prepare_all_args(), initialize_queues_and_events(), and build_pipeline().
  • src/speech_to_speech/utils/thread_manager.py – Implements ThreadManager for handler lifecycle management.
  • src/speech_to_speech/arguments_classes/*.py – Dataclasses defining CLI/JSON arguments for each component.
  • src/speech_to_speech/STT/*_handler.py – Concrete STT implementations (Whisper, Parakeet TDT, etc.).
  • src/speech_to_speech/LLM/*_handler.py – LLM backends including local and API variants.
  • src/speech_to_speech/TTS/*_handler.py – TTS synthesizers (Qwen3-TTS, Kokoro, Pocket, etc.).
  • src/speech_to_speech/pipeline/*.py – Queue definitions and speculative turn tracking.

Summary

  • The speech-to-speech Python API in s2s_pipeline.py provides a modular interface for building voice assistants through four core handlers: VAD, STT, LLM, and TTS.
  • Factory functions (get_stt_handler, get_llm_handler, get_tts_handler) instantiate backends based on configuration dataclasses, while ThreadManager coordinates parallel execution.
  • Typed queues (AudioInItem, STTOutItem, LMOutItem, TTSInItem) and Event flags (stop_event, should_listen) enable safe thread communication and graceful shutdown.
  • The API supports four transport modes (realtime, local, raw-websocket, socket) and can expose an LLM proxy for external API access.
  • Configuration occurs through HfArgumentParser with prefixed arguments (--stt_*, --tts_*) normalized by prepare_all_args().

Frequently Asked Questions

How do I switch between different STT backends in the Python API?

Pass the specific handler kwargs to prepare_all_args() and build_pipeline(). The factory function get_stt_handler selects the concrete implementation based on the stt argument value (e.g., "whisper", "parakeet_tdt", "faster_whisper"). For example, to use Whisper, ensure args.whisper_stt_handler_kwargs contains model_name="openai/whisper-base" and the STT selector is set to "whisper".

What is the difference between realtime and local transport modes?

The realtime mode exposes an OpenAI Realtime API-compatible WebSocket endpoint at ws://localhost:8765/v1/realtime, allowing external clients to connect using standard OpenAI SDKs. The local mode bypasses network transport and connects directly to system microphone and speakers via the local audio handler, making it ideal for single-machine prototyping without network overhead.

How does the pipeline handle graceful shutdown?

The ThreadManager class (in src/speech_to_speech/utils/thread_manager.py) monitors a stop_event threading.Event shared across all handlers. When the process receives SIGINT or SIGTERM, the event is set, causing each handler's run() loop to exit cleanly. The main thread then calls join() on each handler thread via pipeline_manager.wait() to ensure all audio buffers are flushed before termination.

Can I use the speech-to-speech pipeline with a custom LLM server?

Yes. Set --enable_llm_proxy to expose the configured LLM as an OpenAI-compatible HTTP endpoint, or use --llm_backend with an OpenAI-compatible base URL. The LanguageModelHandler supports both responses-api and chat-completions formats, allowing you to proxy requests to custom servers while maintaining the voice pipeline's streaming capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →