How to Integrate Speech-to-Speech into a Python Project: A Complete Developer Guide

You can integrate the Hugging Face speech-to-speech pipeline into any Python project by importing the main() function from s2s_pipeline for CLI usage, or by programmatically constructing the four-stage pipeline (VAD → STT → LLM → TTS) using the build_pipeline() function with argument dataclasses.

The huggingface/speech-to-speech repository provides a modular, real-time voice conversation system that processes audio through four distinct stages. Whether you need a local microphone-to-speaker loop or a networked WebSocket service, you can embed this pipeline directly into existing Python applications without managing complex audio threading yourself.

Understanding the Four-Stage Pipeline Architecture

The repository implements a threaded pipeline where each stage communicates via queues and events managed by initialize_queues_and_events() in src/speech_to_speech/s2s_pipeline.py. Understanding these components helps you configure the integration correctly:

  • VAD (Voice Activity Detection): Detects when users start and stop speaking. Implemented in speech_to_speech/VAD/vad_handler.py and instantiated during pipeline construction.
  • STT (Speech-to-Text): Converts audio to text using handlers for Whisper, Faster-Whisper, Paraformer, or MLX-Audio-Whisper. Selected via module_kwargs.stt and created in s2s_pipeline.get_stt_handler.
  • LLM (Language Model): Generates text responses using Transformers, MLX-LM, or OpenAI-compatible APIs. Built via s2s_pipeline.get_llm_handler.
  • TTS (Text-to-Speech): Synthesizes audio from text using ChatTTS, Facebook MMS, Pocket, Kokoro, or Qwen-3 handlers. Created in s2s_pipeline.get_tts_handler.

All handlers run in separate threads coordinated by the ThreadManager utility in src/speech_to_speech/utils/thread_manager.py.

Method 1: CLI Integration (Fastest Setup)

For rapid prototyping or standalone deployment, invoke the pipeline directly from the command line. The entry point uses HfArgumentParser to process configuration from JSON files or command-line flags defined in src/speech_to_speech/arguments_classes/.

from speech_to_speech.s2s_pipeline import main

if __name__ == "__main__":
    # Reads CLI flags or JSON config automatically

    main()

Run directly via terminal:

python -m speech_to_speech.s2s_pipeline \
    --mode local \
    --stt whisper \
    --tts qwen3 \
    --llm_backend transformers \
    --device cpu \
    --log_level info

This approach handles all argument normalization, queue initialization, and thread management automatically through the prepare_all_args() function.

Method 2: Programmatic Integration (Embedded Usage)

To integrate speech-to-speech into an existing Python application, manually construct the pipeline using the lower-level API. This gives you fine-grained control over model selection, device placement, and communication modes.

from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelHandlerArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSHandlerArguments
from speech_to_speech.arguments_classes.socket_receiver_arguments import SocketReceiverArguments
from speech_to_speech.arguments_classes.socket_sender_arguments import SocketSenderArguments
from speech_to_speech.arguments_classes.websocket_streamer_arguments import WebSocketStreamerArguments
from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments
from speech_to_speech.s2s_pipeline import (
    prepare_all_args,
    initialize_queues_and_events,
    build_pipeline,
)

# 1. Configure module-level settings

module_args = ModuleArguments(
    mode="local",          # Options: "local", "websocket", "realtime"

    stt="whisper",
    tts="qwen3",
    llm_backend="transformers",
    device="cpu",          # or "cuda", "mps" for Apple Silicon

    log_level="info",
)

# 2. Configure individual handler arguments

whisper_args = WhisperSTTHandlerArguments(model_name="openai/whisper-base")
lm_args = LanguageModelHandlerArguments(model_name="Qwen/Qwen3-4B-Instruct-2507")
tts_args = Qwen3TTSHandlerArguments()

# 3. Normalize arguments (applies device optimizations and macOS-specific settings)

prepare_all_args(
    module_args,
    whisper_args,
    WhisperSTTHandlerArguments(),  # Placeholder for paraformer

    WhisperSTTHandlerArguments(),  # Placeholder for faster-whisper

    WhisperSTTHandlerArguments(),  # Placeholder for mlx-audio-whisper

    WhisperSTTHandlerArguments(),  # Placeholder for parakeet-tdt

    lm_args,
    LanguageModelHandlerArguments(),  # Placeholder for responses-api

    tts_args,
    Qwen3TTSHandlerArguments(),  # Placeholder for other TTS handlers

    Qwen3TTSHandlerArguments(),
    Qwen3TTSHandlerArguments(),
    Qwen3TTSHandlerArguments(),
)

# 4. Initialize inter-thread communication queues

queues = initialize_queues_and_events()

# 5. Build the pipeline with all handler configurations

pipeline_manager = build_pipeline(
    module_args,
    socket_receiver_kwargs=SocketReceiverArguments(),
    socket_sender_kwargs=SocketSenderArguments(),
    websocket_streamer_kwargs=WebSocketStreamerArguments(),
    vad_handler_kwargs=VADHandlerArguments(),
    whisper_stt_handler_kwargs=whisper_args,
    faster_whisper_stt_handler_kwargs=FasterWhisperSTTHandlerArguments(),
    paraformer_stt_handler_kwargs=ParaformerSTTHandlerArguments(),
    mlx_audio_whisper_stt_handler_kwargs=MLXAudioWhisperSTTHandlerArguments(),
    parakeet_tdt_stt_handler_kwargs=ParakeetTDTSTTHandlerArguments(),
    language_model_handler_kwargs=lm_args,
    responses_api_language_model_handler_kwargs=ResponsesApiLanguageModelHandlerArguments(),
    chat_tts_handler_kwargs=ChatTTSHandlerArguments(),
    facebook_mms_tts_handler_kwargs=FacebookMMSTTSHandlerArguments(),
    pocket_tts_handler_kwargs=PocketTTSHandlerArguments(),
    kokoro_tts_handler_kwargs=KokoroTTSHandlerArguments(),
    qwen3_tts_handler_kwargs=tts_args,
    queues_and_events=queues,
)

# 6. Start processing (blocks until interrupted)

pipeline_manager.start()
pipeline_manager.wait()

Key implementation details: The prepare_all_args() function in src/speech_to_speech/s2s_pipeline.py handles device detection and argument validation. On macOS, it automatically optimizes for MLX-based backends when available.

Method 3: Networked Deployment (WebSocket and Socket Modes)

For client-server architectures, change module_args.mode to "websocket" or use the socket-based sender/receiver arguments. The repository includes scripts/listen_and_play.py as a reference implementation for TCP audio streaming.

Start the pipeline in WebSocket mode:

python -m speech_to_speech.s2s_pipeline --mode websocket --stt whisper --tts qwen3

Then connect a client using the demo script:

python scripts/listen_and_play.py \
    --host localhost \
    --send_port 12345 \
    --recv_port 12346

The demo script captures raw 16-bit audio, transmits it to the pipeline, and plays back synthesized responses. This is useful for testing when direct microphone access isn't available or when building distributed voice applications.

Configuring Handler Arguments

All configurable parameters reside as dataclasses in src/speech_to_speech/arguments_classes/:

Pass these dataclass instances to build_pipeline() to override defaults without modifying source code.

Summary

  • Import the entry point: Use from speech_to_speech.s2s_pipeline import main for CLI-style integration that handles all setup automatically.
  • Construct manually: Call prepare_all_args(), initialize_queues_and_events(), and build_pipeline() to embed the four-stage pipeline (VAD → STT → LLM → TTS) directly into your application.
  • Choose your mode: Select "local" for direct audio hardware access, "websocket" for browser clients, or "realtime" for OpenAI-compatible API streaming.
  • Leverage argument dataclasses: All configuration lives in src/speech_to_speech/arguments_classes/ with specific handlers for Whisper, Qwen3, Transformers, and socket communication.
  • Enable low-latency: Set module_args.enable_live_transcription=True for real-time transcription feedback during processing.

Frequently Asked Questions

How do I reduce latency when integrating speech-to-speech into my Python application?

Set module_args.enable_live_transcription=True to enable live transcription mode, which streams partial results before the VAD detects speech completion. Additionally, specify device="mps" on Apple Silicon or device="cuda" on NVIDIA GPUs to leverage hardware acceleration in the handlers defined in src/speech_to_speech/LLM/ and src/speech_to_speech/TTS/.

Can I swap the default Whisper STT handler for Faster-Whisper or Paraformer?

Yes. Change module_args.stt to "faster_whisper" or "paraformer", then import and pass the corresponding argument dataclass (e.g., FasterWhisperSTTHandlerArguments or ParaformerSTTHandlerArguments) to the prepare_all_args() and build_pipeline() functions in your integration code.

What is the difference between "local", "websocket", and "realtime" modes in ModuleArguments?

Local mode reads from your system microphone and outputs to speakers directly. WebSocket mode accepts audio via WebSocket connections and returns synthesized audio over the same channel, suitable for web applications. Realtime mode uses the OpenAI Realtime API-compatible handler for cloud-based low-latency processing. Configure this in src/speech_to_speech/arguments_classes/module_arguments.py.

How do I gracefully shut down the pipeline when embedding it in a larger application?

The build_pipeline() function returns a ThreadManager instance (from src/speech_to_speech/utils/thread_manager.py) that registers signal handlers for graceful shutdown. Call pipeline_manager.stop() from your application's shutdown hook, or rely on the automatic Ctrl-C handling implemented in the main() function's try/finally block.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →