How to Integrate Speech-to-Speech into a Python Application: A Complete Guide

You can integrate speech-to-speech into a Python application by using the huggingface/speech-to-speech pipeline's main() entry point for CLI usage, or by programmatically constructing the four-stage pipeline (VAD → STT → LLM → TTS) using the build_pipeline() function with argument dataclasses.

The huggingface/speech-to-speech repository provides a modular, production-ready framework for building real-time voice conversations. It chains together Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) into a seamless pipeline that runs across multiple threads. This guide shows you how to integrate this system into your Python applications using both command-line and programmatic approaches.

Understand the Four-Stage Pipeline Architecture

The pipeline processes audio through four distinct stages connected by queues and events:

  1. VAD (Voice Activity Detection) - Detects when users start/stop speaking. Implemented in speech_to_speech/VAD/vad_handler.py and instantiated in s2s_pipeline._build_pipeline_handlers.

  2. STT (Speech-to-Text) - Converts audio to text using Whisper, Faster-Whisper, Paraformer, or MLX-Audio-Whisper. Selected via module_kwargs.stt and created in s2s_pipeline.get_stt_handler.

  3. LLM (Language Model) - Generates responses using Transformers, MLX-LM, or OpenAI-compatible APIs. Built in s2s_pipeline.get_llm_handler.

  4. TTS (Text-to-Speech) - Synthesizes speech via ChatTTS, Facebook MMS, Pocket, Kokoro, or Qwen-3. Created in s2s_pipeline.get_tts_handler.

All stages communicate via thread-safe queues initialized through initialize_queues_and_events() and managed by the ThreadManager in src/speech_to_speech/utils/thread_manager.py.

Method 1: Quick Integration via Command Line

For rapid prototyping, use the built-in CLI entry point.

This approach handles argument parsing via HfArgumentParser, normalization through prepare_all_args, and pipeline construction automatically.

from speech_to_speech.s2s_pipeline import main

if __name__ == "__main__":
    main()

Run with command-line flags:

python -m speech_to_speech.s2s_pipeline \
    --mode local \
    --stt whisper \
    --tts qwen3 \
    --llm_backend transformers \
    --device cpu \
    --log_level info

The main() function in s2s_pipeline.py orchestrates the entire flow: parsing arguments, building handlers, and starting the ThreadManager.

Method 2: Programmatic Integration

For embedding within existing applications, construct the pipeline manually using dataclasses from src/speech_to_speech/arguments_classes/.

Step 1: Configure Arguments

Import and instantiate the specific argument classes for your chosen backends:

from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.language_model_arguments import LanguageModelHandlerArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSHandlerArguments

module_args = ModuleArguments(
    mode="local",          # Options: "local", "websocket", "realtime"

    stt="whisper",
    tts="qwen3",
    llm_backend="transformers",
    device="cpu",
    log_level="info",
)

whisper_args = WhisperSTTHandlerArguments(model_name="openai/whisper-base")
lm_args = LanguageModelHandlerArguments(model_name="Qwen/Qwen3-4B-Instruct-2507")
tts_args = Qwen3TTSHandlerArguments()

Step 2: Normalize Arguments

Call prepare_all_args() to apply device optimizations and format gen_kwargs:

from speech_to_speech.s2s_pipeline import prepare_all_args

prepare_all_args(
    module_args,
    whisper_args,
    WhisperSTTHandlerArguments(),  # Placeholder for paraformer

    WhisperSTTHandlerArguments(),  # Placeholder for faster-whisper

    WhisperSTTHandlerArguments(),  # Placeholder for mlx-audio-whisper

    WhisperSTTHandlerArguments(),  # Placeholder for parakeet-tdt

    lm_args,
    LanguageModelHandlerArguments(),  # Placeholder for responses-api

    tts_args,
    Qwen3TTSHandlerArguments(),  # Placeholder for other TTS handlers

    Qwen3TTSHandlerArguments(),
    Qwen3TTSHandlerArguments(),
    Qwen3TTSHandlerArguments(),
)

Step 3: Initialize Communication Primitives

Create the queues and events that connect pipeline stages:

from speech_to_speech.s2s_pipeline import initialize_queues_and_events

queues = initialize_queues_and_events()

Step 4: Build and Run the Pipeline

Construct the pipeline with build_pipeline() and start execution:

from speech_to_speech.s2s_pipeline import build_pipeline
from speech_to_speech.arguments_classes.socket_receiver_arguments import SocketReceiverArguments
from speech_to_speech.arguments_classes.socket_sender_arguments import SocketSenderArguments
from speech_to_speech.arguments_classes.websocket_streamer_arguments import WebSocketStreamerArguments
from speech_to_speech.arguments_classes.vad_handler_arguments import VADHandlerArguments
from speech_to_speech.arguments_classes.faster_whisper_stt_arguments import FasterWhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.paraformer_stt_arguments import ParaformerSTTHandlerArguments
from speech_to_speech.arguments_classes.mlx_audio_whisper_stt_arguments import MLXAudioWhisperSTTHandlerArguments
from speech_to_speech.arguments_classes.parakeet_tdt_stt_arguments import ParakeetTDTSTTHandlerArguments
from speech_to_speech.arguments_classes.responses_api_language_model_arguments import ResponsesApiLanguageModelHandlerArguments
from speech_to_speech.arguments_classes.chat_tts_arguments import ChatTTSHandlerArguments
from speech_to_speech.arguments_classes.facebook_mms_tts_arguments import FacebookMMSTTSHandlerArguments
from speech_to_speech.arguments_classes.pocket_tts_arguments import PocketTTSHandlerArguments
from speech_to_speech.arguments_classes.kokoro_tts_arguments import KokoroTTSHandlerArguments

pipeline_manager = build_pipeline(
    module_args,
    SocketReceiverArguments(),
    SocketSenderArguments(),
    WebSocketStreamerArguments(),
    VADHandlerArguments(),
    whisper_args,
    FasterWhisperSTTHandlerArguments(),
    ParaformerSTTHandlerArguments(),
    MLXAudioWhisperSTTHandlerArguments(),
    ParakeetTDTSTTHandlerArguments(),
    lm_args,
    ResponsesApiLanguageModelHandlerArguments(),
    ChatTTSHandlerArguments(),
    FacebookMMSTTSHandlerArguments(),
    PocketTTSHandlerArguments(),
    KokoroTTSHandlerArguments(),
    tts_args,
    queues,
)

pipeline_manager.start()
pipeline_manager.wait()  # Blocks until interrupted

Configure Audio I/O Modes

The pipeline supports three operational modes controlled by module_args.mode:

  • Local mode ("local"): Direct microphone input and speaker output. Best for standalone applications.
  • WebSocket mode ("websocket"): Audio streams over WebSocket connections. Ideal for web applications.
  • Real-time mode ("realtime"): OpenAI-compatible realtime API integration.

For low-latency requirements, enable live transcription with module_args.enable_live_transcription=True.

Test with the Socket Demo

The repository includes scripts/listen_and_play.py for testing TCP socket streaming without hardware loopback.

Start the pipeline in WebSocket mode:

python -m speech_to_speech.s2s_pipeline --mode websocket

Then run the demo client:

python scripts/listen_and_play.py \
    --host localhost \
    --send_port 12345 \
    --recv_port 12346

The demo captures raw 16-bit audio, pushes it to the pipeline via TCP, and plays back synthesized responses.

Summary

  • Four-stage architecture: The pipeline chains VAD → STT → LLM → TTS through thread-safe queues managed by ThreadManager.
  • Two integration paths: Use main() for CLI-driven deployment or build_pipeline() for embedded applications.
  • Flexible backends: Choose from Whisper, Faster-Whisper, or Paraformer for STT; Qwen-3, ChatTTS, or Kokoro for TTS; and Transformers or MLX-LM for language modeling.
  • Multiple modes: Run locally with direct audio hardware, over WebSockets for web apps, or via the OpenAI-compatible realtime API.
  • Configuration via dataclasses: All parameters are type-safe arguments defined in src/speech_to_speech/arguments_classes/.

Frequently Asked Questions

How do I select different STT or TTS backends?

Set the stt and tts parameters in ModuleArguments to your desired backend identifiers (e.g., "whisper", "faster-whisper", "qwen3", "kokoro"), then provide the corresponding handler arguments to prepare_all_args(). The build_pipeline() function instantiates the correct handler classes based on these selections.

Can I run the pipeline on macOS with Apple Silicon?

Yes. The prepare_all_args() function automatically detects macOS and applies optimizations, preferring mlx-lm for the LLM backend and qwen3 for TTS when available. Specify device="mps" or leave it on auto-detect for best performance on Apple Silicon.

What is the difference between local mode and WebSocket mode?

Local mode (mode="local") uses direct system audio through local audio streamers, suitable for desktop applications. WebSocket mode (mode="websocket") accepts audio over WebSocket connections via WebSocketStreamerArguments, enabling remote clients to stream audio to the pipeline over the network.

How do I enable graceful shutdown in my application?

The ThreadManager returned by build_pipeline() registers signal handlers for graceful shutdown. Call pipeline_manager.stop() or send a KeyboardInterrupt (Ctrl+C) to trigger the shutdown sequence, which properly joins all handler threads and clears the inter-thread queues.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →