Using Speech-to-Speech for Real-Time Applications: Architecture and Setup Guide

The huggingface/speech-to-speech framework is expressly designed for real-time applications, delivering sub-second latency through a streaming pipeline that processes microphone input through Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS) components before playing synthesized audio back to the speaker.

The huggingface/speech-to-speech repository provides an open-source implementation of a real-time speech-to-speech pipeline. By leveraging queue-based threading and WebSocket streaming, this framework enables developers to build conversational AI applications that respond to voice input with minimal latency.

Core Architecture for Real-Time Processing

The pipeline achieves real-time performance through a modular, event-driven architecture where components communicate via thread-safe queues (Queue objects) and signaling events (threading.Event). This decouples producer and consumer rates to guarantee non-blocking streaming essential for low-latency interaction.

Audio I/O and Streaming Connections

Audio capture and playback rely on callback-driven streams that push raw PCM chunks into thread-safe queues. The implementation supports two primary connection modes:

Voice Activity Detection (VAD)

The src/speech_to_speech/VAD/vad_handler.py module detects when users start and stop speaking. It supports progressive transcription through the enable_realtime_transcription configuration parameter and realtime_processing_pause settings, allowing live rendering of partial hypotheses during conversation (see module_kwargs.enable_live_transcription).

Speech-to-Text (STT) Streaming

The get_stt_handler factory (defined in src/speech_to_speech/s2s_pipeline.py lines 801-880) instantiates handlers for multiple backends including whisper, mlx-audio-whisper, paraformer, and faster-whisper. These handlers emit conversation.item.input_audio_transcription.delta events for each partial hypothesis, enabling incremental text display before the user finishes speaking.

Language Model (LLM) Processing

The get_llm_handler function (lines 881-934) supports backends like responses-api, chat-completions, transformers, and mlx-lm. The LLM runs in its own thread and streams results via the LMOutputProcessor located in src/speech_to_speech/LLM/lm_output_processor.py, which provides a text_output_queue for real-time text events and speculative turn handling.

Text-to-Speech (TTS) Synthesis

The get_tts_handler factory (lines 938-1012) builds handlers for chatTTS, facebookMMS, pocket, kokoro, and qwen3. Audio chunks stream immediately through the same queue-based mechanism used by the I/O layer, allowing playback to begin before synthesis completes.

Realtime Server Implementation

When configured with mode == "realtime", the system instantiates src/speech_to_speech/api/openai_realtime/server.py, which exposes an OpenAI-compatible Realtime API over WebSocket. The _build_realtime_pipeline_unit function (called within build_pipeline at lines 610-779) creates isolated pipeline units with dedicated queues and events, handling session creation, turn detection, and event routing according to the official Realtime specification.

How to Enable Real-Time Mode

Configuring the framework for real-time speech-to-speech requires specific command-line arguments and backend selections:

  1. Set the runtime mode: Use --mode realtime to activate the WebSocket server and pipeline pool
  2. Configure live transcription: Enable --enable_live_transcription with --live_transcription_update_interval to control update frequency
  3. Select streaming STT: Choose mlx-audio-whisper or paraformer for optimized streaming inference
  4. Choose fast TTS: Select qwen3 or kokoro for sub-second audio synthesis

Running the Realtime Server and Client

To deploy the system, start the server and connect a client following the OpenAI Realtime specification.

Start the server with low-latency backends:

python -m src.speech_to_speech.api.openai_realtime.server \
    --mode realtime \
    --host 0.0.0.0 \
    --port 8765 \
    --stt mlx-audio-whisper \
    --tts qwen3 \
    --enable_live_transcription

Run the provided demo client that records from your microphone and plays synthesized responses:

python -m scripts.listen_and_play_realtime \
    --host 127.0.0.1 \
    --port 8765 \
    --model local \
    --send-rate 16000 \
    --recv-rate 16000 \
    --chunk-size 1024

The client in scripts/listen_and_play_realtime.py creates an AsyncOpenAI client, opens raw audio streams using sounddevice (RawInputStream / RawOutputStream), and handles server events including input_audio_buffer.append and response.output_audio.delta to stream audio as it arrives.

Optimizing Real-Time Performance

Maximize responsiveness for speech-to-speech for real-time applications with these hardware-specific configurations:

  • Use MLX on Apple Silicon: Set --llm_backend mlx-lm and --stt mlx-audio-whisper to leverage GPU acceleration and reduce inference latency
  • Enable macOS optimal settings: The local_mac_optimal_settings flag automatically switches devices to mps and selects the best-performing models (see optimal_mac_settings in s2s_pipeline.py)
  • Disable live transcription for multi-pipeline pools: On macOS, live transcription contends for the global MLX lock; disabling it prevents log flooding and maintains stable throughput (see the guard at line 1050 in s2s_pipeline.py)
  • Adjust realtime_processing_pause: Smaller values provide tighter turn-detection but increase CPU load; balance this parameter based on your hardware capabilities

Summary

  • The huggingface/speech-to-speech framework implements a queue-based, threaded architecture that decouples processing stages to achieve sub-second latency
  • Real-time mode activates via --mode realtime and exposes an OpenAI-compatible WebSocket endpoint at ws://<host>:<port>/v1
  • The pipeline supports progressive transcription through conversation.item.input_audio_transcription.delta events and immediate TTS streaming
  • Optimal performance requires selecting streaming-compatible backends like mlx-audio-whisper for STT and qwen3 for TTS
  • Production deployments on Apple Silicon should use MLX backends and consider disabling live transcription when running multiple pipeline units

Frequently Asked Questions

What latency can I expect when using speech-to-speech for real-time applications?

The framework achieves sub-second latency through its queue-based streaming architecture. By using optimized backends like mlx-audio-whisper for STT and qwen3 for TTS on Apple Silicon, the pipeline processes audio chunks incrementally rather than waiting for complete utterances, enabling responsive conversational flow.

Can I use the speech-to-speech framework without a GPU?

Yes, the framework supports CPU-only execution through backends like faster-whisper for STT and various TTS handlers. However, for real-time applications, GPU acceleration (particularly Apple Silicon with MLX or CUDA-compatible GPUs) is strongly recommended to maintain low latency during inference.

How does the framework handle multiple concurrent conversations?

When running in realtime mode, the system creates a pool of isolated pipeline units via _build_realtime_pipeline_unit in s2s_pipeline.py. Each unit maintains its own queues and events, allowing the server to handle multiple WebSocket connections simultaneously while preserving conversation state for each client.

Is the API compatible with OpenAI's Realtime API?

Yes, the src/speech_to_speech/api/openai_realtime/server.py implements the official OpenAI Realtime specification over WebSocket. Clients can connect to ws://<host>:<port>/v1 and use standard events like input_audio_buffer.append, conversation.item.input_audio_transcription.delta, and response.output_audio.delta, making it compatible with existing OpenAI Realtime client libraries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →