Using Speech-to-Speech for Real-Time Applications: Architecture and Setup Guide
The huggingface/speech-to-speech framework is expressly designed for real-time applications, delivering sub-second latency through a streaming pipeline that processes microphone input through Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS) components before playing synthesized audio back to the speaker.
The huggingface/speech-to-speech repository provides an open-source implementation of a real-time speech-to-speech pipeline. By leveraging queue-based threading and WebSocket streaming, this framework enables developers to build conversational AI applications that respond to voice input with minimal latency.
Core Architecture for Real-Time Processing
The pipeline achieves real-time performance through a modular, event-driven architecture where components communicate via thread-safe queues (Queue objects) and signaling events (threading.Event). This decouples producer and consumer rates to guarantee non-blocking streaming essential for low-latency interaction.
Audio I/O and Streaming Connections
Audio capture and playback rely on callback-driven streams that push raw PCM chunks into thread-safe queues. The implementation supports two primary connection modes:
- Local mode: Uses
src/speech_to_speech/connections/local_audio_streamer.pyfor direct microphone and speaker access on the host machine - WebSocket mode: Uses
src/speech_to_speech/connections/websocket_streamer.pyfor networked clients connecting over WebSocket
Voice Activity Detection (VAD)
The src/speech_to_speech/VAD/vad_handler.py module detects when users start and stop speaking. It supports progressive transcription through the enable_realtime_transcription configuration parameter and realtime_processing_pause settings, allowing live rendering of partial hypotheses during conversation (see module_kwargs.enable_live_transcription).
Speech-to-Text (STT) Streaming
The get_stt_handler factory (defined in src/speech_to_speech/s2s_pipeline.py lines 801-880) instantiates handlers for multiple backends including whisper, mlx-audio-whisper, paraformer, and faster-whisper. These handlers emit conversation.item.input_audio_transcription.delta events for each partial hypothesis, enabling incremental text display before the user finishes speaking.
Language Model (LLM) Processing
The get_llm_handler function (lines 881-934) supports backends like responses-api, chat-completions, transformers, and mlx-lm. The LLM runs in its own thread and streams results via the LMOutputProcessor located in src/speech_to_speech/LLM/lm_output_processor.py, which provides a text_output_queue for real-time text events and speculative turn handling.
Text-to-Speech (TTS) Synthesis
The get_tts_handler factory (lines 938-1012) builds handlers for chatTTS, facebookMMS, pocket, kokoro, and qwen3. Audio chunks stream immediately through the same queue-based mechanism used by the I/O layer, allowing playback to begin before synthesis completes.
Realtime Server Implementation
When configured with mode == "realtime", the system instantiates src/speech_to_speech/api/openai_realtime/server.py, which exposes an OpenAI-compatible Realtime API over WebSocket. The _build_realtime_pipeline_unit function (called within build_pipeline at lines 610-779) creates isolated pipeline units with dedicated queues and events, handling session creation, turn detection, and event routing according to the official Realtime specification.
How to Enable Real-Time Mode
Configuring the framework for real-time speech-to-speech requires specific command-line arguments and backend selections:
- Set the runtime mode: Use
--mode realtimeto activate the WebSocket server and pipeline pool - Configure live transcription: Enable
--enable_live_transcriptionwith--live_transcription_update_intervalto control update frequency - Select streaming STT: Choose
mlx-audio-whisperorparaformerfor optimized streaming inference - Choose fast TTS: Select
qwen3orkokorofor sub-second audio synthesis
Running the Realtime Server and Client
To deploy the system, start the server and connect a client following the OpenAI Realtime specification.
Start the server with low-latency backends:
python -m src.speech_to_speech.api.openai_realtime.server \
--mode realtime \
--host 0.0.0.0 \
--port 8765 \
--stt mlx-audio-whisper \
--tts qwen3 \
--enable_live_transcription
Run the provided demo client that records from your microphone and plays synthesized responses:
python -m scripts.listen_and_play_realtime \
--host 127.0.0.1 \
--port 8765 \
--model local \
--send-rate 16000 \
--recv-rate 16000 \
--chunk-size 1024
The client in scripts/listen_and_play_realtime.py creates an AsyncOpenAI client, opens raw audio streams using sounddevice (RawInputStream / RawOutputStream), and handles server events including input_audio_buffer.append and response.output_audio.delta to stream audio as it arrives.
Optimizing Real-Time Performance
Maximize responsiveness for speech-to-speech for real-time applications with these hardware-specific configurations:
- Use MLX on Apple Silicon: Set
--llm_backend mlx-lmand--stt mlx-audio-whisperto leverage GPU acceleration and reduce inference latency - Enable macOS optimal settings: The
local_mac_optimal_settingsflag automatically switches devices tompsand selects the best-performing models (seeoptimal_mac_settingsins2s_pipeline.py) - Disable live transcription for multi-pipeline pools: On macOS, live transcription contends for the global MLX lock; disabling it prevents log flooding and maintains stable throughput (see the guard at line 1050 in
s2s_pipeline.py) - Adjust
realtime_processing_pause: Smaller values provide tighter turn-detection but increase CPU load; balance this parameter based on your hardware capabilities
Summary
- The huggingface/speech-to-speech framework implements a queue-based, threaded architecture that decouples processing stages to achieve sub-second latency
- Real-time mode activates via
--mode realtimeand exposes an OpenAI-compatible WebSocket endpoint atws://<host>:<port>/v1 - The pipeline supports progressive transcription through
conversation.item.input_audio_transcription.deltaevents and immediate TTS streaming - Optimal performance requires selecting streaming-compatible backends like
mlx-audio-whisperfor STT andqwen3for TTS - Production deployments on Apple Silicon should use MLX backends and consider disabling live transcription when running multiple pipeline units
Frequently Asked Questions
What latency can I expect when using speech-to-speech for real-time applications?
The framework achieves sub-second latency through its queue-based streaming architecture. By using optimized backends like mlx-audio-whisper for STT and qwen3 for TTS on Apple Silicon, the pipeline processes audio chunks incrementally rather than waiting for complete utterances, enabling responsive conversational flow.
Can I use the speech-to-speech framework without a GPU?
Yes, the framework supports CPU-only execution through backends like faster-whisper for STT and various TTS handlers. However, for real-time applications, GPU acceleration (particularly Apple Silicon with MLX or CUDA-compatible GPUs) is strongly recommended to maintain low latency during inference.
How does the framework handle multiple concurrent conversations?
When running in realtime mode, the system creates a pool of isolated pipeline units via _build_realtime_pipeline_unit in s2s_pipeline.py. Each unit maintains its own queues and events, allowing the server to handle multiple WebSocket connections simultaneously while preserving conversation state for each client.
Is the API compatible with OpenAI's Realtime API?
Yes, the src/speech_to_speech/api/openai_realtime/server.py implements the official OpenAI Realtime specification over WebSocket. Clients can connect to ws://<host>:<port>/v1 and use standard events like input_audio_buffer.append, conversation.item.input_audio_transcription.delta, and response.output_audio.delta, making it compatible with existing OpenAI Realtime client libraries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →