# Using Speech-to-Speech for Real-Time Applications: Architecture and Setup Guide

> Explore speech-to-speech for real-time applications. This guide details the architecture and setup for sub-second latency streaming using VAD, STT, LLM, and TTS components.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**The huggingface/speech-to-speech framework is expressly designed for real-time applications, delivering sub-second latency through a streaming pipeline that processes microphone input through Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS) components before playing synthesized audio back to the speaker.**

The huggingface/speech-to-speech repository provides an open-source implementation of a real-time speech-to-speech pipeline. By leveraging queue-based threading and WebSocket streaming, this framework enables developers to build conversational AI applications that respond to voice input with minimal latency.

## Core Architecture for Real-Time Processing

The pipeline achieves real-time performance through a modular, event-driven architecture where components communicate via **thread-safe queues** (`Queue` objects) and **signaling events** (`threading.Event`). This decouples producer and consumer rates to guarantee non-blocking streaming essential for low-latency interaction.

### Audio I/O and Streaming Connections

Audio capture and playback rely on callback-driven streams that push raw PCM chunks into thread-safe queues. The implementation supports two primary connection modes:

- **Local mode**: Uses [`src/speech_to_speech/connections/local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/local_audio_streamer.py) for direct microphone and speaker access on the host machine
- **WebSocket mode**: Uses [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) for networked clients connecting over WebSocket

### Voice Activity Detection (VAD)

The [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) module detects when users start and stop speaking. It supports progressive transcription through the `enable_realtime_transcription` configuration parameter and `realtime_processing_pause` settings, allowing live rendering of partial hypotheses during conversation (see `module_kwargs.enable_live_transcription`).

### Speech-to-Text (STT) Streaming

The `get_stt_handler` factory (defined in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) lines 801-880) instantiates handlers for multiple backends including `whisper`, `mlx-audio-whisper`, `paraformer`, and `faster-whisper`. These handlers emit `conversation.item.input_audio_transcription.delta` events for each partial hypothesis, enabling incremental text display before the user finishes speaking.

### Language Model (LLM) Processing

The `get_llm_handler` function (lines 881-934) supports backends like `responses-api`, `chat-completions`, `transformers`, and `mlx-lm`. The LLM runs in its own thread and streams results via the `LMOutputProcessor` located in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py), which provides a `text_output_queue` for real-time text events and speculative turn handling.

### Text-to-Speech (TTS) Synthesis

The `get_tts_handler` factory (lines 938-1012) builds handlers for `chatTTS`, `facebookMMS`, `pocket`, `kokoro`, and `qwen3`. Audio chunks stream immediately through the same queue-based mechanism used by the I/O layer, allowing playback to begin before synthesis completes.

### Realtime Server Implementation

When configured with `mode == "realtime"`, the system instantiates [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py), which exposes an OpenAI-compatible Realtime API over WebSocket. The `_build_realtime_pipeline_unit` function (called within `build_pipeline` at lines 610-779) creates isolated pipeline units with dedicated queues and events, handling session creation, turn detection, and event routing according to the official Realtime specification.

## How to Enable Real-Time Mode

Configuring the framework for real-time speech-to-speech requires specific command-line arguments and backend selections:

1. **Set the runtime mode**: Use `--mode realtime` to activate the WebSocket server and pipeline pool
2. **Configure live transcription**: Enable `--enable_live_transcription` with `--live_transcription_update_interval` to control update frequency
3. **Select streaming STT**: Choose `mlx-audio-whisper` or `paraformer` for optimized streaming inference
4. **Choose fast TTS**: Select `qwen3` or `kokoro` for sub-second audio synthesis

## Running the Realtime Server and Client

To deploy the system, start the server and connect a client following the OpenAI Realtime specification.

Start the server with low-latency backends:

```bash
python -m src.speech_to_speech.api.openai_realtime.server \
    --mode realtime \
    --host 0.0.0.0 \
    --port 8765 \
    --stt mlx-audio-whisper \
    --tts qwen3 \
    --enable_live_transcription

```

Run the provided demo client that records from your microphone and plays synthesized responses:

```bash
python -m scripts.listen_and_play_realtime \
    --host 127.0.0.1 \
    --port 8765 \
    --model local \
    --send-rate 16000 \
    --recv-rate 16000 \
    --chunk-size 1024

```

The client in [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) creates an **AsyncOpenAI** client, opens raw audio streams using **sounddevice** (`RawInputStream` / `RawOutputStream`), and handles server events including `input_audio_buffer.append` and `response.output_audio.delta` to stream audio as it arrives.

## Optimizing Real-Time Performance

Maximize responsiveness for speech-to-speech for real-time applications with these hardware-specific configurations:

- **Use MLX on Apple Silicon**: Set `--llm_backend mlx-lm` and `--stt mlx-audio-whisper` to leverage GPU acceleration and reduce inference latency
- **Enable macOS optimal settings**: The `local_mac_optimal_settings` flag automatically switches devices to `mps` and selects the best-performing models (see `optimal_mac_settings` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py))
- **Disable live transcription for multi-pipeline pools**: On macOS, live transcription contends for the global MLX lock; disabling it prevents log flooding and maintains stable throughput (see the guard at line 1050 in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py))
- **Adjust `realtime_processing_pause`**: Smaller values provide tighter turn-detection but increase CPU load; balance this parameter based on your hardware capabilities

## Summary

- The huggingface/speech-to-speech framework implements a queue-based, threaded architecture that decouples processing stages to achieve sub-second latency
- Real-time mode activates via `--mode realtime` and exposes an OpenAI-compatible WebSocket endpoint at `ws://<host>:<port>/v1`
- The pipeline supports progressive transcription through `conversation.item.input_audio_transcription.delta` events and immediate TTS streaming
- Optimal performance requires selecting streaming-compatible backends like `mlx-audio-whisper` for STT and `qwen3` for TTS
- Production deployments on Apple Silicon should use MLX backends and consider disabling live transcription when running multiple pipeline units

## Frequently Asked Questions

### What latency can I expect when using speech-to-speech for real-time applications?

The framework achieves sub-second latency through its queue-based streaming architecture. By using optimized backends like `mlx-audio-whisper` for STT and `qwen3` for TTS on Apple Silicon, the pipeline processes audio chunks incrementally rather than waiting for complete utterances, enabling responsive conversational flow.

### Can I use the speech-to-speech framework without a GPU?

Yes, the framework supports CPU-only execution through backends like `faster-whisper` for STT and various TTS handlers. However, for real-time applications, GPU acceleration (particularly Apple Silicon with MLX or CUDA-compatible GPUs) is strongly recommended to maintain low latency during inference.

### How does the framework handle multiple concurrent conversations?

When running in `realtime` mode, the system creates a pool of isolated pipeline units via `_build_realtime_pipeline_unit` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py). Each unit maintains its own queues and events, allowing the server to handle multiple WebSocket connections simultaneously while preserving conversation state for each client.

### Is the API compatible with OpenAI's Realtime API?

Yes, the [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) implements the official OpenAI Realtime specification over WebSocket. Clients can connect to `ws://<host>:<port>/v1` and use standard events like `input_audio_buffer.append`, `conversation.item.input_audio_transcription.delta`, and `response.output_audio.delta`, making it compatible with existing OpenAI Realtime client libraries.