# Real-Time Speech-to-Speech Conversion: A Complete Implementation Guide

> Implement real-time speech-to-speech conversion with the complete pipeline from huggingface. This guide shows live audio transcription, response generation, and speech synthesis for ongoing conversations.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-07

---

**The huggingface/speech-to-speech repository ships a complete, end-to-end pipeline for real-time speech-to-speech conversion that captures live microphone audio, transcribes it, generates responses with a language model, synthesizes speech, and plays it back—all while the conversation is ongoing.**

The repository demonstrates real-time capabilities through an **OpenAI-compatible Realtime server** and a reference client implementation. Located in [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py), the example client streams audio to a local server that coordinates Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) handlers in real-time.

## Real-Time Architecture Overview

The real-time system follows a **client-server architecture** where [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) acts as the client and [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) hosts the WebSocket endpoint.

When the client connects, the repository spins up a **RealtimeServer** that owns a pool of pipeline units (one per concurrent conversation). Each unit contains the full processing chain:

```

VAD → STT → (optional live-transcription) → LLM → LM-output-processor → TTS → Audio output

```

The server receives WebSocket frames, forwards them to the appropriate **PipelineUnit**, and pushes generated events back to the client.

## Entry Point: listen_and_play_realtime.py

The [`listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/listen_and_play_realtime.py) script builds a client that talks to the local **OpenAI-compatible Realtime server** using `AsyncOpenAI(...).realtime.connect`. It streams audio to the server, receives transcription events and assistant audio, and plays the audio back through the default output device.

Key capabilities of this client:

- Handles microphone capture with a **SoundDevice** raw input stream
- Sends audio chunks via the Realtime WebSocket using `input_audio_buffer.append`
- Renders incremental transcription through `conversation.item.input_audio_transcription.delta`
- Plays synthesized audio received via `response.output_audio.delta`

## Realtime Server Implementation

The **RealtimeServer** in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) exposes the `/v1/realtime` WebSocket endpoint and manages the distribution of incoming connections among the pipeline pool.

For each connection, the server:

1. Instantiates a **PipelineUnit** through the `build_pipeline` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)
2. Configures **RealtimeService** ([`src/speech_to_speech/api/openai_realtime/service.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/service.py)) to supply the required `RuntimeConfig` containing chat settings and instructions
3. Streams events between the client and the processing handlers

## Pipeline Construction

The `build_pipeline` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) selects the appropriate mode (`realtime`, `local`, `websocket`, or socket-based). For real-time mode, it:

- Enables live transcription on the VAD handler via `vad_handler_kwargs.enable_realtime_transcription`
- Builds a **RealtimeService** that provides the `RuntimeConfig` (chat size, initial instructions, and OpenAI `RealtimeSessionCreateRequest`)
- Instantiates each handler using the `get_*_handler` factories

## Core Handler Components

The pipeline processes audio through specialized handlers located in the `src/speech_to_speech/` directory:

| Component | Role | Source File |
|-----------|------|-------------|
| **VAD** | Detects voice activity and optionally emits progressive transcription events | [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) |
| **STT** | Converts audio frames into text (supports Whisper, Faster-Whisper, Paraformer) | [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) |
| **LLM** | Generates assistant responses using OpenAI, Hugging Face transformers, or local MLX | [`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py) |
| **LM-output-processor** | Buffers partial audio, optionally streams text, and forwards to TTS | [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py) |
| **TTS** | Synthesizes audio from LLM output (supports Kokoro, Pocket, Facebook-MMS, Qwen-3) | [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) |

The connection layer in [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) handles low-level byte-level streaming for non-realtime WebSocket modes, while the real-time mode uses direct WebSocket communication through the OpenAI-compatible protocol.

## Running the Real-Time Demo

The repository provides a ready-to-run example requiring two terminals: one for the server and one for the client.

### Installation

Install the optional realtime dependencies:

```bash
pip install "speech-to-speech[realtime]"

```

### Starting the Server

Launch the local Realtime server (defaults to `127.0.0.1:8765`):

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --tts qwen3 \
    --stt whisper \
    --llm_backend responses-api \
    --host 127.0.0.1 --port 8765

```

### Launching the Client

In another terminal, run the client that captures microphone input and plays back audio:

```bash
python -m speech_to_speech.scripts.listen_and_play_realtime \
    --host 127.0.0.1 --port 8765 \
    --model local

```

The client prints live transcription (`USER:`) and assistant replies (`ASSISTANT:`) while playing the generated audio in real time.

## Minimal Programmatic Example

Below is a **pure-Python** snippet that mirrors the functionality of [`listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/listen_and_play_realtime.py), suitable for embedding in larger applications:

```python
import asyncio
import base64
from dataclasses import dataclass
from openai import AsyncOpenAI
import sounddevice as sd
from queue import Queue, Empty
from threading import Event, Lock

@dataclass
class Args:
    host: str = "127.0.0.1"
    port: int = 8765
    model: str = "local"
    api_key: str = "test-key"
    send_rate: int = 16000
    recv_rate: int = 16000
    chunk_size: int = 1024
    block_mic_during_playback: bool = False
    print_json: bool = False

args = Args()

def _make_client(a: Args) -> AsyncOpenAI:
    base = f"http://{a.host}:{a.port}/v1"
    ws = f"ws://{a.host}:{a.port}/v1"
    return AsyncOpenUI(a.api_key, base_url=base, websocket_base_url=ws)

def _session_update(a: Args) -> dict:
    def fmt(rate):
        return None if rate == 16000 else {"type": "audio/pcm", "rate": rate}
    return {
        "type": "session.update",
        "session": {
            "type": "realtime",
            "audio": {
                "input": {
                    "turn_detection": {"type": "server_vad", "interrupt_response": True},
                    **(fmt(a.send_rate) or {})
                },
                "output": {**(fmt(a.recv_rate) or {}), "voice": "marin"}
            }
        }
    }

async def realtime_loop(a: Args):
    client = _make_client(a)
    mic_q: Queue[bytes] = Queue(maxsize=128)
    stop = Event()
    playback = bytearray()
    lock = Lock()

    def mic_cb(indata, *_):
        if a.block_mic_during_playback:
            with lock:
                if playback:
                    return
        try:
            mic_q.put_nowait(bytes(indata))
        except Exception:
            pass

    def spk_cb(outdata, *_):
        needed = len(outdata)
        with lock:
            avail = min(needed, len(playback))
            outdata[:avail] = playback[:avail]
            del playback[:avail]
            if avail < needed:
                outdata[avail:] = b"\x00" * (needed - avail)

    in_stream = sd.RawInputStream(
        samplerate=a.send_rate, channels=1, dtype="int16",
        blocksize=a.chunk_size, callback=mic_cb
    )
    out_stream = sd.RawOutputStream(
        samplerate=a.recv_rate, channels=1, dtype="int16",
        blocksize=a.chunk_size, callback=spk_cb
    )

    in_stream.start()
    out_stream.start()
    
    async with client.realtime.connect(model=a.model) as conn:
        await conn.send(_session_update(a))
        
        async def send():
            while not stop.is_set():
                try:
                    chunk = await asyncio.to_thread(mic_q.get, True, 0.1)
                except Empty:
                    continue
                await conn.send({
                    "type": "input_audio_buffer.append",
                    "audio": base64.b64encode(chunk).decode()
                })
        
        async def recv():
            while not stop.is_set():
                ev = await conn.recv()
                if a.print_json:
                    print(ev.model_dump_json())
                if ev.type == "response.output_audio.delta":
                    data = base64.b64decode(ev.delta)
                    with lock:
                        playback.extend(data)
                elif ev.type == "response.done":
                    stop.set()
        
        await asyncio.gather(send(), recv())
    
    in_stream.stop()
    out_stream.stop()

if __name__ == "__main__":
    asyncio.run(realtime_loop(args))

```

> **Note:** The script relies only on the public `openai` package and the repository's internal pipeline; no extra binary components are needed beyond the optional `speech-to-speech[realtime]` extras.

## Summary

The huggingface/speech-to-speech repository provides a **modular, extensible real-time speech-to-speech stack**:

- **[`listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/listen_and_play_realtime.py)** provides a complete working client example that captures microphone input and plays synthesized responses
- **`RealtimeServer`** in [`server.py`](https://github.com/huggingface/speech-to-speech/blob/main/server.py) manages concurrent pipeline units via WebSocket connections at `/v1/realtime`
- The processing flow moves through **VAD → STT → LLM → LM-output-processor → TTS**, with each component swappable via the handler factory system
- Supports multiple backend implementations including **Whisper** for STT, **Qwen-3** for TTS, and various LLM backends (OpenAI, Hugging Face, MLX)
- Implements the **OpenAI Realtime API protocol**, making it compatible with existing OpenAI client libraries

## Frequently Asked Questions

### What is the latency of the real-time speech-to-speech pipeline?

The pipeline achieves low-latency conversion through streaming processing at each stage. The **VAD handler** uses `enable_realtime_transcription` to emit progressive transcription events before the audio segment completes, while the **LM-output-processor** buffers and streams partial audio chunks to the TTS handler as soon as the language model generates tokens. The WebSocket connection in [`websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_streamer.py) maintains persistent connections to minimize network overhead.

### Can I use custom local models instead of OpenAI for the LLM component?

Yes. The `build_pipeline` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) supports multiple LLM backends through the `--llm_backend` argument. You can specify `mlx` for Apple Silicon optimization, `transformers` for Hugging Face models, or `responses-api` for OpenAI compatibility. The [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py) handler abstracts these implementations behind a common interface.

### How does the repository handle voice activity detection in noisy environments?

The **VAD handler** in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py) includes configurable parameters for voice activity detection. When running in real-time mode, it processes raw audio streams to detect speech segments and can optionally emit transcription events progressively. The server-side turn detection (`server_vad` with `interrupt_response: True`) allows the system to handle barge-in and interruptions gracefully.

### Is it possible to stream audio over a network instead of localhost?

Yes. While the example uses `127.0.0.1`, you can specify any host address using the `--host` parameter when starting the server in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py). The client in [`listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/listen_and_play_realtime.py) accepts `--host` and `--port` arguments to connect to remote instances. The WebSocket protocol in [`server.py`](https://github.com/huggingface/speech-to-speech/blob/main/server.py) handles byte-level streaming over standard network connections, though latency will depend on network conditions.