Real-Time Speech-to-Speech Conversion: A Complete Implementation Guide

The huggingface/speech-to-speech repository ships a complete, end-to-end pipeline for real-time speech-to-speech conversion that captures live microphone audio, transcribes it, generates responses with a language model, synthesizes speech, and plays it back—all while the conversation is ongoing.

The repository demonstrates real-time capabilities through an OpenAI-compatible Realtime server and a reference client implementation. Located in scripts/listen_and_play_realtime.py, the example client streams audio to a local server that coordinates Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) handlers in real-time.

Real-Time Architecture Overview

The real-time system follows a client-server architecture where scripts/listen_and_play_realtime.py acts as the client and src/speech_to_speech/api/openai_realtime/server.py hosts the WebSocket endpoint.

When the client connects, the repository spins up a RealtimeServer that owns a pool of pipeline units (one per concurrent conversation). Each unit contains the full processing chain:


VAD → STT → (optional live-transcription) → LLM → LM-output-processor → TTS → Audio output

The server receives WebSocket frames, forwards them to the appropriate PipelineUnit, and pushes generated events back to the client.

Entry Point: listen_and_play_realtime.py

The listen_and_play_realtime.py script builds a client that talks to the local OpenAI-compatible Realtime server using AsyncOpenAI(...).realtime.connect. It streams audio to the server, receives transcription events and assistant audio, and plays the audio back through the default output device.

Key capabilities of this client:

  • Handles microphone capture with a SoundDevice raw input stream
  • Sends audio chunks via the Realtime WebSocket using input_audio_buffer.append
  • Renders incremental transcription through conversation.item.input_audio_transcription.delta
  • Plays synthesized audio received via response.output_audio.delta

Realtime Server Implementation

The RealtimeServer in src/speech_to_speech/api/openai_realtime/server.py exposes the /v1/realtime WebSocket endpoint and manages the distribution of incoming connections among the pipeline pool.

For each connection, the server:

  1. Instantiates a PipelineUnit through the build_pipeline function in src/speech_to_speech/s2s_pipeline.py
  2. Configures RealtimeService (src/speech_to_speech/api/openai_realtime/service.py) to supply the required RuntimeConfig containing chat settings and instructions
  3. Streams events between the client and the processing handlers

Pipeline Construction

The build_pipeline function in src/speech_to_speech/s2s_pipeline.py selects the appropriate mode (realtime, local, websocket, or socket-based). For real-time mode, it:

  • Enables live transcription on the VAD handler via vad_handler_kwargs.enable_realtime_transcription
  • Builds a RealtimeService that provides the RuntimeConfig (chat size, initial instructions, and OpenAI RealtimeSessionCreateRequest)
  • Instantiates each handler using the get_*_handler factories

Core Handler Components

The pipeline processes audio through specialized handlers located in the src/speech_to_speech/ directory:

Component Role Source File
VAD Detects voice activity and optionally emits progressive transcription events src/speech_to_speech/VAD/vad_handler.py
STT Converts audio frames into text (supports Whisper, Faster-Whisper, Paraformer) src/speech_to_speech/STT/whisper_stt_handler.py
LLM Generates assistant responses using OpenAI, Hugging Face transformers, or local MLX src/speech_to_speech/LLM/language_model.py
LM-output-processor Buffers partial audio, optionally streams text, and forwards to TTS src/speech_to_speech/LLM/lm_output_processor.py
TTS Synthesizes audio from LLM output (supports Kokoro, Pocket, Facebook-MMS, Qwen-3) src/speech_to_speech/TTS/qwen3_tts_handler.py

The connection layer in src/speech_to_speech/connections/websocket_streamer.py handles low-level byte-level streaming for non-realtime WebSocket modes, while the real-time mode uses direct WebSocket communication through the OpenAI-compatible protocol.

Running the Real-Time Demo

The repository provides a ready-to-run example requiring two terminals: one for the server and one for the client.

Installation

Install the optional realtime dependencies:

pip install "speech-to-speech[realtime]"

Starting the Server

Launch the local Realtime server (defaults to 127.0.0.1:8765):

python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --tts qwen3 \
    --stt whisper \
    --llm_backend responses-api \
    --host 127.0.0.1 --port 8765

Launching the Client

In another terminal, run the client that captures microphone input and plays back audio:

python -m speech_to_speech.scripts.listen_and_play_realtime \
    --host 127.0.0.1 --port 8765 \
    --model local

The client prints live transcription (USER:) and assistant replies (ASSISTANT:) while playing the generated audio in real time.

Minimal Programmatic Example

Below is a pure-Python snippet that mirrors the functionality of listen_and_play_realtime.py, suitable for embedding in larger applications:

import asyncio
import base64
from dataclasses import dataclass
from openai import AsyncOpenAI
import sounddevice as sd
from queue import Queue, Empty
from threading import Event, Lock

@dataclass
class Args:
    host: str = "127.0.0.1"
    port: int = 8765
    model: str = "local"
    api_key: str = "test-key"
    send_rate: int = 16000
    recv_rate: int = 16000
    chunk_size: int = 1024
    block_mic_during_playback: bool = False
    print_json: bool = False

args = Args()

def _make_client(a: Args) -> AsyncOpenAI:
    base = f"http://{a.host}:{a.port}/v1"
    ws = f"ws://{a.host}:{a.port}/v1"
    return AsyncOpenUI(a.api_key, base_url=base, websocket_base_url=ws)

def _session_update(a: Args) -> dict:
    def fmt(rate):
        return None if rate == 16000 else {"type": "audio/pcm", "rate": rate}
    return {
        "type": "session.update",
        "session": {
            "type": "realtime",
            "audio": {
                "input": {
                    "turn_detection": {"type": "server_vad", "interrupt_response": True},
                    **(fmt(a.send_rate) or {})
                },
                "output": {**(fmt(a.recv_rate) or {}), "voice": "marin"}
            }
        }
    }

async def realtime_loop(a: Args):
    client = _make_client(a)
    mic_q: Queue[bytes] = Queue(maxsize=128)
    stop = Event()
    playback = bytearray()
    lock = Lock()

    def mic_cb(indata, *_):
        if a.block_mic_during_playback:
            with lock:
                if playback:
                    return
        try:
            mic_q.put_nowait(bytes(indata))
        except Exception:
            pass

    def spk_cb(outdata, *_):
        needed = len(outdata)
        with lock:
            avail = min(needed, len(playback))
            outdata[:avail] = playback[:avail]
            del playback[:avail]
            if avail < needed:
                outdata[avail:] = b"\x00" * (needed - avail)

    in_stream = sd.RawInputStream(
        samplerate=a.send_rate, channels=1, dtype="int16",
        blocksize=a.chunk_size, callback=mic_cb
    )
    out_stream = sd.RawOutputStream(
        samplerate=a.recv_rate, channels=1, dtype="int16",
        blocksize=a.chunk_size, callback=spk_cb
    )

    in_stream.start()
    out_stream.start()
    
    async with client.realtime.connect(model=a.model) as conn:
        await conn.send(_session_update(a))
        
        async def send():
            while not stop.is_set():
                try:
                    chunk = await asyncio.to_thread(mic_q.get, True, 0.1)
                except Empty:
                    continue
                await conn.send({
                    "type": "input_audio_buffer.append",
                    "audio": base64.b64encode(chunk).decode()
                })
        
        async def recv():
            while not stop.is_set():
                ev = await conn.recv()
                if a.print_json:
                    print(ev.model_dump_json())
                if ev.type == "response.output_audio.delta":
                    data = base64.b64decode(ev.delta)
                    with lock:
                        playback.extend(data)
                elif ev.type == "response.done":
                    stop.set()
        
        await asyncio.gather(send(), recv())
    
    in_stream.stop()
    out_stream.stop()

if __name__ == "__main__":
    asyncio.run(realtime_loop(args))

Note: The script relies only on the public openai package and the repository's internal pipeline; no extra binary components are needed beyond the optional speech-to-speech[realtime] extras.

Summary

The huggingface/speech-to-speech repository provides a modular, extensible real-time speech-to-speech stack:

  • listen_and_play_realtime.py provides a complete working client example that captures microphone input and plays synthesized responses
  • RealtimeServer in server.py manages concurrent pipeline units via WebSocket connections at /v1/realtime
  • The processing flow moves through VAD → STT → LLM → LM-output-processor → TTS, with each component swappable via the handler factory system
  • Supports multiple backend implementations including Whisper for STT, Qwen-3 for TTS, and various LLM backends (OpenAI, Hugging Face, MLX)
  • Implements the OpenAI Realtime API protocol, making it compatible with existing OpenAI client libraries

Frequently Asked Questions

What is the latency of the real-time speech-to-speech pipeline?

The pipeline achieves low-latency conversion through streaming processing at each stage. The VAD handler uses enable_realtime_transcription to emit progressive transcription events before the audio segment completes, while the LM-output-processor buffers and streams partial audio chunks to the TTS handler as soon as the language model generates tokens. The WebSocket connection in websocket_streamer.py maintains persistent connections to minimize network overhead.

Can I use custom local models instead of OpenAI for the LLM component?

Yes. The build_pipeline function in s2s_pipeline.py supports multiple LLM backends through the --llm_backend argument. You can specify mlx for Apple Silicon optimization, transformers for Hugging Face models, or responses-api for OpenAI compatibility. The language_model.py handler abstracts these implementations behind a common interface.

How does the repository handle voice activity detection in noisy environments?

The VAD handler in vad_handler.py includes configurable parameters for voice activity detection. When running in real-time mode, it processes raw audio streams to detect speech segments and can optionally emit transcription events progressively. The server-side turn detection (server_vad with interrupt_response: True) allows the system to handle barge-in and interruptions gracefully.

Is it possible to stream audio over a network instead of localhost?

Yes. While the example uses 127.0.0.1, you can specify any host address using the --host parameter when starting the server in s2s_pipeline.py. The client in listen_and_play_realtime.py accepts --host and --port arguments to connect to remote instances. The WebSocket protocol in server.py handles byte-level streaming over standard network connections, though latency will depend on network conditions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →