Real-Time Speech-to-Speech Conversion: A Complete Implementation Guide
The huggingface/speech-to-speech repository ships a complete, end-to-end pipeline for real-time speech-to-speech conversion that captures live microphone audio, transcribes it, generates responses with a language model, synthesizes speech, and plays it back—all while the conversation is ongoing.
The repository demonstrates real-time capabilities through an OpenAI-compatible Realtime server and a reference client implementation. Located in scripts/listen_and_play_realtime.py, the example client streams audio to a local server that coordinates Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Models (LLM), and Text-to-Speech (TTS) handlers in real-time.
Real-Time Architecture Overview
The real-time system follows a client-server architecture where scripts/listen_and_play_realtime.py acts as the client and src/speech_to_speech/api/openai_realtime/server.py hosts the WebSocket endpoint.
When the client connects, the repository spins up a RealtimeServer that owns a pool of pipeline units (one per concurrent conversation). Each unit contains the full processing chain:
VAD → STT → (optional live-transcription) → LLM → LM-output-processor → TTS → Audio output
The server receives WebSocket frames, forwards them to the appropriate PipelineUnit, and pushes generated events back to the client.
Entry Point: listen_and_play_realtime.py
The listen_and_play_realtime.py script builds a client that talks to the local OpenAI-compatible Realtime server using AsyncOpenAI(...).realtime.connect. It streams audio to the server, receives transcription events and assistant audio, and plays the audio back through the default output device.
Key capabilities of this client:
- Handles microphone capture with a SoundDevice raw input stream
- Sends audio chunks via the Realtime WebSocket using
input_audio_buffer.append - Renders incremental transcription through
conversation.item.input_audio_transcription.delta - Plays synthesized audio received via
response.output_audio.delta
Realtime Server Implementation
The RealtimeServer in src/speech_to_speech/api/openai_realtime/server.py exposes the /v1/realtime WebSocket endpoint and manages the distribution of incoming connections among the pipeline pool.
For each connection, the server:
- Instantiates a PipelineUnit through the
build_pipelinefunction insrc/speech_to_speech/s2s_pipeline.py - Configures RealtimeService (
src/speech_to_speech/api/openai_realtime/service.py) to supply the requiredRuntimeConfigcontaining chat settings and instructions - Streams events between the client and the processing handlers
Pipeline Construction
The build_pipeline function in src/speech_to_speech/s2s_pipeline.py selects the appropriate mode (realtime, local, websocket, or socket-based). For real-time mode, it:
- Enables live transcription on the VAD handler via
vad_handler_kwargs.enable_realtime_transcription - Builds a RealtimeService that provides the
RuntimeConfig(chat size, initial instructions, and OpenAIRealtimeSessionCreateRequest) - Instantiates each handler using the
get_*_handlerfactories
Core Handler Components
The pipeline processes audio through specialized handlers located in the src/speech_to_speech/ directory:
| Component | Role | Source File |
|---|---|---|
| VAD | Detects voice activity and optionally emits progressive transcription events | src/speech_to_speech/VAD/vad_handler.py |
| STT | Converts audio frames into text (supports Whisper, Faster-Whisper, Paraformer) | src/speech_to_speech/STT/whisper_stt_handler.py |
| LLM | Generates assistant responses using OpenAI, Hugging Face transformers, or local MLX | src/speech_to_speech/LLM/language_model.py |
| LM-output-processor | Buffers partial audio, optionally streams text, and forwards to TTS | src/speech_to_speech/LLM/lm_output_processor.py |
| TTS | Synthesizes audio from LLM output (supports Kokoro, Pocket, Facebook-MMS, Qwen-3) | src/speech_to_speech/TTS/qwen3_tts_handler.py |
The connection layer in src/speech_to_speech/connections/websocket_streamer.py handles low-level byte-level streaming for non-realtime WebSocket modes, while the real-time mode uses direct WebSocket communication through the OpenAI-compatible protocol.
Running the Real-Time Demo
The repository provides a ready-to-run example requiring two terminals: one for the server and one for the client.
Installation
Install the optional realtime dependencies:
pip install "speech-to-speech[realtime]"
Starting the Server
Launch the local Realtime server (defaults to 127.0.0.1:8765):
python -m speech_to_speech.s2s_pipeline \
--mode realtime \
--tts qwen3 \
--stt whisper \
--llm_backend responses-api \
--host 127.0.0.1 --port 8765
Launching the Client
In another terminal, run the client that captures microphone input and plays back audio:
python -m speech_to_speech.scripts.listen_and_play_realtime \
--host 127.0.0.1 --port 8765 \
--model local
The client prints live transcription (USER:) and assistant replies (ASSISTANT:) while playing the generated audio in real time.
Minimal Programmatic Example
Below is a pure-Python snippet that mirrors the functionality of listen_and_play_realtime.py, suitable for embedding in larger applications:
import asyncio
import base64
from dataclasses import dataclass
from openai import AsyncOpenAI
import sounddevice as sd
from queue import Queue, Empty
from threading import Event, Lock
@dataclass
class Args:
host: str = "127.0.0.1"
port: int = 8765
model: str = "local"
api_key: str = "test-key"
send_rate: int = 16000
recv_rate: int = 16000
chunk_size: int = 1024
block_mic_during_playback: bool = False
print_json: bool = False
args = Args()
def _make_client(a: Args) -> AsyncOpenAI:
base = f"http://{a.host}:{a.port}/v1"
ws = f"ws://{a.host}:{a.port}/v1"
return AsyncOpenUI(a.api_key, base_url=base, websocket_base_url=ws)
def _session_update(a: Args) -> dict:
def fmt(rate):
return None if rate == 16000 else {"type": "audio/pcm", "rate": rate}
return {
"type": "session.update",
"session": {
"type": "realtime",
"audio": {
"input": {
"turn_detection": {"type": "server_vad", "interrupt_response": True},
**(fmt(a.send_rate) or {})
},
"output": {**(fmt(a.recv_rate) or {}), "voice": "marin"}
}
}
}
async def realtime_loop(a: Args):
client = _make_client(a)
mic_q: Queue[bytes] = Queue(maxsize=128)
stop = Event()
playback = bytearray()
lock = Lock()
def mic_cb(indata, *_):
if a.block_mic_during_playback:
with lock:
if playback:
return
try:
mic_q.put_nowait(bytes(indata))
except Exception:
pass
def spk_cb(outdata, *_):
needed = len(outdata)
with lock:
avail = min(needed, len(playback))
outdata[:avail] = playback[:avail]
del playback[:avail]
if avail < needed:
outdata[avail:] = b"\x00" * (needed - avail)
in_stream = sd.RawInputStream(
samplerate=a.send_rate, channels=1, dtype="int16",
blocksize=a.chunk_size, callback=mic_cb
)
out_stream = sd.RawOutputStream(
samplerate=a.recv_rate, channels=1, dtype="int16",
blocksize=a.chunk_size, callback=spk_cb
)
in_stream.start()
out_stream.start()
async with client.realtime.connect(model=a.model) as conn:
await conn.send(_session_update(a))
async def send():
while not stop.is_set():
try:
chunk = await asyncio.to_thread(mic_q.get, True, 0.1)
except Empty:
continue
await conn.send({
"type": "input_audio_buffer.append",
"audio": base64.b64encode(chunk).decode()
})
async def recv():
while not stop.is_set():
ev = await conn.recv()
if a.print_json:
print(ev.model_dump_json())
if ev.type == "response.output_audio.delta":
data = base64.b64decode(ev.delta)
with lock:
playback.extend(data)
elif ev.type == "response.done":
stop.set()
await asyncio.gather(send(), recv())
in_stream.stop()
out_stream.stop()
if __name__ == "__main__":
asyncio.run(realtime_loop(args))
Note: The script relies only on the public
openaipackage and the repository's internal pipeline; no extra binary components are needed beyond the optionalspeech-to-speech[realtime]extras.
Summary
The huggingface/speech-to-speech repository provides a modular, extensible real-time speech-to-speech stack:
listen_and_play_realtime.pyprovides a complete working client example that captures microphone input and plays synthesized responsesRealtimeServerinserver.pymanages concurrent pipeline units via WebSocket connections at/v1/realtime- The processing flow moves through VAD → STT → LLM → LM-output-processor → TTS, with each component swappable via the handler factory system
- Supports multiple backend implementations including Whisper for STT, Qwen-3 for TTS, and various LLM backends (OpenAI, Hugging Face, MLX)
- Implements the OpenAI Realtime API protocol, making it compatible with existing OpenAI client libraries
Frequently Asked Questions
What is the latency of the real-time speech-to-speech pipeline?
The pipeline achieves low-latency conversion through streaming processing at each stage. The VAD handler uses enable_realtime_transcription to emit progressive transcription events before the audio segment completes, while the LM-output-processor buffers and streams partial audio chunks to the TTS handler as soon as the language model generates tokens. The WebSocket connection in websocket_streamer.py maintains persistent connections to minimize network overhead.
Can I use custom local models instead of OpenAI for the LLM component?
Yes. The build_pipeline function in s2s_pipeline.py supports multiple LLM backends through the --llm_backend argument. You can specify mlx for Apple Silicon optimization, transformers for Hugging Face models, or responses-api for OpenAI compatibility. The language_model.py handler abstracts these implementations behind a common interface.
How does the repository handle voice activity detection in noisy environments?
The VAD handler in vad_handler.py includes configurable parameters for voice activity detection. When running in real-time mode, it processes raw audio streams to detect speech segments and can optionally emit transcription events progressively. The server-side turn detection (server_vad with interrupt_response: True) allows the system to handle barge-in and interruptions gracefully.
Is it possible to stream audio over a network instead of localhost?
Yes. While the example uses 127.0.0.1, you can specify any host address using the --host parameter when starting the server in s2s_pipeline.py. The client in listen_and_play_realtime.py accepts --host and --port arguments to connect to remote instances. The WebSocket protocol in server.py handles byte-level streaming over standard network connections, though latency will depend on network conditions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →