How to Connect OpenAI Realtime Clients to the Speech-to-Speech WebSocket Server
Use the RealtimeClient SDK from the huggingface/speech-to-speech repository to establish an authenticated WebSocket connection, register audio callbacks, and stream PCM audio frames to the OpenAI Realtime endpoint.
The huggingface/speech-to-speech repository provides a complete OpenAI Realtime API implementation that bridges Python clients to a WebSocket server. Connecting OpenAI Realtime clients requires understanding three architectural layers: the WebSocket service layer that manages the protocol, the client SDK that abstracts low-level socket operations, and the application layer that wires everything together. This guide walks through each layer with code references from the actual source.
WebSocket Service Architecture
The repository's Realtime stack is organized into distinct layers that handle different responsibilities.
WebSocket Service Layer
The service layer follows the OpenAI Realtime protocol specification. It handles JSON-based control messages, binary audio frames, and keep-alive pings to maintain low-latency connections.
The demo server in demo/server.py starts this service and exposes the /realtime WebSocket endpoint. This server is protocol-compatible with OpenAI's official Realtime API, meaning any compliant client can connect.
Client SDK Layer
The openai_realtime package provides a thin Python wrapper that transforms raw WebSocket calls into high-level events. The RealtimeClient class manages handshake, authentication, and message routing.
Key events exposed by the SDK:
on_turn_start— fired when the server detects the beginning of a conversational turnon_turn_end— fired when the server finishes generating its responseon_output_audio— delivers synthesized speech bytes from the LLMon_error— handles protocol violations or connection failures
The implementation is validated in tests/openai_realtime/test_openai_client.py, which exercises the full client lifecycle.
Application Layer
User-facing scripts glue the SDK to speech pipelines. The canonical reference is listen_and_play_realtime.py, which creates a RealtimeClient, streams microphone input, and plays back real-time responses.
Step-by-Step Connection Process
1. Configure API Credentials
The client authenticates using the OPENAI_API_KEY environment variable. The constructor also accepts an explicit token parameter.
import os
api_key = os.getenv("OPENAI_API_KEY")
if not api_key:
raise RuntimeError("Set OPENAI_API_KEY environment variable")
This pattern appears in listen_and_play_realtime.py at line referencing os.getenv("OPENAI_API_KEY").
2. Initialize the RealtimeClient
Instantiate RealtimeClient with your API key. The endpoint URL (wss://api.openai.com/v1/realtime) is baked into the SDK default, though you can override it for custom deployments.
from openai_realtime import RealtimeClient
client = RealtimeClient(api_key=api_key)
Under the hood, this creates an aiohttp WebSocket session and registers protocol handlers in openai_realtime/client.py.
3. Register Audio and Event Callbacks
Define coroutines that handle server events. At minimum, implement on_output_audio to receive synthesized speech.
import numpy as np
import sounddevice as sd
async def on_output_audio(chunk: bytes):
"""Play incoming PCM audio at 16kHz."""
audio_array = np.frombuffer(chunk, dtype=np.int16)
sd.play(audio_array, samplerate=16000)
client.on_output_audio = on_output_audio
Optional callbacks include on_turn_start, on_turn_end, and on_error for UI synchronization and error recovery.
4. Connect and Perform Handshake
The connect() method establishes the WebSocket and waits for the session.created response containing the session_id.
await client.connect() # Blocks until handshake completes
This logic is exercised in tests/openai_realtime/test_realtime_service.py, which validates the message exchange format.
5. Start a Turn and Stream Audio
Signal intent to speak with start_turn(), then stream PCM frames at 16 kHz. The SDK batches frames and respects the server's max_input_frames limit.
await client.start_turn()
# Stream microphone frames (16-bit PCM, 16000 Hz)
async for frame in capture_microphone():
await client.send_audio(frame)
await client.end_turn() # Signal completion
The send_audio() method in openai_realtime/client.py handles frame batching and protocol encoding.
6. Handle Server Responses
While your turn ends, the server begins streaming output_audio events. Your on_output_audio callback receives chunks as they arrive, enabling true real-time playback.
The test suite in tests/openai_realtime/test_smart_turn.py validates this turn-handling lifecycle.
Complete Minimal Example
import os
import asyncio
import numpy as np
import sounddevice as sd
from openai_realtime import RealtimeClient
async def main():
# Authentication
api_key = os.getenv("OPENAI_API_KEY")
if not api_key:
raise RuntimeError("Set OPENAI_API_KEY environment variable")
# Initialize client
client = RealtimeClient(api_key=api_key)
# Define audio playback callback
async def on_output_audio(chunk: bytes):
audio = np.frombuffer(chunk, dtype=np.int16)
sd.play(audio, samplerate=16000)
sd.wait()
client.on_output_audio = on_output_audio
# Connect and start session
await client.connect()
await client.start_turn()
# Stream 5 seconds of microphone audio
stream = sd.RawInputStream(
samplerate=16000,
blocksize=1600, # 100ms chunks
dtype=np.int16,
channels=1
)
stream.start()
for _ in range(50): # 5 seconds
frame, _ = stream.read(1600)
await client.send_audio(frame.tobytes())
await asyncio.sleep(0.01)
stream.stop()
await client.end_turn()
# Allow time for server response
await asyncio.sleep(5)
if __name__ == "__main__":
asyncio.run(main())
Production Reference: listen_and_play_realtime.py
The repository ships a production-ready script that extends the minimal example with voice activity detection and pipeline integration:
export OPENAI_API_KEY=sk-************************************
python scripts/listen_and_play_realtime.py
This script in scripts/listen_and_play_realtime.py adds:
- Voice Activity Detection (VAD) to auto-detect speech endpoints
- Whisper streaming transcription for real-time STT
- Pipeline integration with Parakeet, MEL-O, or other TTS backends
Key Design Characteristics
Async-first architecture — The SDK uses asyncio and aiohttp for non-blocking I/O, essential for sub-100ms audio latency.
Modular event dispatch — Each protocol message type routes to a dedicated coroutine, enabling custom behaviors like logging or analytics injection.
Audio pipeline abstraction — Utilities in utils/audio.py and utils/thread_manager.py decouple capture/playback from WebSocket logic for cross-platform portability.
Error Handling and Resilience
Protocol errors trigger the on_error callback with structured error codes. The client optionally implements exponential backoff reconnection, validated in tests/openai_realtime/test_websocket_session_lifecycle.py.
Common error scenarios include:
- Authentication failures (invalid or expired API key)
- Rate limiting (HTTP 429 equivalents over WebSocket)
- Protocol violations (malformed message sequences)
File Reference Guide
| File | Purpose |
|---|---|
scripts/listen_and_play_realtime.py |
Production reference implementation with VAD and pipelines |
tests/openai_realtime/test_openai_client.py |
Unit tests for RealtimeClient handshake and messaging |
tests/openai_realtime/test_realtime_service.py |
Integration tests for server-side protocol compliance |
tests/openai_realtime/test_websocket_session_lifecycle.py |
Connection resilience and reconnection logic |
tests/openai_realtime/test_smart_turn.py |
Turn lifecycle validation |
demo/server.py |
Minimal WebSocket server for local testing |
src/speech_to_speech/baseHandler.py |
Core event routing to STT/TTS pipelines |
Summary
- Obtain an OpenAI API key and pass it to
RealtimeClientinitialization - Register callbacks for
on_output_audioand optional turn management events - Call
connect()to perform the WebSocket handshake and receivesession_id - Use
start_turn()/end_turn()to bracket audio streaming, sending PCM frames viasend_audio() - Reference
listen_and_play_realtime.pyfor production patterns including VAD and pipeline integration
Frequently Asked Questions
What audio format does the RealtimeClient expect?
The client expects 16-bit PCM audio at 16000 Hz, delivered as bytes objects. The SDK handles internal framing and protocol encoding. Downsample or convert your source audio using utils/audio.py utilities if needed.
Can I connect to a custom WebSocket server instead of OpenAI's endpoint?
Yes. Pass the base_url parameter to RealtimeClient to override the default wss://api.openai.com/v1/realtime. The server must implement the same JSON/binary message protocol verified in tests/openai_realtime/test_realtime_service.py.
How does the SDK handle connection drops?
The RealtimeClient includes optional exponential backoff reconnection, exercised in test_websocket_session_lifecycle.py. Set auto_reconnect=True in the constructor and provide an on_error callback to monitor recovery status.
Is WebRTC supported for browser clients?
The repository includes WebRTC forwarding in tests/openai_realtime/test_webrtc.py. This wraps the Realtime WebSocket with a media server bridge, enabling browser-based clients to use WebRTC audio tracks while the backend speaks the native Realtime protocol.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →