WebRTC Transport Architecture in Speech-to-Speech: How It Differs from WebSocket Mode
The WebRTC transport architecture streams audio as RTP Opus packets over a peer connection while routing events through a reliable data channel, whereas WebSocket mode transmits base64-encoded PCM audio inside JSON messages, resulting in higher overhead but simpler setup.
The huggingface/speech-to-speech repository implements the OpenAI Realtime API with a pluggable transport layer that supports both WebSocket and WebRTC modes. Understanding the WebRTC transport architecture reveals significant differences in audio handling, latency characteristics, and interruption capabilities compared to the default WebSocket implementation.
Architectural Comparison: WebRTC vs WebSocket Transport
The repository abstracts transport functionality through the SessionTransport base class, with two concrete implementations handling fundamentally different network protocols.
| Aspect | WebSocket Transport | WebRTC Transport |
|---|---|---|
| Transport Class | WebSocketTransport in transports.py |
WebRTCSession in webrtc_session.py |
| Audio Format | Base64-encoded PCM deltas inside JSON events | RTP Opus packets on dedicated media track |
| Event Channel | Same WebSocket connection (JSON messages) | Reliable ordered data channel named "oai-events" |
| Client Handling | Must buffer and decode base64 chunks | Receives continuous stream playable via HTML5 audio |
| Setup | HTTP upgrade to /v1/realtime |
SDP offer/answer exchange via /v1/realtime/calls |
| Barge-in | No-op (client buffers audio) | True flush via discard_pending_audio() |
| Dependencies | None (always available) | Requires aiortc (speech-to-speech[webrtc]) |
WebSocket Transport: JSON-Over-WebSocket
The WebSocketTransport class (src/speech_to_speech/api/openai_realtime/transports.py) implements the default transport mechanism using standard WebSocket connections. All OpenAI-compatible events (session.created, response.created, etc.) travel as JSON messages through the same socket.
Audio data follows a specific encoding path:
- The
RealtimeServiceencodes raw PCM bytes usingservice.encode_audio_chunk() - Audio transmits as base64-encoded deltas inside
audio.deltaevents - The client must reconstruct the audio stream by decoding these chunks
Because the client maintains its own audio buffer, the discard_pending_audio() method performs no operation—interrupting playback requires client-side handling.
# Abstract base defining the transport interface
class SessionTransport(ABC):
async def send_events(self, events: list[ServerEvent]) -> None: ...
async def send_audio_chunk(self, service, session_id, pcm: bytes) -> None: ...
def discard_pending_audio(self) -> None: ...
async def close(self) -> None: ...
WebRTC Transport: Media Tracks and Data Channels
The WebRTCSession class (src/speech_to_speech/api/openai_realtime/webrtc_session.py) establishes a full RTCPeerConnection with two distinct channels:
- Media Track (
PipelineAudioTrack): Handles outbound audio as RTP Opus packets, eliminating base64 encoding overhead - Data Channel (
oai-events): Transmits JSON events using the same schema as WebSocket mode
The implementation manages sample rate conversion through PcmResampler:
- Incoming audio resamples from 48 kHz to the pipeline rate (16 kHz)
- Outgoing audio resamples from 16 kHz to 48 kHz, with
PipelineAudioTrack.recv()pacing RTP timing
Unlike WebSocket mode, the server controls the audio buffer, enabling true barge-in functionality.
Barge-in and Audio Interruption
A critical architectural difference lies in interruption handling:
- WebSocket: The
discard_pending_audio()method is a no-op because the client has already received and buffered base64-encoded chunks. Truncation requires client-side logic to discard pending audio. - WebRTC: The server can flush pending audio by clearing the
PipelineAudioTrackbuffer viatransport.discard_pending_audio(), immediately stopping transmission and enabling conversational interruptions.
Transport Selection and Endpoints
The server selects transport implementations through distinct FastAPI endpoints defined in websocket_router.py:
- WebSocket Mode: The
/v1/realtimeendpoint creates aWebSocketTransportand claims a pipeline unit immediately - WebRTC Mode: The
/v1/realtime/callsendpoint accepts a POST request with an SDP offer, constructs anRTCPeerConnection, and returns the SDP answer
Both modes utilize the same _send_loop_for pipeline logic through the SessionTransport abstraction, ensuring consistent behavior regardless of underlying protocol.
Configuration and Dependencies
WebRTC support requires explicit installation:
pip install "speech-to-speech[webrtc]"
This installs the aiortc library and enables the WebRTC transport classes. Configuration uses the SPEECH_TO_SPEECH_ICE_SERVERS environment variable (JSON list) to populate the RTCConfiguration for NAT traversal.
Practical Implementation Examples
WebSocket Client Implementation
import websockets
import json
async def websocket_client():
uri = "ws://localhost:8000/v1/realtime"
async with websockets.connect(uri) as ws:
# Receive session.created event
print(await ws.recv())
# Send base64-encoded PCM audio
await ws.send(json.dumps({
"type": "input_audio_buffer.append",
"audio": "<base64-pcm-chunk>"
}))
# Receive audio deltas and events
while True:
msg = json.loads(await ws.recv())
if msg["type"] == "audio.delta":
# Decode base64 and play
pass
WebRTC Client Implementation
import aiohttp
from aiortc import RTCPeerConnection, RTCSessionDescription
import json
async def webrtc_client():
pc = RTCPeerConnection()
# Create and send SDP offer
offer = await pc.createOffer()
await pc.setLocalDescription(offer)
async with aiohttp.ClientSession() as session:
async with session.post(
"http://localhost:8000/v1/realtime/calls",
data=pc.localDescription.sdp,
headers={"Content-Type": "application/sdp"}
) as resp:
answer_sdp = await resp.text()
# Apply SDP answer
await pc.setRemoteDescription(
RTCSessionDescription(sdp=answer_sdp, type="answer")
)
# Create data channel for events
dc = pc.createDataChannel("oai-events")
dc.send(json.dumps({
"type": "session.update",
"session": {"voice": "alloy"}
}))
# Audio streams automatically on media track
Summary
- WebRTC transport architecture uses RTP Opus packets and a separate data channel (
oai-events), while WebSocket sends base64 PCM inside JSON messages - Source files:
transports.py(WebSocket),webrtc_session.py(WebRTC), andwebsocket_router.py(endpoints) implement theSessionTransportabstraction - Audio handling: WebRTC resamples between 48 kHz and 16 kHz via
PcmResampler, eliminating base64 encoding overhead - Interruption: Only WebRTC supports server-side
discard_pending_audio()for true barge-in capabilities - Setup: WebRTC requires SDP negotiation via
/v1/realtime/callsand theaiortcdependency, while WebSocket uses simple HTTP upgrade to/v1/realtime
Frequently Asked Questions
When should I choose WebRTC over WebSocket transport?
Choose WebRTC when you need lower latency streaming, smaller payload sizes (no base64 encoding), or server-controlled audio interruption (barge-in). Choose WebSocket for simpler client implementations, broader compatibility, or when aiortc dependencies are problematic in your environment.
How does audio resampling work in WebRTC mode?
The WebRTCSession class uses PcmResampler to convert between the WebRTC-standard 48 kHz sample rate and the pipeline's native 16 kHz. Incoming audio resamples from 48 kHz to 16 kHz for processing, while outgoing audio resamples from 16 kHz to 48 kHz before transmission via PipelineAudioTrack.
Why does WebSocket mode not support true barge-in?
In WebSocket mode, audio transmits as base64-encoded audio.delta events that the client buffers immediately upon receipt. Once sent, the server cannot recall these events, making discard_pending_audio() a no-op. The client must implement its own truncation logic, whereas WebRTC's PipelineAudioTrack maintains a server-side buffer that can be cleared instantly.
What environment configuration is required for WebRTC support?
You must install the extra dependency using pip install "speech-to-speech[webrtc]" and configure the SPEECH_TO_SPEECH_ICE_SERVERS environment variable with a JSON list of ICE servers (e.g., STUN/TURN servers) to enable NAT traversal during peer connection establishment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →