WebRTC Transport Architecture in Speech-to-Speech: How It Differs from WebSocket Mode

The WebRTC transport architecture streams audio as RTP Opus packets over a peer connection while routing events through a reliable data channel, whereas WebSocket mode transmits base64-encoded PCM audio inside JSON messages, resulting in higher overhead but simpler setup.

The huggingface/speech-to-speech repository implements the OpenAI Realtime API with a pluggable transport layer that supports both WebSocket and WebRTC modes. Understanding the WebRTC transport architecture reveals significant differences in audio handling, latency characteristics, and interruption capabilities compared to the default WebSocket implementation.

Architectural Comparison: WebRTC vs WebSocket Transport

The repository abstracts transport functionality through the SessionTransport base class, with two concrete implementations handling fundamentally different network protocols.

Aspect WebSocket Transport WebRTC Transport
Transport Class WebSocketTransport in transports.py WebRTCSession in webrtc_session.py
Audio Format Base64-encoded PCM deltas inside JSON events RTP Opus packets on dedicated media track
Event Channel Same WebSocket connection (JSON messages) Reliable ordered data channel named "oai-events"
Client Handling Must buffer and decode base64 chunks Receives continuous stream playable via HTML5 audio
Setup HTTP upgrade to /v1/realtime SDP offer/answer exchange via /v1/realtime/calls
Barge-in No-op (client buffers audio) True flush via discard_pending_audio()
Dependencies None (always available) Requires aiortc (speech-to-speech[webrtc])

WebSocket Transport: JSON-Over-WebSocket

The WebSocketTransport class (src/speech_to_speech/api/openai_realtime/transports.py) implements the default transport mechanism using standard WebSocket connections. All OpenAI-compatible events (session.created, response.created, etc.) travel as JSON messages through the same socket.

Audio data follows a specific encoding path:

  • The RealtimeService encodes raw PCM bytes using service.encode_audio_chunk()
  • Audio transmits as base64-encoded deltas inside audio.delta events
  • The client must reconstruct the audio stream by decoding these chunks

Because the client maintains its own audio buffer, the discard_pending_audio() method performs no operation—interrupting playback requires client-side handling.


# Abstract base defining the transport interface

class SessionTransport(ABC):
    async def send_events(self, events: list[ServerEvent]) -> None: ...
    async def send_audio_chunk(self, service, session_id, pcm: bytes) -> None: ...
    def discard_pending_audio(self) -> None: ...
    async def close(self) -> None: ...

WebRTC Transport: Media Tracks and Data Channels

The WebRTCSession class (src/speech_to_speech/api/openai_realtime/webrtc_session.py) establishes a full RTCPeerConnection with two distinct channels:

  1. Media Track (PipelineAudioTrack): Handles outbound audio as RTP Opus packets, eliminating base64 encoding overhead
  2. Data Channel (oai-events): Transmits JSON events using the same schema as WebSocket mode

The implementation manages sample rate conversion through PcmResampler:

  • Incoming audio resamples from 48 kHz to the pipeline rate (16 kHz)
  • Outgoing audio resamples from 16 kHz to 48 kHz, with PipelineAudioTrack.recv() pacing RTP timing

Unlike WebSocket mode, the server controls the audio buffer, enabling true barge-in functionality.

Barge-in and Audio Interruption

A critical architectural difference lies in interruption handling:

  • WebSocket: The discard_pending_audio() method is a no-op because the client has already received and buffered base64-encoded chunks. Truncation requires client-side logic to discard pending audio.
  • WebRTC: The server can flush pending audio by clearing the PipelineAudioTrack buffer via transport.discard_pending_audio(), immediately stopping transmission and enabling conversational interruptions.

Transport Selection and Endpoints

The server selects transport implementations through distinct FastAPI endpoints defined in websocket_router.py:

  1. WebSocket Mode: The /v1/realtime endpoint creates a WebSocketTransport and claims a pipeline unit immediately
  2. WebRTC Mode: The /v1/realtime/calls endpoint accepts a POST request with an SDP offer, constructs an RTCPeerConnection, and returns the SDP answer

Both modes utilize the same _send_loop_for pipeline logic through the SessionTransport abstraction, ensuring consistent behavior regardless of underlying protocol.

Configuration and Dependencies

WebRTC support requires explicit installation:

pip install "speech-to-speech[webrtc]"

This installs the aiortc library and enables the WebRTC transport classes. Configuration uses the SPEECH_TO_SPEECH_ICE_SERVERS environment variable (JSON list) to populate the RTCConfiguration for NAT traversal.

Practical Implementation Examples

WebSocket Client Implementation

import websockets
import json

async def websocket_client():
    uri = "ws://localhost:8000/v1/realtime"
    async with websockets.connect(uri) as ws:
        # Receive session.created event

        print(await ws.recv())
        
        # Send base64-encoded PCM audio

        await ws.send(json.dumps({
            "type": "input_audio_buffer.append",
            "audio": "<base64-pcm-chunk>"
        }))
        
        # Receive audio deltas and events

        while True:
            msg = json.loads(await ws.recv())
            if msg["type"] == "audio.delta":
                # Decode base64 and play

                pass

WebRTC Client Implementation

import aiohttp
from aiortc import RTCPeerConnection, RTCSessionDescription
import json

async def webrtc_client():
    pc = RTCPeerConnection()
    
    # Create and send SDP offer

    offer = await pc.createOffer()
    await pc.setLocalDescription(offer)
    
    async with aiohttp.ClientSession() as session:
        async with session.post(
            "http://localhost:8000/v1/realtime/calls",
            data=pc.localDescription.sdp,
            headers={"Content-Type": "application/sdp"}
        ) as resp:
            answer_sdp = await resp.text()
    
    # Apply SDP answer

    await pc.setRemoteDescription(
        RTCSessionDescription(sdp=answer_sdp, type="answer")
    )
    
    # Create data channel for events

    dc = pc.createDataChannel("oai-events")
    dc.send(json.dumps({
        "type": "session.update",
        "session": {"voice": "alloy"}
    }))
    
    # Audio streams automatically on media track

Summary

  • WebRTC transport architecture uses RTP Opus packets and a separate data channel (oai-events), while WebSocket sends base64 PCM inside JSON messages
  • Source files: transports.py (WebSocket), webrtc_session.py (WebRTC), and websocket_router.py (endpoints) implement the SessionTransport abstraction
  • Audio handling: WebRTC resamples between 48 kHz and 16 kHz via PcmResampler, eliminating base64 encoding overhead
  • Interruption: Only WebRTC supports server-side discard_pending_audio() for true barge-in capabilities
  • Setup: WebRTC requires SDP negotiation via /v1/realtime/calls and the aiortc dependency, while WebSocket uses simple HTTP upgrade to /v1/realtime

Frequently Asked Questions

When should I choose WebRTC over WebSocket transport?

Choose WebRTC when you need lower latency streaming, smaller payload sizes (no base64 encoding), or server-controlled audio interruption (barge-in). Choose WebSocket for simpler client implementations, broader compatibility, or when aiortc dependencies are problematic in your environment.

How does audio resampling work in WebRTC mode?

The WebRTCSession class uses PcmResampler to convert between the WebRTC-standard 48 kHz sample rate and the pipeline's native 16 kHz. Incoming audio resamples from 48 kHz to 16 kHz for processing, while outgoing audio resamples from 16 kHz to 48 kHz before transmission via PipelineAudioTrack.

Why does WebSocket mode not support true barge-in?

In WebSocket mode, audio transmits as base64-encoded audio.delta events that the client buffers immediately upon receipt. Once sent, the server cannot recall these events, making discard_pending_audio() a no-op. The client must implement its own truncation logic, whereas WebRTC's PipelineAudioTrack maintains a server-side buffer that can be cleared instantly.

What environment configuration is required for WebRTC support?

You must install the extra dependency using pip install "speech-to-speech[webrtc]" and configure the SPEECH_TO_SPEECH_ICE_SERVERS environment variable with a JSON list of ICE servers (e.g., STUN/TURN servers) to enable NAT traversal during peer connection establishment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →