# WebRTC Transport Architecture in Speech-to-Speech: How It Differs from WebSocket Mode

> Explore the WebRTC transport architecture for speech-to-speech, differentiating it from WebSocket mode by understanding RTP packet streaming and data channel routing for efficient audio transmission.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: architecture
- Published: 2026-07-31

---

**The WebRTC transport architecture streams audio as RTP Opus packets over a peer connection while routing events through a reliable data channel, whereas WebSocket mode transmits base64-encoded PCM audio inside JSON messages, resulting in higher overhead but simpler setup.**

The huggingface/speech-to-speech repository implements the OpenAI Realtime API with a pluggable transport layer that supports both WebSocket and WebRTC modes. Understanding the WebRTC transport architecture reveals significant differences in audio handling, latency characteristics, and interruption capabilities compared to the default WebSocket implementation.

## Architectural Comparison: WebRTC vs WebSocket Transport

The repository abstracts transport functionality through the `SessionTransport` base class, with two concrete implementations handling fundamentally different network protocols.

| Aspect | WebSocket Transport | WebRTC Transport |
|--------|---------------------|------------------|
| **Transport Class** | `WebSocketTransport` in [`transports.py`](https://github.com/huggingface/speech-to-speech/blob/main/transports.py) | `WebRTCSession` in [`webrtc_session.py`](https://github.com/huggingface/speech-to-speech/blob/main/webrtc_session.py) |
| **Audio Format** | Base64-encoded PCM deltas inside JSON events | RTP Opus packets on dedicated media track |
| **Event Channel** | Same WebSocket connection (JSON messages) | Reliable ordered data channel named `"oai-events"` |
| **Client Handling** | Must buffer and decode base64 chunks | Receives continuous stream playable via HTML5 audio |
| **Setup** | HTTP upgrade to `/v1/realtime` | SDP offer/answer exchange via `/v1/realtime/calls` |
| **Barge-in** | No-op (client buffers audio) | True flush via `discard_pending_audio()` |
| **Dependencies** | None (always available) | Requires `aiortc` (`speech-to-speech[webrtc]`) |

## WebSocket Transport: JSON-Over-WebSocket

The `WebSocketTransport` class ([`src/speech_to_speech/api/openai_realtime/transports.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/transports.py)) implements the default transport mechanism using standard WebSocket connections. All OpenAI-compatible events (`session.created`, `response.created`, etc.) travel as JSON messages through the same socket.

Audio data follows a specific encoding path:

- The `RealtimeService` encodes raw PCM bytes using `service.encode_audio_chunk()`
- Audio transmits as **base64-encoded deltas** inside `audio.delta` events
- The client must reconstruct the audio stream by decoding these chunks

Because the client maintains its own audio buffer, the `discard_pending_audio()` method performs no operation—interrupting playback requires client-side handling.

```python

# Abstract base defining the transport interface

class SessionTransport(ABC):
    async def send_events(self, events: list[ServerEvent]) -> None: ...
    async def send_audio_chunk(self, service, session_id, pcm: bytes) -> None: ...
    def discard_pending_audio(self) -> None: ...
    async def close(self) -> None: ...

```

## WebRTC Transport: Media Tracks and Data Channels

The `WebRTCSession` class ([`src/speech_to_speech/api/openai_realtime/webrtc_session.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/webrtc_session.py)) establishes a full **RTCPeerConnection** with two distinct channels:

1. **Media Track (`PipelineAudioTrack`)**: Handles outbound audio as **RTP Opus packets**, eliminating base64 encoding overhead
2. **Data Channel (`oai-events`)**: Transmits JSON events using the same schema as WebSocket mode

The implementation manages sample rate conversion through `PcmResampler`:

- Incoming audio resamples from **48 kHz** to the pipeline rate (**16 kHz**)
- Outgoing audio resamples from 16 kHz to 48 kHz, with `PipelineAudioTrack.recv()` pacing RTP timing

Unlike WebSocket mode, the server controls the audio buffer, enabling **true barge-in** functionality.

## Barge-in and Audio Interruption

A critical architectural difference lies in interruption handling:

- **WebSocket**: The `discard_pending_audio()` method is a no-op because the client has already received and buffered base64-encoded chunks. Truncation requires client-side logic to discard pending audio.
- **WebRTC**: The server can flush pending audio by clearing the `PipelineAudioTrack` buffer via `transport.discard_pending_audio()`, immediately stopping transmission and enabling conversational interruptions.

## Transport Selection and Endpoints

The server selects transport implementations through distinct FastAPI endpoints defined in [`websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_router.py):

1. **WebSocket Mode**: The `/v1/realtime` endpoint creates a `WebSocketTransport` and claims a pipeline unit immediately
2. **WebRTC Mode**: The `/v1/realtime/calls` endpoint accepts a POST request with an SDP offer, constructs an `RTCPeerConnection`, and returns the SDP answer

Both modes utilize the same `_send_loop_for` pipeline logic through the `SessionTransport` abstraction, ensuring consistent behavior regardless of underlying protocol.

## Configuration and Dependencies

WebRTC support requires explicit installation:

```bash
pip install "speech-to-speech[webrtc]"

```

This installs the `aiortc` library and enables the WebRTC transport classes. Configuration uses the `SPEECH_TO_SPEECH_ICE_SERVERS` environment variable (JSON list) to populate the `RTCConfiguration` for NAT traversal.

## Practical Implementation Examples

### WebSocket Client Implementation

```python
import websockets
import json

async def websocket_client():
    uri = "ws://localhost:8000/v1/realtime"
    async with websockets.connect(uri) as ws:
        # Receive session.created event

        print(await ws.recv())
        
        # Send base64-encoded PCM audio

        await ws.send(json.dumps({
            "type": "input_audio_buffer.append",
            "audio": "<base64-pcm-chunk>"
        }))
        
        # Receive audio deltas and events

        while True:
            msg = json.loads(await ws.recv())
            if msg["type"] == "audio.delta":
                # Decode base64 and play

                pass

```

### WebRTC Client Implementation

```python
import aiohttp
from aiortc import RTCPeerConnection, RTCSessionDescription
import json

async def webrtc_client():
    pc = RTCPeerConnection()
    
    # Create and send SDP offer

    offer = await pc.createOffer()
    await pc.setLocalDescription(offer)
    
    async with aiohttp.ClientSession() as session:
        async with session.post(
            "http://localhost:8000/v1/realtime/calls",
            data=pc.localDescription.sdp,
            headers={"Content-Type": "application/sdp"}
        ) as resp:
            answer_sdp = await resp.text()
    
    # Apply SDP answer

    await pc.setRemoteDescription(
        RTCSessionDescription(sdp=answer_sdp, type="answer")
    )
    
    # Create data channel for events

    dc = pc.createDataChannel("oai-events")
    dc.send(json.dumps({
        "type": "session.update",
        "session": {"voice": "alloy"}
    }))
    
    # Audio streams automatically on media track

```

## Summary

- **WebRTC transport architecture** uses RTP Opus packets and a separate data channel (`oai-events`), while WebSocket sends base64 PCM inside JSON messages
- **Source files**: [`transports.py`](https://github.com/huggingface/speech-to-speech/blob/main/transports.py) (WebSocket), [`webrtc_session.py`](https://github.com/huggingface/speech-to-speech/blob/main/webrtc_session.py) (WebRTC), and [`websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_router.py) (endpoints) implement the `SessionTransport` abstraction
- **Audio handling**: WebRTC resamples between 48 kHz and 16 kHz via `PcmResampler`, eliminating base64 encoding overhead
- **Interruption**: Only WebRTC supports server-side `discard_pending_audio()` for true barge-in capabilities
- **Setup**: WebRTC requires SDP negotiation via `/v1/realtime/calls` and the `aiortc` dependency, while WebSocket uses simple HTTP upgrade to `/v1/realtime`

## Frequently Asked Questions

### When should I choose WebRTC over WebSocket transport?

Choose **WebRTC** when you need lower latency streaming, smaller payload sizes (no base64 encoding), or server-controlled audio interruption (barge-in). Choose **WebSocket** for simpler client implementations, broader compatibility, or when `aiortc` dependencies are problematic in your environment.

### How does audio resampling work in WebRTC mode?

The `WebRTCSession` class uses `PcmResampler` to convert between the WebRTC-standard 48 kHz sample rate and the pipeline's native 16 kHz. Incoming audio resamples from 48 kHz to 16 kHz for processing, while outgoing audio resamples from 16 kHz to 48 kHz before transmission via `PipelineAudioTrack`.

### Why does WebSocket mode not support true barge-in?

In WebSocket mode, audio transmits as base64-encoded `audio.delta` events that the client buffers immediately upon receipt. Once sent, the server cannot recall these events, making `discard_pending_audio()` a no-op. The client must implement its own truncation logic, whereas WebRTC's `PipelineAudioTrack` maintains a server-side buffer that can be cleared instantly.

### What environment configuration is required for WebRTC support?

You must install the extra dependency using `pip install "speech-to-speech[webrtc]"` and configure the `SPEECH_TO_SPEECH_ICE_SERVERS` environment variable with a JSON list of ICE servers (e.g., STUN/TURN servers) to enable NAT traversal during peer connection establishment.