# WebSocket vs Socket Modes in HuggingFace Speech-to-Speech: Key Differences

> Understand WebSocket vs Socket modes in HuggingFace Speech-to-Speech. WebSocket offers structured events for OpenAI API, while Socket provides minimal TCP for raw audio.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-11

---

**WebSocket mode implements the full OpenAI Realtime API with structured events and interruption handling, while Socket mode offers a minimal TCP transport for raw 16 kHz PCM audio only.**

The `huggingface/speech-to-speech` repository provides multiple transport modes for running the speech-to-speech pipeline. Understanding the difference between **WebSocket** and **Socket** modes is essential for choosing the right integration method for your application.

## Protocol and Transport Layer Differences

The fundamental distinction lies in the underlying network protocol and the level of abstraction provided to clients.

### WebSocket Mode Architecture

**WebSocket mode** operates over the WebSocket protocol and supports two distinct sub-modes: the full **OpenAI Realtime API** (`--mode realtime`) and raw PCM streaming (`--mode websocket`). According to the source code in [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py), the `WebSocketStreamer` class manages the complete Realtime event loop, handling session updates, transcription events, tool calls, and audio deltas.

When running in `realtime` mode, the server exposes the `/v1/realtime` endpoint and implements the full OpenAI Realtime protocol. This includes structured JSON events for conversation state management, voice activity detection, and function calling capabilities.

### Socket Mode Architecture

**Socket mode** (`--mode socket`) utilizes plain TCP sockets without the WebSocket handshake overhead. As implemented in [`src/speech_to_speech/connections/socket_sender.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_sender.py) and [`src/speech_to_speech/connections/socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_receiver.py), this mode uses the `SocketSender` and `SocketReceiver` classes to forward raw audio bytes bidirectionally. The protocol is intentionally minimal: it streams only raw 16 kHz int16 mono PCM bytes with no message framing or structured events.

## Feature Comparison

The two modes serve different use cases based on their feature sets:

| Feature | WebSocket Mode | Socket Mode |
|---------|---------------|-------------|
| **Transport Protocol** | WebSocket (RFC 6455) | Plain TCP |
| **Audio Format** | Raw PCM or encoded (via Realtime API) | Raw 16 kHz int16 mono PCM only |
| **Event Structure** | JSON events (session.update, conversation.item.create, etc.) | No structured events |
| **Interruption Handling** | Supported via `input_audio_buffer.clear` events | Not supported |
| **Live Transcription** | Available through `conversation.item.input_audio_transcription` | Not available |
| **Tool Calling** | Full function calling support | Not supported |
| **CLI Flag** | `--mode realtime` or `--mode websocket` | `--mode socket` |

## Code Examples

### Starting WebSocket (Realtime) Mode

To launch the server with full OpenAI Realtime API compatibility:

```bash
speech-to-speech \
    --mode realtime \
    --ws_host 0.0.0.0 \
    --ws_port 8765

```

### Starting Raw PCM WebSocket Mode

For raw PCM streaming without Realtime protocol overhead:

```bash
speech-to-speech \
    --mode websocket \
    --ws_host 0.0.0.0 \
    --ws_port 8765

```

### Starting Socket (TCP) Mode

For minimal TCP-based transport:

```bash
speech-to-speech \
    --mode socket \
    --recv_host 0.0.0.0 \
    --send_host 0.0.0.0

```

### Client Connection Examples

**Connecting to WebSocket mode** using the OpenAI Python client:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed",
)

with client.realtime.connect(model="local") as conn:
    conn.send({"type": "session.update", "session": {"type": "realtime"}})
    for event in conn:
        print(event.type)

```

**Connecting to Socket mode** using the provided helper script:

```bash
python scripts/listen_and_play.py --host <SERVER_IP>

```

## Source Code Architecture

The implementation differences are visible in the connection handler classes defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py), which parses the `--mode` flag to instantiate the appropriate transport layer.

- **WebSocket implementation**: [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) contains the `WebSocketStreamer` class that manages the Realtime protocol state machine and event routing.

- **Socket implementation**: The TCP mode splits send and receive responsibilities between [`src/speech_to_speech/connections/socket_sender.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_sender.py) (`SocketSender` class) and [`src/speech_to_speech/connections/socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_receiver.py) (`SocketReceiver` class), both handling raw byte streams without protocol parsing.

The CLI arguments for WebSocket configuration are defined in [`src/speech_to_speech/arguments_classes/websocket_streamer_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/websocket_streamer_arguments.py), including `--ws_host` and `--ws_port` parameters.

## Summary

- **WebSocket mode** provides a rich, event-driven API compatible with the OpenAI Realtime specification, supporting interruptions, transcripts, and tool calls.
- **Socket mode** offers a lightweight, low-latency transport for raw PCM audio without protocol overhead.
- WebSocket mode is ideal for browser clients and voice assistants requiring full conversational state management.
- Socket mode suits embedded systems or legacy pipelines needing simple audio streaming.
- The `--mode` flag in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) determines which transport classes are instantiated at runtime.

## Frequently Asked Questions

### When should I use WebSocket mode versus Socket mode?

**WebSocket mode** is the correct choice when building interactive applications that require real-time transcription visibility, interruption handling, or integration with OpenAI-compatible clients. **Socket mode** works best for simple pipelines where you only need to send audio and receive generated audio without conversation state management, such as embedded systems or custom TCP-based audio processors.

### Does Socket mode support the OpenAI Realtime protocol?

No. Socket mode deliberately omits the Realtime API feature set. As implemented in [`socket_sender.py`](https://github.com/huggingface/speech-to-speech/blob/main/socket_sender.py) and [`socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/socket_receiver.py), it exchanges only raw 16 kHz int16 mono PCM bytes without JSON event framing. This makes it incompatible with standard OpenAI clients but more efficient for custom implementations that handle their own protocol layer.

### What audio format is required for Socket mode?

Socket mode requires **raw PCM audio** at 16 kHz sample rate, 16-bit integer depth, and mono channel configuration. The `SocketReceiver` class in [`src/speech_to_speech/connections/socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_receiver.py) expects this specific format and forwards raw bytes directly to the pipeline without conversion or validation.

### Can I use standard OpenAI SDK clients with WebSocket mode?

Yes. When running with `--mode realtime`, the server exposes the `/v1/realtime` endpoint that is fully compatible with the OpenAI Realtime API. You can connect using the official OpenAI Python or JavaScript SDKs by pointing the websocket base URL to your local server instance, as shown in the client example above.