WebSocket vs Socket Modes in HuggingFace Speech-to-Speech: Key Differences
WebSocket mode implements the full OpenAI Realtime API with structured events and interruption handling, while Socket mode offers a minimal TCP transport for raw 16 kHz PCM audio only.
The huggingface/speech-to-speech repository provides multiple transport modes for running the speech-to-speech pipeline. Understanding the difference between WebSocket and Socket modes is essential for choosing the right integration method for your application.
Protocol and Transport Layer Differences
The fundamental distinction lies in the underlying network protocol and the level of abstraction provided to clients.
WebSocket Mode Architecture
WebSocket mode operates over the WebSocket protocol and supports two distinct sub-modes: the full OpenAI Realtime API (--mode realtime) and raw PCM streaming (--mode websocket). According to the source code in src/speech_to_speech/connections/websocket_streamer.py, the WebSocketStreamer class manages the complete Realtime event loop, handling session updates, transcription events, tool calls, and audio deltas.
When running in realtime mode, the server exposes the /v1/realtime endpoint and implements the full OpenAI Realtime protocol. This includes structured JSON events for conversation state management, voice activity detection, and function calling capabilities.
Socket Mode Architecture
Socket mode (--mode socket) utilizes plain TCP sockets without the WebSocket handshake overhead. As implemented in src/speech_to_speech/connections/socket_sender.py and src/speech_to_speech/connections/socket_receiver.py, this mode uses the SocketSender and SocketReceiver classes to forward raw audio bytes bidirectionally. The protocol is intentionally minimal: it streams only raw 16 kHz int16 mono PCM bytes with no message framing or structured events.
Feature Comparison
The two modes serve different use cases based on their feature sets:
| Feature | WebSocket Mode | Socket Mode |
|---|---|---|
| Transport Protocol | WebSocket (RFC 6455) | Plain TCP |
| Audio Format | Raw PCM or encoded (via Realtime API) | Raw 16 kHz int16 mono PCM only |
| Event Structure | JSON events (session.update, conversation.item.create, etc.) | No structured events |
| Interruption Handling | Supported via input_audio_buffer.clear events |
Not supported |
| Live Transcription | Available through conversation.item.input_audio_transcription |
Not available |
| Tool Calling | Full function calling support | Not supported |
| CLI Flag | --mode realtime or --mode websocket |
--mode socket |
Code Examples
Starting WebSocket (Realtime) Mode
To launch the server with full OpenAI Realtime API compatibility:
speech-to-speech \
--mode realtime \
--ws_host 0.0.0.0 \
--ws_port 8765
Starting Raw PCM WebSocket Mode
For raw PCM streaming without Realtime protocol overhead:
speech-to-speech \
--mode websocket \
--ws_host 0.0.0.0 \
--ws_port 8765
Starting Socket (TCP) Mode
For minimal TCP-based transport:
speech-to-speech \
--mode socket \
--recv_host 0.0.0.0 \
--send_host 0.0.0.0
Client Connection Examples
Connecting to WebSocket mode using the OpenAI Python client:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8765/v1",
websocket_base_url="ws://localhost:8765/v1",
api_key="not-needed",
)
with client.realtime.connect(model="local") as conn:
conn.send({"type": "session.update", "session": {"type": "realtime"}})
for event in conn:
print(event.type)
Connecting to Socket mode using the provided helper script:
python scripts/listen_and_play.py --host <SERVER_IP>
Source Code Architecture
The implementation differences are visible in the connection handler classes defined in src/speech_to_speech/arguments_classes/module_arguments.py, which parses the --mode flag to instantiate the appropriate transport layer.
-
WebSocket implementation:
src/speech_to_speech/connections/websocket_streamer.pycontains theWebSocketStreamerclass that manages the Realtime protocol state machine and event routing. -
Socket implementation: The TCP mode splits send and receive responsibilities between
src/speech_to_speech/connections/socket_sender.py(SocketSenderclass) andsrc/speech_to_speech/connections/socket_receiver.py(SocketReceiverclass), both handling raw byte streams without protocol parsing.
The CLI arguments for WebSocket configuration are defined in src/speech_to_speech/arguments_classes/websocket_streamer_arguments.py, including --ws_host and --ws_port parameters.
Summary
- WebSocket mode provides a rich, event-driven API compatible with the OpenAI Realtime specification, supporting interruptions, transcripts, and tool calls.
- Socket mode offers a lightweight, low-latency transport for raw PCM audio without protocol overhead.
- WebSocket mode is ideal for browser clients and voice assistants requiring full conversational state management.
- Socket mode suits embedded systems or legacy pipelines needing simple audio streaming.
- The
--modeflag inmodule_arguments.pydetermines which transport classes are instantiated at runtime.
Frequently Asked Questions
When should I use WebSocket mode versus Socket mode?
WebSocket mode is the correct choice when building interactive applications that require real-time transcription visibility, interruption handling, or integration with OpenAI-compatible clients. Socket mode works best for simple pipelines where you only need to send audio and receive generated audio without conversation state management, such as embedded systems or custom TCP-based audio processors.
Does Socket mode support the OpenAI Realtime protocol?
No. Socket mode deliberately omits the Realtime API feature set. As implemented in socket_sender.py and socket_receiver.py, it exchanges only raw 16 kHz int16 mono PCM bytes without JSON event framing. This makes it incompatible with standard OpenAI clients but more efficient for custom implementations that handle their own protocol layer.
What audio format is required for Socket mode?
Socket mode requires raw PCM audio at 16 kHz sample rate, 16-bit integer depth, and mono channel configuration. The SocketReceiver class in src/speech_to_speech/connections/socket_receiver.py expects this specific format and forwards raw bytes directly to the pipeline without conversion or validation.
Can I use standard OpenAI SDK clients with WebSocket mode?
Yes. When running with --mode realtime, the server exposes the /v1/realtime endpoint that is fully compatible with the OpenAI Realtime API. You can connect using the official OpenAI Python or JavaScript SDKs by pointing the websocket base URL to your local server instance, as shown in the client example above.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →