WebSocket vs Socket Modes in HuggingFace Speech-to-Speech: Key Differences

WebSocket mode implements the full OpenAI Realtime API with structured events and interruption handling, while Socket mode offers a minimal TCP transport for raw 16 kHz PCM audio only.

The huggingface/speech-to-speech repository provides multiple transport modes for running the speech-to-speech pipeline. Understanding the difference between WebSocket and Socket modes is essential for choosing the right integration method for your application.

Protocol and Transport Layer Differences

The fundamental distinction lies in the underlying network protocol and the level of abstraction provided to clients.

WebSocket Mode Architecture

WebSocket mode operates over the WebSocket protocol and supports two distinct sub-modes: the full OpenAI Realtime API (--mode realtime) and raw PCM streaming (--mode websocket). According to the source code in src/speech_to_speech/connections/websocket_streamer.py, the WebSocketStreamer class manages the complete Realtime event loop, handling session updates, transcription events, tool calls, and audio deltas.

When running in realtime mode, the server exposes the /v1/realtime endpoint and implements the full OpenAI Realtime protocol. This includes structured JSON events for conversation state management, voice activity detection, and function calling capabilities.

Socket Mode Architecture

Socket mode (--mode socket) utilizes plain TCP sockets without the WebSocket handshake overhead. As implemented in src/speech_to_speech/connections/socket_sender.py and src/speech_to_speech/connections/socket_receiver.py, this mode uses the SocketSender and SocketReceiver classes to forward raw audio bytes bidirectionally. The protocol is intentionally minimal: it streams only raw 16 kHz int16 mono PCM bytes with no message framing or structured events.

Feature Comparison

The two modes serve different use cases based on their feature sets:

Feature WebSocket Mode Socket Mode
Transport Protocol WebSocket (RFC 6455) Plain TCP
Audio Format Raw PCM or encoded (via Realtime API) Raw 16 kHz int16 mono PCM only
Event Structure JSON events (session.update, conversation.item.create, etc.) No structured events
Interruption Handling Supported via input_audio_buffer.clear events Not supported
Live Transcription Available through conversation.item.input_audio_transcription Not available
Tool Calling Full function calling support Not supported
CLI Flag --mode realtime or --mode websocket --mode socket

Code Examples

Starting WebSocket (Realtime) Mode

To launch the server with full OpenAI Realtime API compatibility:

speech-to-speech \
    --mode realtime \
    --ws_host 0.0.0.0 \
    --ws_port 8765

Starting Raw PCM WebSocket Mode

For raw PCM streaming without Realtime protocol overhead:

speech-to-speech \
    --mode websocket \
    --ws_host 0.0.0.0 \
    --ws_port 8765

Starting Socket (TCP) Mode

For minimal TCP-based transport:

speech-to-speech \
    --mode socket \
    --recv_host 0.0.0.0 \
    --send_host 0.0.0.0

Client Connection Examples

Connecting to WebSocket mode using the OpenAI Python client:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed",
)

with client.realtime.connect(model="local") as conn:
    conn.send({"type": "session.update", "session": {"type": "realtime"}})
    for event in conn:
        print(event.type)

Connecting to Socket mode using the provided helper script:

python scripts/listen_and_play.py --host <SERVER_IP>

Source Code Architecture

The implementation differences are visible in the connection handler classes defined in src/speech_to_speech/arguments_classes/module_arguments.py, which parses the --mode flag to instantiate the appropriate transport layer.

The CLI arguments for WebSocket configuration are defined in src/speech_to_speech/arguments_classes/websocket_streamer_arguments.py, including --ws_host and --ws_port parameters.

Summary

  • WebSocket mode provides a rich, event-driven API compatible with the OpenAI Realtime specification, supporting interruptions, transcripts, and tool calls.
  • Socket mode offers a lightweight, low-latency transport for raw PCM audio without protocol overhead.
  • WebSocket mode is ideal for browser clients and voice assistants requiring full conversational state management.
  • Socket mode suits embedded systems or legacy pipelines needing simple audio streaming.
  • The --mode flag in module_arguments.py determines which transport classes are instantiated at runtime.

Frequently Asked Questions

When should I use WebSocket mode versus Socket mode?

WebSocket mode is the correct choice when building interactive applications that require real-time transcription visibility, interruption handling, or integration with OpenAI-compatible clients. Socket mode works best for simple pipelines where you only need to send audio and receive generated audio without conversation state management, such as embedded systems or custom TCP-based audio processors.

Does Socket mode support the OpenAI Realtime protocol?

No. Socket mode deliberately omits the Realtime API feature set. As implemented in socket_sender.py and socket_receiver.py, it exchanges only raw 16 kHz int16 mono PCM bytes without JSON event framing. This makes it incompatible with standard OpenAI clients but more efficient for custom implementations that handle their own protocol layer.

What audio format is required for Socket mode?

Socket mode requires raw PCM audio at 16 kHz sample rate, 16-bit integer depth, and mono channel configuration. The SocketReceiver class in src/speech_to_speech/connections/socket_receiver.py expects this specific format and forwards raw bytes directly to the pipeline without conversion or validation.

Can I use standard OpenAI SDK clients with WebSocket mode?

Yes. When running with --mode realtime, the server exposes the /v1/realtime endpoint that is fully compatible with the OpenAI Realtime API. You can connect using the official OpenAI Python or JavaScript SDKs by pointing the websocket base URL to your local server instance, as shown in the client example above.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →