Hugging Face Speech-to-Speech Run Modes: Local, Socket, WebSocket, and Realtime Explained

The Hugging Face speech-to-speech library supports four distinct run modes—local, socket, websocket, and realtime—each wiring different audio transport implementations into the pipeline via the --mode argument, from local microphone streaming to OpenAI-compatible Realtime API servers.

The speech-to-speech repository provides an end-to-end voice conversion pipeline with flexible transport options for different deployment architectures. You control how audio flows into and out of the system using the --mode argument defined in src/speech_to_speech/arguments_classes/module_arguments.py, which determines how the build_pipeline function in src/speech_to_speech/s2s_pipeline.py instantiates specific handler classes.

Understanding the Four Speech-to-Speech Run Modes

Each mode configures a distinct transport layer implementation that determines where audio is captured from and where output is sent.

Local Mode: Direct Microphone and Speaker Access

Local mode uses the LocalAudioStreamer class to read audio directly from your machine’s microphone and write generated speech to local speakers. In src/speech_to_speech/s2s_pipeline.py, the branch if module_kwargs.mode == "local" (lines 42‑53) creates this streamer and adds it to the comms_handlers list.

This is the simplest configuration for testing the pipeline on a single machine without network overhead. Use this mode when you want to speak into your laptop’s mic and hear the response immediately through connected headphones or speakers.

Socket Mode: Raw TCP for External Processes

Socket mode implements raw TCP socket communication using SocketReceiver and SocketSender. When mode is set to socket, the else block in build_pipeline (lines 107‑127) instantiates these classes with default host 0.0.0.0, receive port 12346, and chunk size of 1024 bytes.

Use this mode when audio capture or playback is handled by a separate process, such as a custom C++ frontend, a hardware device streaming raw PCM, or a remote virtual machine. The sockets decouple I/O from the Python pipeline, allowing you to send audio data over TCP and receive generated speech on a separate port.

WebSocket Mode: Browser-Compatible Streaming

WebSocket mode leverages the WebSocketStreamer class to provide asynchronous bidirectional communication over HTTP. The elif module_kwargs.mode == "websocket" branch (lines 53‑66) registers this handler as the sole communications handler, defaulting to host 0.0.0.0 and port 12345.

This mode is ideal for web-based clients and JavaScript frontends. Browser applications can connect to ws://host:port/v1/realtime to stream audio to the pipeline and receive synthesized responses without dealing with raw TCP socket management.

Realtime Mode: OpenAI-Compatible Concurrent Sessions

Realtime mode implements the OpenAI Realtime API specification using a RealtimeServer and a pool of isolated pipeline units. The elif module_kwargs.mode == "realtime" branch (lines 66‑81) builds these units via _build_realtime_pipeline_unit, each containing its own VAD, STT, LLM, and TTS components. The server starts on lines 95‑101 and exposes the /v1/realtime endpoint.

Use this mode when you need full OpenAI-compatible realtime behavior, support for multiple concurrent websocket sessions, or advanced features like turn-based streaming and tool use. Each incoming websocket connection is routed to a free pipeline unit from the pool.

When to Use Each Speech-to-Speech Mode

Select your mode based on where your audio originates and how many clients you need to support:

  • local – Use for simple local prototyping and debugging on a single laptop where you just want to talk to the model and hear the result immediately.
  • socket – Use when integrating with an external audio source or custom client that streams raw PCM over TCP, such as IoT devices or separate capture services.
  • websocket – Use for browser-based UIs and JavaScript applications that require standard WebSocket connections rather than raw TCP sockets.
  • realtime – Use for production deployments requiring OpenAI Realtime API compatibility, turn handling, or support for many simultaneous websocket clients.

Implementation Examples

Running a Local Demo

Start the pipeline with direct microphone and speaker access:

python -m speech_to_speech.run \
    --mode local \
    --stt whisper \
    --llm responses-api \
    --tts pocket

This executes the local branch in build_pipeline, instantiating LocalAudioStreamer.

Streaming via TCP Sockets

Launch the server to accept raw PCM over TCP:


# Server side

python -m speech_to_speech.run \
    --mode socket \
    --socket_receiver_kwargs.recv_host 0.0.0.0 \
    --socket_receiver_kwargs.recv_port 12346 \
    --socket_sender_kwargs.send_host 0.0.0.0 \
    --socket_sender_kwargs.send_port 12347 \
    --stt whisper

This triggers the socket else block in build_pipeline, creating SocketReceiver and SocketSender instances.

Connecting from a Browser

Start the WebSocket endpoint for JavaScript clients:

python -m speech_to_speech.run \
    --mode websocket \
    --ws_host 0.0.0.0 \
    --ws_port 12345 \
    --stt whisper

Then connect from the browser:

const ws = new WebSocket("ws://localhost:12345/v1/realtime");

The websocket branch in build_pipeline instantiates WebSocketStreamer.

Deploying the OpenAI Realtime Server

Launch the realtime server with three concurrent pipeline units:

python -m speech_to_speech.run \
    --mode realtime \
    --num_pipelines 3 \
    --ws_host 0.0.0.0 \
    --ws_port 12345 \
    --stt whisper

This executes the realtime branch, calling _build_realtime_pipeline_unit to create isolated pipeline instances and starting RealtimeServer for OpenAI-compatible routing.

Key Source Files

These files define the mode-specific transport implementations:

Summary

  • Local mode wires LocalAudioStreamer for direct microphone/speaker I/O, ideal for single-machine demos.
  • Socket mode configures SocketReceiver and SocketSender for raw TCP communication with external processes or hardware.
  • WebSocket mode uses WebSocketStreamer to support browser-based JavaScript clients with standard websocket connections.
  • Realtime mode deploys RealtimeServer with a pool of isolated pipeline units to provide OpenAI Realtime API compatibility and concurrent session support.

Frequently Asked Questions

Can I switch between modes without changing the model configuration?

Yes. The --mode argument only affects the transport layer (comms_handlers) in build_pipeline. Your STT, LLM, and TTS model selections remain independent of whether you use local, socket, websocket, or realtime transport, allowing you to test locally then deploy to production using the same model weights.

Why does socket mode use two different ports?

The SocketReceiver and SocketSender operate on separate TCP ports (default 12346 for receiving audio, configurable for sending) to maintain unidirectional data flow separation. This design allows you to route captured audio from one device and send generated audio to another, or to integrate with systems that handle input and output through different network endpoints.

How many concurrent clients can the realtime mode handle?

The realtime mode handles concurrency through a pool of isolated pipeline units sized by the --num_pipelines argument. Each websocket connection gets routed to a free unit by RealtimeServer; if all units are busy, new connections wait until a unit becomes available. Scale by increasing --num_pipelines based on your GPU memory and compute capacity.

Is the WebSocket mode compatible with the OpenAI Realtime API?

No. Only realtime mode implements the OpenAI Realtime API specification on the /v1/realtime endpoint. The websocket mode uses WebSocketStreamer for generic bidirectional audio streaming without OpenAI-specific message formats, session management, or turn detection. Use realtime mode specifically when you need OpenAI client compatibility.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →