# Hugging Face Speech-to-Speech Run Modes: Local, Socket, WebSocket, and Realtime Explained

> Explore Hugging Face speech-to-speech run modes local, socket, websocket, and realtime. Understand their differences and choose the best mode for your audio streaming needs.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-30

---

**The Hugging Face `speech-to-speech` library supports four distinct run modes—local, socket, websocket, and realtime—each wiring different audio transport implementations into the pipeline via the `--mode` argument, from local microphone streaming to OpenAI-compatible Realtime API servers.**

The `speech-to-speech` repository provides an end-to-end voice conversion pipeline with flexible transport options for different deployment architectures. You control how audio flows into and out of the system using the `--mode` argument defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py), which determines how the `build_pipeline` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) instantiates specific handler classes.

## Understanding the Four Speech-to-Speech Run Modes

Each mode configures a distinct transport layer implementation that determines where audio is captured from and where output is sent.

### Local Mode: Direct Microphone and Speaker Access

**Local mode** uses the `LocalAudioStreamer` class to read audio directly from your machine’s microphone and write generated speech to local speakers. In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the branch `if module_kwargs.mode == "local"` (lines 42‑53) creates this streamer and adds it to the `comms_handlers` list.

This is the simplest configuration for testing the pipeline on a single machine without network overhead. Use this mode when you want to speak into your laptop’s mic and hear the response immediately through connected headphones or speakers.

### Socket Mode: Raw TCP for External Processes

**Socket mode** implements raw TCP socket communication using `SocketReceiver` and `SocketSender`. When mode is set to `socket`, the `else` block in `build_pipeline` (lines 107‑127) instantiates these classes with default host `0.0.0.0`, receive port `12346`, and chunk size of 1024 bytes.

Use this mode when audio capture or playback is handled by a separate process, such as a custom C++ frontend, a hardware device streaming raw PCM, or a remote virtual machine. The sockets decouple I/O from the Python pipeline, allowing you to send audio data over TCP and receive generated speech on a separate port.

### WebSocket Mode: Browser-Compatible Streaming

**WebSocket mode** leverages the `WebSocketStreamer` class to provide asynchronous bidirectional communication over HTTP. The `elif module_kwargs.mode == "websocket"` branch (lines 53‑66) registers this handler as the sole communications handler, defaulting to host `0.0.0.0` and port `12345`.

This mode is ideal for web-based clients and JavaScript frontends. Browser applications can connect to `ws://host:port/v1/realtime` to stream audio to the pipeline and receive synthesized responses without dealing with raw TCP socket management.

### Realtime Mode: OpenAI-Compatible Concurrent Sessions

**Realtime mode** implements the OpenAI Realtime API specification using a `RealtimeServer` and a pool of isolated pipeline units. The `elif module_kwargs.mode == "realtime"` branch (lines 66‑81) builds these units via `_build_realtime_pipeline_unit`, each containing its own VAD, STT, LLM, and TTS components. The server starts on lines 95‑101 and exposes the `/v1/realtime` endpoint.

Use this mode when you need full OpenAI-compatible realtime behavior, support for multiple concurrent websocket sessions, or advanced features like turn-based streaming and tool use. Each incoming websocket connection is routed to a free pipeline unit from the pool.

## When to Use Each Speech-to-Speech Mode

Select your mode based on where your audio originates and how many clients you need to support:

- **`local`** – Use for simple local prototyping and debugging on a single laptop where you just want to talk to the model and hear the result immediately.
- **`socket`** – Use when integrating with an external audio source or custom client that streams raw PCM over TCP, such as IoT devices or separate capture services.
- **`websocket`** – Use for browser-based UIs and JavaScript applications that require standard WebSocket connections rather than raw TCP sockets.
- **`realtime`** – Use for production deployments requiring OpenAI Realtime API compatibility, turn handling, or support for many simultaneous websocket clients.

## Implementation Examples

### Running a Local Demo

Start the pipeline with direct microphone and speaker access:

```bash
python -m speech_to_speech.run \
    --mode local \
    --stt whisper \
    --llm responses-api \
    --tts pocket

```

*This executes the `local` branch in `build_pipeline`, instantiating `LocalAudioStreamer`.*

### Streaming via TCP Sockets

Launch the server to accept raw PCM over TCP:

```bash

# Server side

python -m speech_to_speech.run \
    --mode socket \
    --socket_receiver_kwargs.recv_host 0.0.0.0 \
    --socket_receiver_kwargs.recv_port 12346 \
    --socket_sender_kwargs.send_host 0.0.0.0 \
    --socket_sender_kwargs.send_port 12347 \
    --stt whisper

```

*This triggers the socket `else` block in `build_pipeline`, creating `SocketReceiver` and `SocketSender` instances.*

### Connecting from a Browser

Start the WebSocket endpoint for JavaScript clients:

```bash
python -m speech_to_speech.run \
    --mode websocket \
    --ws_host 0.0.0.0 \
    --ws_port 12345 \
    --stt whisper

```

Then connect from the browser:

```javascript
const ws = new WebSocket("ws://localhost:12345/v1/realtime");

```

*The `websocket` branch in `build_pipeline` instantiates `WebSocketStreamer`.*

### Deploying the OpenAI Realtime Server

Launch the realtime server with three concurrent pipeline units:

```bash
python -m speech_to_speech.run \
    --mode realtime \
    --num_pipelines 3 \
    --ws_host 0.0.0.0 \
    --ws_port 12345 \
    --stt whisper

```

*This executes the `realtime` branch, calling `_build_realtime_pipeline_unit` to create isolated pipeline instances and starting `RealtimeServer` for OpenAI-compatible routing.*

## Key Source Files

These files define the mode-specific transport implementations:

- **Mode argument definition**: [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py)
- **Pipeline builder**: [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (contains `build_pipeline` and `_build_realtime_pipeline_unit`)
- **Local audio**: [`src/speech_to_speech/connections/local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/local_audio_streamer.py)
- **TCP sockets**: [`src/speech_to_speech/connections/socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_receiver.py) and [`src/speech_to_speech/connections/socket_sender.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_sender.py)
- **WebSocket transport**: [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py)
- **Realtime API**: [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) and [`src/speech_to_speech/api/openai_realtime/websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/websocket_router.py)

## Summary

- **Local mode** wires `LocalAudioStreamer` for direct microphone/speaker I/O, ideal for single-machine demos.
- **Socket mode** configures `SocketReceiver` and `SocketSender` for raw TCP communication with external processes or hardware.
- **WebSocket mode** uses `WebSocketStreamer` to support browser-based JavaScript clients with standard websocket connections.
- **Realtime mode** deploys `RealtimeServer` with a pool of isolated pipeline units to provide OpenAI Realtime API compatibility and concurrent session support.

## Frequently Asked Questions

### Can I switch between modes without changing the model configuration?

Yes. The `--mode` argument only affects the transport layer (`comms_handlers`) in `build_pipeline`. Your STT, LLM, and TTS model selections remain independent of whether you use local, socket, websocket, or realtime transport, allowing you to test locally then deploy to production using the same model weights.

### Why does socket mode use two different ports?

The `SocketReceiver` and `SocketSender` operate on separate TCP ports (default 12346 for receiving audio, configurable for sending) to maintain unidirectional data flow separation. This design allows you to route captured audio from one device and send generated audio to another, or to integrate with systems that handle input and output through different network endpoints.

### How many concurrent clients can the realtime mode handle?

The realtime mode handles concurrency through a pool of isolated pipeline units sized by the `--num_pipelines` argument. Each websocket connection gets routed to a free unit by `RealtimeServer`; if all units are busy, new connections wait until a unit becomes available. Scale by increasing `--num_pipelines` based on your GPU memory and compute capacity.

### Is the WebSocket mode compatible with the OpenAI Realtime API?

No. Only **`realtime`** mode implements the OpenAI Realtime API specification on the `/v1/realtime` endpoint. The **`websocket`** mode uses `WebSocketStreamer` for generic bidirectional audio streaming without OpenAI-specific message formats, session management, or turn detection. Use `realtime` mode specifically when you need OpenAI client compatibility.