# How to Implement Custom Audio Streaming for Embedded Devices with Speech-to-Speech

> Implement custom audio streaming for embedded devices using huggingface/speech-to-speech. Stream 16 kHz PCM audio via raw TCP sockets without complex protocols.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**The huggingface/speech-to-speech repository enables custom audio streaming for embedded devices through raw TCP socket mode, which transmits 16 kHz PCM audio over two independent ports (12345 for input, 12346 for output) without requiring WebSocket libraries or complex protocol stacks.**

The huggingface/speech-to-speech repository ships a modular voice-agent pipeline optimized for low-latency speech-to-speech applications. While the framework supports three distinct transport modes, implementing custom audio streaming for embedded devices requires selecting the **Socket mode**, which replaces WebSocket overhead with standard POSIX TCP sockets ideal for microcontrollers and single-board computers.

## Why Choose Socket Mode for Embedded Hardware

The repository exposes three transport configurations, but only one minimizes resource consumption for constrained environments:

- **Realtime mode**: Uses WebSocket with the OpenAI Realtime protocol. Suitable for full-featured clients requiring tool calls and complex session management, but carries unnecessary overhead for microcontrollers.
- **WebSocket mode**: Transmits raw PCM over WebSocket. Lighter than Realtime mode but still requires a WebSocket client implementation.
- **Socket mode**: Transmits raw PCM over standard TCP sockets. Requires only basic socket APIs available in C, MicroPython, Rust, or Arduino, making it optimal for ESP-32 devices, Raspberry Pi boards, and custom ARM-based systems.

Socket mode eliminates the need for HTTP upgrade headers, framing masks, or text-based protocols. The server-side implementation in [`src/speech_to_speech/connections/socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_receiver.py) and [`src/speech_to_speech/connections/socket_sender.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_sender.py) handles all audio buffering and control signaling, allowing your embedded client to focus solely on capturing and playing PCM bytes.

## Architecture of the Socket Transport Layer

The socket implementation splits bidirectional audio into two unidirectional TCP connections managed by dedicated classes.

### SocketReceiver: Ingesting Audio from Embedded Clients

The **SocketReceiver** class (defined in [`src/speech_to_speech/connections/socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_receiver.py)) listens on port **12345** for incoming TCP connections. When an embedded client connects, the receiver:

1. Sets the `should_listen` threading event to activate the Voice Activity Detection (VAD) pipeline.
2. Reads raw PCM chunks (default **1024 bytes** or 512 samples) from the socket.
3. Enqueues audio data onto `queue_out` for processing by the STT → LLM → TTS pipeline.
4. Injects a **PIPELINE_END** control message (defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py)) when the client disconnects, triggering graceful pipeline shutdown.

The receiver expects PCM format: **16 kHz sample rate, signed 16-bit little-endian, mono channel**.

### SocketSender: Delivering Generated Audio to Devices

The **SocketSender** class (defined in [`src/speech_to_speech/connections/socket_sender.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_sender.py)) listens on port **12346** for outbound connections. This component:

1. Monitors `queue_in` for generated audio from the TTS module.
2. Converts NumPy arrays or byte buffers to raw PCM.
3. Streams bytes back to the embedded client via the TCP socket.
4. Intercepts the **AUDIO_RESPONSE_DONE** control message to re-enable listening by setting `should_listen`, allowing the conversation to continue.

Both classes use the queue abstractions defined in [`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py) (specifically `AudioInItem` and `AudioOutItem`) and filter control messages using utilities from [`src/speech_to_speech/pipeline/control.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/control.py).

## Implementing the Embedded Client

Your embedded implementation must open two simultaneous TCP connections to the server and handle non-blocking I/O for real-time performance.

### Python Client for Linux Single-Board Computers

For prototyping on Raspberry Pi or similar boards, use this Python implementation with the `sounddevice` library for audio hardware abstraction:

```python
import socket
import sounddevice as sd
import numpy as np

HOST = "192.168.1.100"    # Server IP address

INPUT_PORT = 12345        # To SocketReceiver

OUTPUT_PORT = 12346       # From SocketSender

SAMPLE_RATE = 16000
CHUNK = 512               # Frames per packet

def audio_callback(indata, frames, time, status):
    """Capture microphone and stream to server."""
    tx_sock.sendall(indata.tobytes())

# Establish outbound connection (microphone → server)

tx_sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
tx_sock.connect((HOST, INPUT_PORT))

# Establish inbound connection (server → speakers)

rx_sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
rx_sock.connect((HOST, OUTPUT_PORT))
rx_sock.setblocking(False)

# Start microphone stream

stream = sd.InputStream(
    samplerate=SAMPLE_RATE,
    dtype="int16",
    channels=1,
    callback=audio_callback,
    blocksize=CHUNK,
)
stream.start()

print("Streaming audio. Press Ctrl-C to stop.")
try:
    while True:
        try:
            data = rx_sock.recv(4096)
            if data:
                audio = np.frombuffer(data, dtype=np.int16).reshape(-1, 1)
                sd.play(audio, SAMPLE_RATE, blocking=False)
        except BlockingIOError:
            pass
except KeyboardInterrupt:
    pass
finally:
    stream.stop()
    tx_sock.close()
    rx_sock.close()

```

This client maintains separate sockets for upstream (microphone) and downstream (speaker) audio, matching the server-side separation of concerns in [`socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/socket_receiver.py) and [`socket_sender.py`](https://github.com/huggingface/speech-to-speech/blob/main/socket_sender.py).

### C Implementation for Embedded Linux

For resource-constrained Linux systems where Python overhead is unacceptable, implement the client in C using POSIX sockets and pthreads:

```c
#include <arpa/inet.h>
#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>

#define SAMPLE_RATE 16000
#define CHUNK_SAMPLES 512
#define CHUNK_BYTES (CHUNK_SAMPLES * sizeof(int16_t))

int tx_sock, rx_sock;

void *capture_thread(void *arg) {
    int16_t buffer[CHUNK_SAMPLES];
    while (1) {
        // Capture from ADC/microphone driver into buffer
        // write() blocks until buffer transmitted
        write(tx_sock, buffer, CHUNK_BYTES);
    }
    return NULL;
}

void *playback_thread(void *arg) {
    int16_t buffer[CHUNK_SAMPLES];
    ssize_t n;
    while ((n = read(rx_sock, buffer, CHUNK_BYTES)) > 0) {
        // Send buffer to DAC/speaker driver
    }
    return NULL;
}

int main(void) {
    struct sockaddr_in srv_addr;
    
    // Setup TX socket (to server port 12345)
    tx_sock = socket(AF_INET, SOCK_STREAM, 0);
    srv_addr.sin_family = AF_INET;
    srv_addr.sin_port = htons(12345);
    inet_pton(AF_INET, "192.168.1.100", &srv_addr.sin_addr);
    connect(tx_sock, (struct sockaddr *)&srv_addr, sizeof(srv_addr));
    
    // Setup RX socket (from server port 12346)
    rx_sock = socket(AF_INET, SOCK_STREAM, 0);
    srv_addr.sin_port = htons(12346);
    connect(rx_sock, (struct sockaddr *)&srv_addr, sizeof(srv_addr));
    
    pthread_t cap_tid, play_tid;
    pthread_create(&cap_tid, NULL, capture_thread, NULL);
    pthread_create(&play_tid, NULL, playback_thread, NULL);
    
    pthread_join(cap_tid, NULL);
    pthread_join(play_tid, NULL);
    return 0;
}

```

This implementation uses standard POSIX APIs available in embedded Linux distributions and RTOS environments with BSD socket support.

### MicroPython for ESP-32 Microcontrollers

For bare-metal microcontrollers like the ESP-32, use MicroPython with the I2S peripheral for direct digital audio acquisition:

```python
import network
import socket
import machine

HOST = "192.168.1.100"
INPUT_PORT = 12345
OUTPUT_PORT = 12346
SAMPLE_RATE = 16000
CHUNK_BYTES = 1024  # 512 samples * 2 bytes

# Wi-Fi connection

wlan = network.WLAN(network.STA_IF)
wlan.active(True)
wlan.connect("SSID", "PASSWORD")
while not wlan.isconnected():
    pass

# TCP sockets

tx = socket.socket()
tx.connect((HOST, INPUT_PORT))

rx = socket.socket()
rx.connect((HOST, OUTPUT_PORT))
rx.setblocking(False)

# I2S configuration for INMP441 or similar MEMS microphone

i2s = machine.I2S(
    0,
    sck=machine.Pin(26),
    ws=machine.Pin(25),
    sd=machine.Pin(22),
    mode=machine.I2S.RX,
    bits=16,
    format=machine.I2S.MONO,
    rate=SAMPLE_RATE,
    ibuf=CHUNK_BYTES,
)

while True:
    # Read from microphone and transmit

    buf = bytearray(CHUNK_BYTES)
    if i2s.readinto(buf) == CHUNK_BYTES:
        tx.send(buf)
    
    # Receive from server and queue for DAC

    try:
        audio = rx.recv(CHUNK_BYTES)
        if audio:
            # Push to PWM or I2S DAC

            pass
    except OSError as e:
        if e.args[0] != 11:  # EAGAIN

            raise

```

This MicroPython client demonstrates that implementing custom audio streaming for embedded devices requires only socket and I2S primitives, with no external dependencies beyond the standard library.

## Key Source Files and Protocol Details

Understanding the internal implementation helps debug connection issues and optimize buffer sizes:

- **[`src/speech_to_speech/connections/socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_receiver.py)**: Implements the server-side listener on port 12345. Manages the `should_listen` event and handles **PIPELINE_END** control messages when clients disconnect.
- **[`src/speech_to_speech/connections/socket_sender.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_sender.py)**: Implements the outbound streamer on port 12346. Monitors for **AUDIO_RESPONSE_DONE** to toggle listening states.
- **[`src/speech_to_speech/pipeline/queue_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/queue_types.py)**: Defines `AudioInItem` and `AudioOutItem` type aliases used for inter-thread communication.
- **[`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py)**: Contains control constants: `PIPELINE_END`, `AUDIO_RESPONSE_DONE`, and `SESSION_END`.
- **[`src/speech_to_speech/pipeline/control.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/control.py)**: Provides `is_control_message()` helper to distinguish audio bytes from control signals.
- **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)**: Orchestrates the full pipeline; launches receiver and sender threads based on the `--mode socket` flag.

## Deployment Workflow

To integrate your embedded hardware with the speech-to-speech pipeline:

1. **Start the server** in socket mode:
   ```bash
   python -m speech_to_speech --mode socket \
       --stt parakeet-tdt \
       --tts qwen3 \
       --host 0.0.0.0
   ```

2. **Connect your embedded client** to ports 12345 and 12346 on the server IP.

3. **Stream audio**: The client transmits 16-bit PCM chunks continuously. The pipeline processes speech through VAD → STT → LLM → TTS.

4. **Handle turn-taking**: When the server finishes generating audio, it sends the **AUDIO_RESPONSE_DONE** signal. Your client should resume transmitting microphone data immediately to initiate the next turn.

The server automatically manages session lifecycle through control messages defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py), ensuring robust handling of network interruptions without pipeline stalls.

## Summary

- **Socket mode** provides the lightest transport for embedded devices, using raw TCP instead of WebSocket protocols.
- **Two independent sockets** handle bidirectional audio: port 12345 for input (handled by **SocketReceiver**) and port 12346 for output (handled by **SocketSender**).
- **PCM format requirements** are strict: 16 kHz, signed 16-bit little-endian, mono. Any deviation requires resampling on the client.
- **Control messages** (**PIPELINE_END**, **AUDIO_RESPONSE_DONE**, **SESSION_END**) coordinate conversation flow and session management without additional wire protocol complexity.
- **Implementation flexibility** allows clients in C, Python, MicroPython, or Rust, provided they can open standard TCP sockets and handle the specified audio format.

## Frequently Asked Questions

### What audio format does the socket mode require?

The socket transport expects **16 kHz sample rate, signed 16-bit little-endian, mono PCM**. Each sample occupies 2 bytes. The default chunk size is 1024 bytes (512 samples), though the `SocketReceiver` implementation accepts arbitrary packet sizes and buffers them internally.

### Can I use WebSocket instead of raw sockets on my microcontroller?

While possible, WebSocket mode introduces unnecessary overhead for embedded targets. Raw sockets require approximately 4KB of RAM for buffering versus 20-50KB for a minimal WebSocket client. Unless you specifically need the OpenAI Realtime protocol features, **Socket mode** provides lower latency and simpler implementation on constrained hardware.

### How does the server coordinate turn-taking in conversations?

The server uses the **AUDIO_RESPONSE_DONE** control message injected into the pipeline queue. When **SocketSender** detects this message (defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py)), it triggers the `should_listen` threading event, which reactivates the VAD module in **SocketReceiver**. Your client does not need to manage state; simply continue streaming microphone data and the server will naturally interleave listening and speaking phases.

### What happens if my embedded client disconnects unexpectedly?

**SocketReceiver** monitors the TCP connection for closure or read failures. Upon disconnection, it automatically injects a **PIPELINE_END** message into the queue, which the pipeline control loop interprets as a signal to flush buffers and terminate the current session gracefully. When your client reconnects, the server spawns a new receiver thread and resumes normal operation.