How to Implement Custom Audio Streaming for Embedded Devices with Speech-to-Speech
The huggingface/speech-to-speech repository enables custom audio streaming for embedded devices through raw TCP socket mode, which transmits 16 kHz PCM audio over two independent ports (12345 for input, 12346 for output) without requiring WebSocket libraries or complex protocol stacks.
The huggingface/speech-to-speech repository ships a modular voice-agent pipeline optimized for low-latency speech-to-speech applications. While the framework supports three distinct transport modes, implementing custom audio streaming for embedded devices requires selecting the Socket mode, which replaces WebSocket overhead with standard POSIX TCP sockets ideal for microcontrollers and single-board computers.
Why Choose Socket Mode for Embedded Hardware
The repository exposes three transport configurations, but only one minimizes resource consumption for constrained environments:
- Realtime mode: Uses WebSocket with the OpenAI Realtime protocol. Suitable for full-featured clients requiring tool calls and complex session management, but carries unnecessary overhead for microcontrollers.
- WebSocket mode: Transmits raw PCM over WebSocket. Lighter than Realtime mode but still requires a WebSocket client implementation.
- Socket mode: Transmits raw PCM over standard TCP sockets. Requires only basic socket APIs available in C, MicroPython, Rust, or Arduino, making it optimal for ESP-32 devices, Raspberry Pi boards, and custom ARM-based systems.
Socket mode eliminates the need for HTTP upgrade headers, framing masks, or text-based protocols. The server-side implementation in src/speech_to_speech/connections/socket_receiver.py and src/speech_to_speech/connections/socket_sender.py handles all audio buffering and control signaling, allowing your embedded client to focus solely on capturing and playing PCM bytes.
Architecture of the Socket Transport Layer
The socket implementation splits bidirectional audio into two unidirectional TCP connections managed by dedicated classes.
SocketReceiver: Ingesting Audio from Embedded Clients
The SocketReceiver class (defined in src/speech_to_speech/connections/socket_receiver.py) listens on port 12345 for incoming TCP connections. When an embedded client connects, the receiver:
- Sets the
should_listenthreading event to activate the Voice Activity Detection (VAD) pipeline. - Reads raw PCM chunks (default 1024 bytes or 512 samples) from the socket.
- Enqueues audio data onto
queue_outfor processing by the STT → LLM → TTS pipeline. - Injects a PIPELINE_END control message (defined in
src/speech_to_speech/pipeline/messages.py) when the client disconnects, triggering graceful pipeline shutdown.
The receiver expects PCM format: 16 kHz sample rate, signed 16-bit little-endian, mono channel.
SocketSender: Delivering Generated Audio to Devices
The SocketSender class (defined in src/speech_to_speech/connections/socket_sender.py) listens on port 12346 for outbound connections. This component:
- Monitors
queue_infor generated audio from the TTS module. - Converts NumPy arrays or byte buffers to raw PCM.
- Streams bytes back to the embedded client via the TCP socket.
- Intercepts the AUDIO_RESPONSE_DONE control message to re-enable listening by setting
should_listen, allowing the conversation to continue.
Both classes use the queue abstractions defined in src/speech_to_speech/pipeline/queue_types.py (specifically AudioInItem and AudioOutItem) and filter control messages using utilities from src/speech_to_speech/pipeline/control.py.
Implementing the Embedded Client
Your embedded implementation must open two simultaneous TCP connections to the server and handle non-blocking I/O for real-time performance.
Python Client for Linux Single-Board Computers
For prototyping on Raspberry Pi or similar boards, use this Python implementation with the sounddevice library for audio hardware abstraction:
import socket
import sounddevice as sd
import numpy as np
HOST = "192.168.1.100" # Server IP address
INPUT_PORT = 12345 # To SocketReceiver
OUTPUT_PORT = 12346 # From SocketSender
SAMPLE_RATE = 16000
CHUNK = 512 # Frames per packet
def audio_callback(indata, frames, time, status):
"""Capture microphone and stream to server."""
tx_sock.sendall(indata.tobytes())
# Establish outbound connection (microphone → server)
tx_sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
tx_sock.connect((HOST, INPUT_PORT))
# Establish inbound connection (server → speakers)
rx_sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
rx_sock.connect((HOST, OUTPUT_PORT))
rx_sock.setblocking(False)
# Start microphone stream
stream = sd.InputStream(
samplerate=SAMPLE_RATE,
dtype="int16",
channels=1,
callback=audio_callback,
blocksize=CHUNK,
)
stream.start()
print("Streaming audio. Press Ctrl-C to stop.")
try:
while True:
try:
data = rx_sock.recv(4096)
if data:
audio = np.frombuffer(data, dtype=np.int16).reshape(-1, 1)
sd.play(audio, SAMPLE_RATE, blocking=False)
except BlockingIOError:
pass
except KeyboardInterrupt:
pass
finally:
stream.stop()
tx_sock.close()
rx_sock.close()
This client maintains separate sockets for upstream (microphone) and downstream (speaker) audio, matching the server-side separation of concerns in socket_receiver.py and socket_sender.py.
C Implementation for Embedded Linux
For resource-constrained Linux systems where Python overhead is unacceptable, implement the client in C using POSIX sockets and pthreads:
#include <arpa/inet.h>
#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#define SAMPLE_RATE 16000
#define CHUNK_SAMPLES 512
#define CHUNK_BYTES (CHUNK_SAMPLES * sizeof(int16_t))
int tx_sock, rx_sock;
void *capture_thread(void *arg) {
int16_t buffer[CHUNK_SAMPLES];
while (1) {
// Capture from ADC/microphone driver into buffer
// write() blocks until buffer transmitted
write(tx_sock, buffer, CHUNK_BYTES);
}
return NULL;
}
void *playback_thread(void *arg) {
int16_t buffer[CHUNK_SAMPLES];
ssize_t n;
while ((n = read(rx_sock, buffer, CHUNK_BYTES)) > 0) {
// Send buffer to DAC/speaker driver
}
return NULL;
}
int main(void) {
struct sockaddr_in srv_addr;
// Setup TX socket (to server port 12345)
tx_sock = socket(AF_INET, SOCK_STREAM, 0);
srv_addr.sin_family = AF_INET;
srv_addr.sin_port = htons(12345);
inet_pton(AF_INET, "192.168.1.100", &srv_addr.sin_addr);
connect(tx_sock, (struct sockaddr *)&srv_addr, sizeof(srv_addr));
// Setup RX socket (from server port 12346)
rx_sock = socket(AF_INET, SOCK_STREAM, 0);
srv_addr.sin_port = htons(12346);
connect(rx_sock, (struct sockaddr *)&srv_addr, sizeof(srv_addr));
pthread_t cap_tid, play_tid;
pthread_create(&cap_tid, NULL, capture_thread, NULL);
pthread_create(&play_tid, NULL, playback_thread, NULL);
pthread_join(cap_tid, NULL);
pthread_join(play_tid, NULL);
return 0;
}
This implementation uses standard POSIX APIs available in embedded Linux distributions and RTOS environments with BSD socket support.
MicroPython for ESP-32 Microcontrollers
For bare-metal microcontrollers like the ESP-32, use MicroPython with the I2S peripheral for direct digital audio acquisition:
import network
import socket
import machine
HOST = "192.168.1.100"
INPUT_PORT = 12345
OUTPUT_PORT = 12346
SAMPLE_RATE = 16000
CHUNK_BYTES = 1024 # 512 samples * 2 bytes
# Wi-Fi connection
wlan = network.WLAN(network.STA_IF)
wlan.active(True)
wlan.connect("SSID", "PASSWORD")
while not wlan.isconnected():
pass
# TCP sockets
tx = socket.socket()
tx.connect((HOST, INPUT_PORT))
rx = socket.socket()
rx.connect((HOST, OUTPUT_PORT))
rx.setblocking(False)
# I2S configuration for INMP441 or similar MEMS microphone
i2s = machine.I2S(
0,
sck=machine.Pin(26),
ws=machine.Pin(25),
sd=machine.Pin(22),
mode=machine.I2S.RX,
bits=16,
format=machine.I2S.MONO,
rate=SAMPLE_RATE,
ibuf=CHUNK_BYTES,
)
while True:
# Read from microphone and transmit
buf = bytearray(CHUNK_BYTES)
if i2s.readinto(buf) == CHUNK_BYTES:
tx.send(buf)
# Receive from server and queue for DAC
try:
audio = rx.recv(CHUNK_BYTES)
if audio:
# Push to PWM or I2S DAC
pass
except OSError as e:
if e.args[0] != 11: # EAGAIN
raise
This MicroPython client demonstrates that implementing custom audio streaming for embedded devices requires only socket and I2S primitives, with no external dependencies beyond the standard library.
Key Source Files and Protocol Details
Understanding the internal implementation helps debug connection issues and optimize buffer sizes:
src/speech_to_speech/connections/socket_receiver.py: Implements the server-side listener on port 12345. Manages theshould_listenevent and handles PIPELINE_END control messages when clients disconnect.src/speech_to_speech/connections/socket_sender.py: Implements the outbound streamer on port 12346. Monitors for AUDIO_RESPONSE_DONE to toggle listening states.src/speech_to_speech/pipeline/queue_types.py: DefinesAudioInItemandAudioOutItemtype aliases used for inter-thread communication.src/speech_to_speech/pipeline/messages.py: Contains control constants:PIPELINE_END,AUDIO_RESPONSE_DONE, andSESSION_END.src/speech_to_speech/pipeline/control.py: Providesis_control_message()helper to distinguish audio bytes from control signals.src/speech_to_speech/s2s_pipeline.py: Orchestrates the full pipeline; launches receiver and sender threads based on the--mode socketflag.
Deployment Workflow
To integrate your embedded hardware with the speech-to-speech pipeline:
-
Start the server in socket mode:
python -m speech_to_speech --mode socket \ --stt parakeet-tdt \ --tts qwen3 \ --host 0.0.0.0 -
Connect your embedded client to ports 12345 and 12346 on the server IP.
-
Stream audio: The client transmits 16-bit PCM chunks continuously. The pipeline processes speech through VAD → STT → LLM → TTS.
-
Handle turn-taking: When the server finishes generating audio, it sends the AUDIO_RESPONSE_DONE signal. Your client should resume transmitting microphone data immediately to initiate the next turn.
The server automatically manages session lifecycle through control messages defined in src/speech_to_speech/pipeline/messages.py, ensuring robust handling of network interruptions without pipeline stalls.
Summary
- Socket mode provides the lightest transport for embedded devices, using raw TCP instead of WebSocket protocols.
- Two independent sockets handle bidirectional audio: port 12345 for input (handled by SocketReceiver) and port 12346 for output (handled by SocketSender).
- PCM format requirements are strict: 16 kHz, signed 16-bit little-endian, mono. Any deviation requires resampling on the client.
- Control messages (PIPELINE_END, AUDIO_RESPONSE_DONE, SESSION_END) coordinate conversation flow and session management without additional wire protocol complexity.
- Implementation flexibility allows clients in C, Python, MicroPython, or Rust, provided they can open standard TCP sockets and handle the specified audio format.
Frequently Asked Questions
What audio format does the socket mode require?
The socket transport expects 16 kHz sample rate, signed 16-bit little-endian, mono PCM. Each sample occupies 2 bytes. The default chunk size is 1024 bytes (512 samples), though the SocketReceiver implementation accepts arbitrary packet sizes and buffers them internally.
Can I use WebSocket instead of raw sockets on my microcontroller?
While possible, WebSocket mode introduces unnecessary overhead for embedded targets. Raw sockets require approximately 4KB of RAM for buffering versus 20-50KB for a minimal WebSocket client. Unless you specifically need the OpenAI Realtime protocol features, Socket mode provides lower latency and simpler implementation on constrained hardware.
How does the server coordinate turn-taking in conversations?
The server uses the AUDIO_RESPONSE_DONE control message injected into the pipeline queue. When SocketSender detects this message (defined in src/speech_to_speech/pipeline/messages.py), it triggers the should_listen threading event, which reactivates the VAD module in SocketReceiver. Your client does not need to manage state; simply continue streaming microphone data and the server will naturally interleave listening and speaking phases.
What happens if my embedded client disconnects unexpectedly?
SocketReceiver monitors the TCP connection for closure or read failures. Upon disconnection, it automatically injects a PIPELINE_END message into the queue, which the pipeline control loop interprets as a signal to flush buffers and terminate the current session gracefully. When your client reconnects, the server spawns a new receiver thread and resumes normal operation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →